Assessment design
How Do You Assess a Candidate Who Mostly Supervises AI Agents?
A role that mostly supervises AI agents is assessed with a review exercise, not a build: hand the candidate an agent run that looks finished, with three or four defects seeded in and the sources that settle them, then grade the review they hand back. Four columns: which defects got named, what got checked and against what, what got rejected on a ground stated in one sentence, what got escalated instead of shipped. Same packet for every candidate, 45 to 60 minutes. No single number stands for the person.
The takeThe exercise is the cheap half of this hire. The expensive half is whether saying no costs the person anything, and on most teams it does: an escalation lands on a dashboard as slowness, while a queue that clears with nobody stopping it reads as a good week. The AI Act pairs competence with authority in one sentence, and no work sample grants the second. The reviewer who caught every seeded defect in the exercise is probably the one waving runs through a year later, because that is what the job teaches them stopping costs. The rubric is not what makes a supervisor. The room around it is.
Where Olive fits
Open a role and see what the work shows
Two of the six things Olive reads from a session are the acts this exercise is built to observe: output rejection, counted per answer actually taken up and only where a reason is stated, and verification, meaning a claim tested outside the conversation rather than a claim of one. A human reviewer writes each finding against the moment it rests on, and the candidate is granted the same report.
Rank your shortlistWhat does an agent supervisor actually do all day?
Four acts, and none of them is production. You write the brief and the limits the agent runs under. You read what comes back against something other than itself. You reject specific parts on stated grounds. You decide what never ships without a second human. Assess those four, in that order, because they are what the person does on an ordinary Tuesday.
Most exercises available today test a fifth thing: whether the candidate can plan, implement, debug and architect an agent. That is the builder's job. If your hire will never open the agent's code, an architecture exercise tells you about someone else, and the person who would have been excellent at the actual work fails it for the wrong reason. Write down what the agents in this role already do, step by step, before you write the exercise. Working out what AI actually does in the role is the same first move you would make before writing the job post.
The volume is real enough to hire against. In Stack Overflow's 2025 developer survey, 14.1% of respondents said they use AI agents at work daily and a further 9% weekly, out of 31,877 answering that question 5.
The failure this job exists to prevent has a name and a literature behind it. A systematic review of automation bias describes users over-accepting computer output as "a heuristic replacement of vigilant information seeking and processing"; a meta-analysis of four clinical decision-support studies inside that review found erroneous advice was more likely to be followed in the decision-support groups than in the controls, and that a system in error raised the risk of an incorrect decision being made by 26% 3. The tool is new. The way people defer to it is not.
Supervision is also turning into a named qualification rather than an implied one. Under Regulation (EU) 2024/1689, the EU's AI Act, deployers of high-risk systems "shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support" (Article 26(2)) 6. Which systems are in scope, and from when, is a question for counsel. The hiring consequence is the part worth acting on now: authority is granted, not brought to the interview. If the role cannot stop a run without three approvals, no assessment fixes that, and you are testing someone you have already set up to lose.
Build the exercise from one agent run that already looks finished
Give the candidate the four things the job gives them: the original request, the brief the agent was handed, the agent's full run (its steps, its intermediate claims, its finished deliverable), and the source material that settles whether any of it is true. Seed three or four real defects into the run. What they hand back is the review, not a tidier version of the work.
Seed by class, not by difficulty. Each of these is something an agent does that reads as competence:
- A claim sourced to a real document that does not say it. The link resolves. The paragraph behind it says something adjacent.
- A step reported as done that was not done. The run log says the file was validated; nothing in the trace validates anything.
- A correct answer to a slightly wrong question. The brief asked for last quarter and the agent answered for the trailing year, competently.
- A number right in its arithmetic and wrong in its base. The growth rate is computed cleanly off the wrong denominator.
- An act that should have stopped for a human. A refund issued, a record amended, an external email drafted and queued.
Cap it at 45 to 60 minutes in one sitting, and hand every candidate the identical packet so the round stays comparable. Then ask three questions in writing at the end: what did you check and against what, what are you rejecting and on what ground, and what would you not ship without another person looking at it.
Add the specification half if you have ten more minutes, because it is where the strongest candidates separate: rewrite the brief so this run would have failed loudly instead of quietly. A supervisor who can only find defects after the fact is doing half the job. One who writes "cite the page number, and stop if a source cannot be opened" into the brief has removed a class of failure from every future run. Testing whether a candidate notices when AI gets something wrong is the review half; the brief they write is the other one.
What counts as a credible check in your field?
One act that touches something outside the agent's own output. Reading the deliverable more carefully is not a check, and asking the agent whether it is confident is not one either. Name the act your field accepts before you write the rubric, because it decides what a reviewer can mark down as done. It varies more than people expect:
- Software engineering. The branch checked out and the test run against the real fixture, not the one the agent wrote to pass.
- Financial analysis. The figure re-derived from the filing rather than from the agent's summary of the filing.
- Legal operations. The case opened and the paragraph read. Leading AI legal research tools were measured hallucinating between 17% and 33% of the time, and a citation that resolves is not a citation that supports 4.
- Marketing. The statistic traced back to its sample, until the candidate can state the n and who was surveyed.
- Healthcare revenue cycle. The denial code read against the payer's own published policy, not against the agent's account of it.
- Supply chain. Three quotes re-normalized onto one set of terms before any of them are compared.
- Data and analytics. The query re-run with the filter the agent described, and the row count compared against what the write-up claims.
Two shortcuts will be offered to you, and both fail. The first is the agent's own reasoning. In a controlled study of AI-assisted decisions, explanations increased the chance that people accepted the recommendation regardless of whether it was correct 2, so a fluent rationale attached to a wrong answer is worse than a bare wrong answer, not better.
The second is the candidate's own sense of how the run went. In METR's randomized trial, 16 experienced open-source developers working across 246 real issues took 19% longer to complete them when allowed AI tools; they had forecast a 24% speedup beforehand, and still believed they had been sped up by 20% afterward 1. Self-report ran in the opposite direction to the stopwatch. This is the same reason hiring for verification rather than production has become its own problem: producing has gotten cheap, and knowing whether the product is right has not.
How do you score it so two reviewers agree?
Score observable acts, one column each, marked separately by two people. Which seeded defects were named. Which rejections carry a ground you could restate in one sentence. Which checks were actually run, and against what. What was escalated instead of shipped. No single number for the person: the columns are what you compare across candidates, and the columns are what you can defend later.
A ground you can restate is the whole test. "Rejected the market-size paragraph because the only source behind it is a vendor's own 40-person customer survey" is a yes. "Showed good judgment" is not a column, it is an impression, and two reviewers will fill it differently every time. Writing a rubric two reviewers score the same way is mostly this one discipline applied to every row.
Have both reviewers fill the grid before they talk. Where they disagree, the usual cause is a badly written column rather than an ambiguous candidate, so rewrite the column and re-mark. Accountability appears to matter here in a way that is worth building into the process: the automation-bias review found that external manipulations of accountability gave mixed results, while people who perceived themselves as accountable made fewer automation-bias errors 3. The cheap version of that is a name attached to every pass, on both sides of the table.
One trap in the marking. An over-refusing candidate is not the safe answer. Someone who blocks every run costs exactly the throughput the agent was bought for, and refusal without a ground is a different way of not reading. Score refusals that name what is wrong, and count them against the parts of the run actually taken up, so three well-aimed rejections beat thirty reflexive ones.
What this exercise still won't tell you
Whether they are still checking in month seven. A seeded defect is easier to find than a real one, because the candidate already knows something is wrong somewhere; the job's actual difficulty is staying alert across a long line of runs that are mostly fine. One exercise measures an afternoon of attention. Say that out loud in the debrief rather than treating one packet as a character reading.
The packet leaks, too. A dozen candidates in, it will be sitting in a group chat somewhere and the exercise starts measuring recall. Build the second case at the same time as the first, keep one rubric across both, and rotate when the round gets busy.
Tell the candidate what is being assessed before they start, in one short paragraph: the run has defects in it, the review is the deliverable, and the rubric marks checks, grounds and escalations. Nothing here needs to be a trick. A candidate who knows the review is the work does the job better, which is the point of the exercise, and a hidden rubric only measures who guessed your intent.
Two things not to score, because they will pull the grade off the job. Volume of AI use is not a virtue: a candidate who did one step by hand because the agent was the wrong instrument for it made a supervision decision, and it should count as one. And a slow candidate who ran three real checks has told you more than a fast one who skimmed and approved. Speed is a legitimate business constraint, but it belongs in the time cap, not in the rubric.
Common questions
Is this the same as hiring someone to build AI agents?
No, and the two exercises look nothing alike. A builder is assessed on planning, implementation, debugging and architecture. A supervisor is assessed on the brief they write, the checks they run against the output, the rejections they can ground in a sentence, and what they escalate rather than ship. Most published agent exercises score the builder, because that role came first. If your hire will never open the agent's code, an architecture test fails the right people for the wrong reason.
How long should the exercise take?
Forty-five to sixty minutes, in one sitting, with the same packet for every candidate. That is long enough for three or four seeded defects and a written review, and short enough that completion does not collapse. If you want the specification half as well, add ten minutes for rewriting the brief rather than extending the review. Anything past about ninety minutes stops being a work sample and starts being unpaid work, and the candidates with current jobs are the ones who drop.
What do you seed into the agent run?
Three or four defects, each from a different class: a claim sourced to a real document that does not say it, a step reported as done that was not done, a correct answer to a slightly wrong question, a number computed cleanly off the wrong base, and an action that should have stopped for a human. Keep them plausible. A defect that announces itself measures reading speed, and a defect no reviewer can settle from the source packet measures luck.
Can you run this for a non-technical role?
Yes, and the shape does not change. What changes is the credible check. In marketing it is the statistic traced back to its sample size and population. In revenue cycle it is the denial code read against the payer's published policy. In legal operations it is the case opened and the paragraph read. In supply chain it is three quotes re-normalized onto one set of terms. Name that act for your field first, then write the rubric around it.
Should the candidate run the agent themselves during the exercise?
For the specification half, yes if you can arrange it. For the review half, no. A fixed artifact is what makes the round comparable: if every candidate generates their own run, they are each reviewing different work and the marks stop meaning the same thing. The practical split is a supplied run for the graded review, and their own brief for the part that shows how they would have set the run up.
What does it mean if the candidate catches nothing?
It means they did not check, which is the finding. Ask before you conclude anything harsher: some candidates read the packet as a proofreading task because the instructions were vague about what the deliverable was. That is a fixable instruction problem, and it is worth ruling out on the first few runs. Where the instructions were clear and the review still came back approving everything with a tidy summary, you have watched the job's central failure happen once, cheaply.
References
- 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org Randomized trial with 16 experienced open-source developers across 246 issues: they took 19% longer when allowed AI tools, having forecast a 24% speedup, and still believed afterward that they had been sped up by 20%. Supports the claim that a supervisor's own sense of how a run went is not a measurement.
- 2. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance ✓ arxiv.org "Explanations increased the chance that humans will accept the AI's recommendation, regardless of its correctness." Supports the claim that an agent's own stated reasoning is not a check on its output.
- 3. Automation bias: a systematic review of frequency, effect mediators, and mitigators ✓ pmc.ncbi.nlm.nih.gov Defines automation bias as over-accepting computer output as "a heuristic replacement of vigilant information seeking and processing"; a meta-analysis of four clinical decision-support studies found erroneous advice more likely to be followed in the decision-support groups than in controls, with an in-error system raising the risk of an incorrect decision by 26%; on mitigators, external manipulations of accountability gave mixed results while people who perceived themselves as accountable made fewer automation-bias errors.
- 4. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools ✓ arxiv.org Measured the legal research tools from LexisNexis and Thomson Reuters hallucinating "between 17% and 33% of the time". Supports the legal-operations credible check: a citation that resolves is not a citation that supports.
- 5. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co Of 31,877 respondents answering the AI agents question, 14.1% said they use AI agents at work daily and 9% weekly. Supports the claim that supervising agent output is ordinary enough work to hire against.
- 6. Article 26: Obligations of Deployers of High-Risk AI Systems, Regulation (EU) 2024/1689 (Artificial Intelligence Act) ✓ artificialintelligenceact.eu Article 26(2): "Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support." Supports the claim that oversight competence and authority are named obligations in EU law, quoted without any assertion about which systems are in scope or when.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.