Assessment design

Measure How Fast They Work With AI, or How Well They Judge It?

When you assess how a candidate works with AI, measure judgment, not speed, and let the clock cap the exercise instead of scoring it. Throughput is honest only where a wrong output meets a check that already runs every time and undoing it costs an edit: a first-draft brief, a throwaway script. Where the output reaches a client, a board or a counterparty, the fastest candidate is the most expensive one. Score the clock only in genuinely paced work, and always alongside a defect rate.

The takeVerification has an accounting problem. The error it prevents never happens, so it lands in nobody's numbers, while the afternoon it costs lands in everybody's. That asymmetry is why speed keeps winning arguments it should lose, and why the fast hire reads as the stronger hire right up until the first unchecked figure surfaces in front of someone who matters. Nobody has priced what a careful colleague quietly absorbs, so the case for judgment stays an argument against speed's number. Make the argument anyway. That number is mostly describing your task.

Where Olive fits

Open a role and see what the work shows

If you build this yourself, the expensive parts are the answer key that says which output should have been refused and the evidence trail behind each call. Olive returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each written by a human reviewer and anchored to a timestamped moment in the session, with completion speed among the things it never marks.

Rank your shortlist

Should you measure speed or judgment?

Measure judgment. Speed earns a score only where a wrong output is caught by a check that already runs every time and undone with an edit: a first-draft brief, a throwaway script, an internal summary someone reads before anyone acts on it. Once a confident wrong answer leaves the building, the fast candidate is the one who costs you the most, and the stopwatch said nothing about that.

Speed and judgment are not two ends of one dial, but under a time limit they trade against each other in exactly one direction. Opening the source behind a claim costs minutes. Recomputing a number costs minutes. Rejecting a draft and restating the brief costs the most of all. A candidate who does all three finishes later than one who accepts the first output and formats it, so if the clock is the score, you have just selected the second person.

The two hiring mistakes are not symmetrical either. A careful hire who is slow is visible inside a month: the work lands late, someone notices, and it can be coached. A fast hire with no verification habit is visible when a client asks where a number came from, and by then it has been in three decks and a board pack. That asymmetry is the whole argument, and it is the same one behind hiring for verification rather than production.

Throughput has a genuine case, worth stating in full before taking it apart. In a controlled experiment, developers given an AI pair programmer completed the task (implementing an HTTP server in JavaScript) 55.8% faster than the control group 1. Notice the shape of that task: well specified, greenfield, with a known-good answer and an obvious way to check it. Where your work looks like that, the person who produces twice as much at the same defect rate is genuinely better, and pretending otherwise is precious.

Run the error-cost test on your own work

Take one wrong output (a plausible, confident, incorrect thing a competent person on your team could ship this week) and answer three questions about it. Who catches it, and does that check run every time or only when someone is already suspicious? How long before someone outside the team acts on it? What does undoing it cost: an edit, a revert, a restatement, or a call to a customer?

Three cheap answers and throughput is the honest metric. One expensive answer and it is judgment, because the error cost is set by the worst path a wrong output can take, not by the average one. Most managers who run this find the expensive answer is question one, and it is usually the same discovery: the check they were counting on does not exist. A test suite and a code review are real checks that run on every commit. "The team lead reads everything" is an intention.

The same test sorts six kinds of work sharply, and not along the lines an org chart would suggest.

WorkThe wrong outputWho catches itUndoing itWhat to weight
Marketing copyAn off-brief draftThe editor, before publishAn editThroughput
Software in a tested repoA wrong patchThe suite and the reviewer, every commitA revertThroughput, if the suite is real
Data analysisA misread result in a deckNobody, unless someone re-runs itAn internal retractionJudgment
Financial analysisA valuation input in a board memoThe board, after the decisionA restatementJudgment
Revenue cycleAn appeal filed on a claim that should have been concededThe payerA missed deadline and a written-off claimJudgment
Legal operationsA position taken on a customer's contractThe counterpartyA renegotiationJudgment

Two rows out of six. That ratio is the reason the default answer is judgment, and it holds inside a single team as well as across an industry: the same marketer whose draft is cheap to fix is expensive the moment the draft carries a statistic, a compliance claim or a competitor comparison. Sort by the artifact, not by the department, and the judgment question stops being a technical-roles question.

Why does speed keep winning the argument?

A stopwatch produces one of the two and not the other. Judgment needs a rubric, a reviewer and an answer key. Time to completion needs a start and an end, so it becomes the number in the debrief, the number in the vendor demo and the number a manager repeats in the hiring meeting, while measuring the task's difficulty at least as much as the person doing it.

It is also the number people are worst at estimating about themselves. In a randomized trial, 16 experienced open-source developers worked 246 real issues drawn from their own repositories, randomized issue by issue to allow or forbid AI tools. With the tools allowed they took 19% longer. They had forecast a 24% speedup beforehand, and after living through the slowdown they still believed the tools had sped them up by 20% 2.

Two things follow for anyone assessing candidates. Never accept a candidate's account of their own speed: the people in that study were experienced, working in code they knew, and they were wrong about themselves in the direction that flatters. And never carry a speed result across task shapes or across years: a 2023 experiment measured a 55.8% speedup on a greenfield task with a known answer 1, and a 2025 study measured a 19% slowdown on mature repositories 2. Two settings, two years, two models: a number that swings that far is describing the conditions it was taken under, not the people.

So use the clock as a constraint instead. State a cap, tell the candidate the cap before they start, and treat finishing inside it as a floor rather than a score. Someone who produces nothing in the time has told you something real. Someone who finishes in half the time has told you they took the assistant's first answer, or that your case was too easy, and both of those are findings about the exercise.

What does good judgment actually look like?

Six acts, every one of them visible in a work session and none of them visible in the finished document: what got framed before anything was generated, which claim got a source demanded, what the person kept instead of handing over, what existed between the brief and the deliverable, what got refused and on what grounds, and what got checked against something outside the conversation.

Each one has an observable and a failure mode, which is what makes them markable rather than impressionistic:

  • Framing. The first message states what is being decided and what would make an answer wrong. Failure: the first message asks for the deliverable.
  • Evidence. A source is demanded for one named claim and actually opened. Failure: claims arrive as facts and nothing is asked for behind them.
  • Delegation. Some work is deliberately kept: a figure recomputed by hand, a judgment made without asking. Failure: everything of consequence goes to the assistant.
  • Structure. A plan, a criteria list or an outline exists before the deliverable and visibly shapes it. Failure: brief in, answer out.
  • Refusal. A direction is rejected with the reason stated, per answer taken up. Failure: every output is accepted and tidied.
  • Verification. A claim is tested outside the chat and the result changes a number, a recommendation or a stated limit. Failure: fluent description treated as fact.

Refusal and verification deserve the most weight, because the failure they guard against is not incompetence. A systematic review of clinical decision support found clinicians over-rode their own correct decisions in favour of erroneous advice in around 6% of cases, and that when the system was wrong it raised the risk of an incorrect decision by 26% 3. Those systems are not the assistants your team uses and the effect sizes will not transfer. Take the mechanism, not the number. A confident, fluent, wrong answer suppresses judgment the person already had, which is why you have to watch someone meet one rather than ask how they would.

Write those six rows down and you have a rubric instead of a vibe. Two reviewers marking the same session on those six rows should reach the same outcome without conferring; when they do not, the rubric is what needs rewriting, not the candidate.

How do you measure both in one exercise?

Give one task with a stated time cap, plant a claim in the source material that the assistant will confidently get wrong, and mark the acts rather than the clock. Time enters as a pass or fail constraint everyone knows about in advance. Judgment enters as a per-behavior outcome marked against an answer key that says which output should have been refused, and why.

The planted claim is the whole instrument. It has to be something the model resolves fluently and the material contradicts: a growth figure a footnote undercuts, a payer policy that does not say what the summary says, a quoted statistic whose source does not support it. A candidate who never opens the material never meets it, and finishes early. That is the case doing its job: the early finish and the missed claim are one finding, not two.

Keep three words for the outcome of each row (demonstrated, partly demonstrated, not demonstrated) and resist converting them into a total. An average across six behaviors is a claim about a person that none of the six rows supports, and it puts the fast-and-shallow candidate level with the slow-and-rigorous one, which is the exact confusion you started with.

If the clock decides anything (a cutoff, a disqualification, a tiebreak), it has become part of a selection procedure, and a selection procedure that screens people out unevenly has to be job-related and consistent with business necessity for the job you are using it on 4. The federal guidelines are specific about what makes a work sample defensible: it is supported by content validity to the extent it is a representative sample of the content of the job, and its manner, setting, level and complexity should closely approximate the work situation 5. A stopwatch on work nobody does under time pressure fails both tests, and a hard timed score is also the piece most likely to need an accommodation. Write down why the cap is the length it is, before a candidate asks.

Where you run the exercise matters less than what you mark, though not by much: a take-home and a live working session put the clock in different places, and an observer makes timing feel decisive whether or not you score it. The same caution applies to vendors. If an assessment reports completion time as a headline metric, ask what that metric was validated against. A validated assessment and a bias-audited one answer different questions, and no speed number answers either. Where an assessment is not on the table at all, the interview questions that show how someone works with AI get you a weaker version of the same signal.

See what gets scored

Common questions

Is time to completion ever worth scoring?

As a cap and a floor, not as a rank. State the limit up front, and treat finishing inside it as the minimum. Score the clock itself only where the job is genuinely paced (a support queue, a trading desk, a newsroom on deadline), and even then measure it alongside defect rate rather than instead of it. Throughput measured without a quality term rewards the person who ships the first output the assistant produced, which is the behavior most likely to cost you money later.

How do I check judgment without building a whole rubric?

Plant one error and ask three questions afterwards. Give the candidate source material containing a claim the assistant will confidently repeat and the material contradicts. Then ask: which claim did you check, how did you check it, and what did you decide not to use? A candidate who found it will answer with specifics in under a minute. A candidate who did not will describe a general process. This is not a validated instrument and should not be a gate on its own, but it costs ten minutes and it separates people.

My team ships marketing content. Should I just measure speed?

Mostly yes, with one carve-out. Draft volume, variant generation and first-pass copy are cheap to catch and cheap to undo, so throughput is honest there. The carve-out is any claim that leaves the building intact: a statistic, a compliance line, a competitor comparison, a customer quote. Those are unrecoverable once published and they are exactly what an assistant produces most fluently. Measure speed on the drafting and judgment on the claims, and make the split explicit in the exercise so you know which one you are looking at.

What if the work is high volume and high stakes at the same time?

Split the queue rather than the metric. Claims handling, revenue cycle and legal intake all have a large cheap tier and a small expensive one, and a single throughput number across both hides the only cases that matter. Decide which items can be worked fast and which route to someone who verifies, then assess for the tier you are hiring into. If one person covers both, judgment is the hiring metric, because the cheap tier can be learned in a week and the expensive one cannot.

Can I just ask candidates how they check the model's output?

You will get a rehearsed answer, because the question telegraphs what a good one sounds like. Ask instead about a specific recent instance: the last time an assistant gave them something wrong, what tipped them off, and what they did next. Vague answers are the tell. Even then you are hearing an account rather than watching the behavior, which is why a work sample with a planted claim outperforms any version of the question.

References

  1. 1. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot Sida Peng, Eirini Kalliamvakou, Peter Cihon, Mert Demirer (arXiv:2302.06590), 2023. arxiv.org In a controlled experiment implementing an HTTP server in JavaScript, the group with an AI pair programmer completed the task 55.8% faster than the control group.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced open-source developers, 246 real issues from their own repositories, randomized to allow or disallow AI tools: 19% longer with the tools allowed, against a forecast 24% speedup and a post-hoc belief that they had been sped up by 20%.
  3. 3. Automation bias: a systematic review of frequency, effect mediators, and mitigators Journal of the American Medical Informatics Association (Goddard, Roudsari, Wyatt), 2012. pmc.ncbi.nlm.nih.gov Clinicians over-rode their own correct decisions in favour of erroneous advice in about 6% of cases; pooled across four studies, incorrect decision-support advice raised the risk of an incorrect decision by 26% (RR 1.26, 95% CI 1.11-1.44).
  4. 4. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Work samples and simulations are selection procedures, and a selection procedure must be job-related and consistent with business necessity.
  5. 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 Legal Information Institute, Cornell Law School, 1978. law.cornell.edu Content validity holds to the extent the procedure is a representative sample of the content of the job, and the manner, setting, level and complexity of a sampled work behavior should closely approximate the work situation.

5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.