Assessment design

How Do You Hire Someone Who Can Tell When AI Is Wrong?

Test the checking, not the production. Hand the candidate a real task from the field you're hiring for, an assistant that will do all of it, and source material with something wrong in it that only checking catches. Grade four acts from the record: which claim they checked, what against, whether before drafting or after, and what changed in the deliverable. Write it per field, because verification doesn't transfer. Seed two problems rather than one, and budget a second case, because the packet leaks.

The takeWhat the exercise really measures is whether someone goes outside the conversation when nothing forces them to. That's a disposition rather than a credential, and it is the one thing a resume from a strong school has never reported. Which is also why the shortage reads as a talent problem and isn't one: the habit is cheap to teach and nearly impossible to interview for. I suspect most teams will keep hiring on output for another year and absorb the difference in review time, which is the more expensive way to buy the same thing.

Where Olive fits

Open a role and see what the work shows

Building this yourself, the expensive parts are the authored case and the evidence behind each judgment, per occupation. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, and verification, which asks whether a claim was tested against something outside the conversation and whether the result changed anything), each anchored to a moment in the session and granted to the candidate as well.

Rank your shortlist

Why can't your interns tell when the output is wrong?

Because verification was never the part of the job they were hired into. Production used to be the apprenticeship. You wrote the draft, the memo, the query, and someone senior checked it. An assistant does that first pass now, so the entry-level work that remains is the checking, and nobody has practiced it. The skill didn't disappear; it moved to the front of a career that used to end with it.

That shift is measured rather than anecdotal. Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 first-hand examples of using generative AI at work, and found the nature of critical thinking moving toward information verification, response integration and task stewardship, with higher confidence in the tool associated with less critical thinking, and higher confidence in oneself associated with more 1.

The hiring side is measured too. Brynjolfsson, Chandar and Chen find employment for 22-to-25-year-olds in the most AI-exposed occupations sitting 19% below where it would be had it kept pace with their less-exposed peers, with the adjustment running through reduced hiring rather than separations, and no comparable gap among experienced workers 2. The entry-level seat is being priced on something other than output, which is a reason to test something other than output.

And the errors are ordinary, not exotic. In Stack Overflow's 2025 developer survey the most-cited frustration with AI tools was "AI solutions that are almost right, but not quite" at 66%, and more developers actively distrusted the accuracy of AI output (46%) than trusted it (33%) 3. Almost-right is the hard case: it survives a read-through and fails downstream, and reading carefully is exactly what a strong intern has been trained to do. What a new grad who has never worked without AI is actually missing is not tool skill. It's the habit of going outside the conversation.

What does a verification exercise look like in your field?

It looks like the check your best people already run before they sign anything, staged so you can watch someone do it. A financial analyst reconciles to the filing. An engineer runs the test. A marketer traces the statistic back to the sample it came from. Same rubric, different instrument, and an exercise written for one field measures almost nothing in another.

Six versions, with what counts as checking spelled out:

  • Financial analysis. The packet's summary page carries a segment growth figure the filing behind it does not support. Checking means opening the 10-K and re-adding the segment, not asking the assistant whether the number looks right.
  • Software engineering. The docstring and the shipped signature disagree about argument order, and the fixture uses equal values so the tests pass. Checking means running it against a case the fixture doesn't cover.
  • Marketing. The most on-message statistic in the research folder comes from a vendor's own 42-person customer survey and is written up as market-wide. Checking means finding the sample before the claim reaches the brief.
  • Legal operations. A memo cites a real case with a pin cite to a paragraph that says something else. The citation resolves; only reading the paragraph settles it. Stanford researchers benchmarking purpose-built legal research tools found Lexis+ AI and Ask Practical Law AI returning incorrect information more than 17% of the time, and Westlaw's AI-Assisted Research more than 34% 4. Retrieval is not verification.
  • Data and analytics. A fluent explanation of a result the dataset does not support. Checking means re-running the cut, not restating the explanation more carefully.
  • Healthcare revenue cycle. A denial coded as a coverage dispute that is really a timely-filing miss. Checking means opening the remit and the submission date before drafting an appeal that cannot win.

Write one, then calibrate it against two people already doing the job before it decides anything. One of them should find the problem inside ten minutes, and neither should call it a trick. That calibration is also the honest budget line: one authored case per role, plus the second you'll need when the first leaks. There is no generic version, because a generic online assessment is the exact thing the assistant has already solved.

How do you grade a check?

Four columns, filled from the record rather than from impression: which claim was checked, what it was checked against, when (before the deliverable was drafted or after), and what in the final answer is different because of it. Two graders should fill those in separately and agree. "Showed good judgment" cannot be scored twice the same way; "re-added the segment against the filing" is a yes or a no.

Grade selection hardest. Checking everything is unavailable in real work, so the question is which claim they picked and whether it was the one the recommendation rested on. A candidate who verified three peripheral facts and shipped the load-bearing one unchecked did worse than one who checked only the load-bearing claim and said so in the deliverable. Set the time limit tight enough that triage is forced, then score the triage. See how Olive measures this.

Do not substitute a conversation for the record. METR ran a randomized trial with 16 experienced open-source developers on 246 real issues in repositories they knew well; with AI tools allowed the work took 19% longer, and afterwards the developers still believed the tools had sped them up by 20% 5. If people inside the work cannot feel the difference between 19% slower and 20% faster, then "tell me how you check AI output" is a story rather than evidence.

There is a compliance argument for building it as work instead of as questions. The Uniform Guidelines hold that content validity runs to the extent the procedure is a representative sample of the content of the job, and that a procedure resting on inferences about mental processes cannot be supported by content validity alone 6. Asking someone how they would check a confident claim is an inference about a mental process. Handing them the claim, the source and thirty minutes is a work sample. A rubric written for AI-assisted answers takes an hour and outlives every task you swap in behind it.

What should you stop testing?

Output speed, tool inventories and prompt technique. A candidate who lists six assistants has told you about their browser tabs. A timed production task now measures whose model was faster. And the volume of AI a person used is not a virtue: someone who judged the assistant was the wrong instrument for a step and did it by hand has demonstrated the thing you are hiring for.

Multiple-choice AI-literacy tests fail from the other end. They ask about verification in the abstract, where the right answer is obvious and costs nothing. Nobody picks "ship it unchecked" on a quiz. The price of checking is the whole test: the twenty minutes, the file that has to be opened, the deadline sitting on top of both. A quiz removes the price and then reports that everyone is willing to pay it.

Separate the hiring question from the training one before you spend a loop on it. Verification is teachable, often faster than managers expect, and the choice between training the habit and hiring for it turns on how long your review loop can carry someone who cannot yet do it. Hire for it where a wrong answer ships without a second reader; train it where a senior person sees everything anyway.

What does this exercise miss?

One planted problem is one observation, and a competent person can miss one thing on a bad Tuesday. It tests the error you thought to plant, in the field you know well enough to plant it in. A candidate can catch your seeded arithmetic and still hand every framing decision to the model on Monday. Two problems of different kinds in the same task roughly halve the coin flip and cost one extra afternoon to author.

It also leaks. A dozen candidates in, the packet is posted somewhere, and a task everyone has seen measures recall. Budget the second case at the start rather than discovering the need at candidate thirteen, and keep one rubric across both: the rubric transfers, the task never does.

Unwatched, you see the deliverable and not the checking, which is the entire signal. Ask for the check explicitly: a short note naming what was verified, against what, and what it showed. That is still a claim about an act rather than the act, but it beats grading a clean document. Running the task live fixes the record and costs you the time verification actually takes. Whether candidates get AI help during a live interview is a separate problem, and solving it by compressing the task defeats this one.

And a work sample tests one person on one afternoon, watched. It says nothing about whether they keep checking in month four, when nobody is looking and the deadline is real. No hiring process reaches that. The review loop does, and building one that expects unchecked output is the other half of this problem.

See what gets scored

Common questions

How long should a verification exercise take?

Long enough that checking has a cost. Under thirty minutes, triage never bites and everyone verifies everything; past ninety, completion drops and you select for free time rather than judgment. Forty-five to sixty minutes, with a task slightly too large for it, is the shape that forces a real choice about what to check. Say the cap in the invite and mean it, and pay for anything longer. An unpaid two-hour task filters for candidates who can afford two hours.

Should you tell candidates the material contains an error?

Tell them the source material may contain mistakes, and that flagging what they could not settle is part of the deliverable. Do not say how many or where. Saying nothing produces two populations (the ones who assumed the packet was clean and the ones who assumed a trap), and their results are not comparable. Saying exactly where turns it into a scavenger hunt. The middle position is also how the job works: real material has errors in it, and nothing is labeled.

What if the candidate checks nothing and the deliverable is still good?

Then you learned the thing you set out to learn. A clean-looking deliverable produced with no check is the failure the exercise exists to surface, and it isn't rescued by the output happening to be fine. It was fine because the assistant was right this time. Score the record, not the artifact. In the debrief, ask what they would have had to see to change the recommendation. A candidate with no answer to that was never going to catch anything.

Can one exercise cover every role you hire for?

No. What counts as checking is set by the field: the filing for an analyst, a test run for an engineer, the underlying sample for a marketer, the paragraph behind the pin cite for a paralegal. Hand an engineer the market-sizing packet and you have measured reading comprehension. The rubric (which claim, what instrument, when, what changed) carries across all of them unedited. The task never does, and that authoring cost is why most teams stop at a generic quiz.

Is this fair to someone who has never held a job?

It is fairer than the resume, which for an early-career candidate mostly reports where they went to school. A verification exercise asks for a behavior rather than a history, and the behavior is learnable, so someone with no relevant employment can demonstrate it in an afternoon. Two conditions make it fair in practice: the field knowledge the task requires has to be teachable inside the task itself, and the accommodation path has to be real for candidates who need one.

Does this replace the interview or sit alongside it?

Alongside, and it usually replaces one round rather than adding one. The exercise produces a record; the interview is where you read it back: which claim they picked, what they left unchecked, what they would have done with more time. That conversation is worth more than any standalone question about AI use, because it is anchored to something that happened. Adding a round without removing one is how loops reach six stages and candidates stop finishing them.

References

  1. 1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers Microsoft Research and Carnegie Mellon University (Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks, Wilson), CHI 2025, 2025. microsoft.com 319 knowledge workers and 936 first-hand examples: the nature of critical thinking shifts toward information verification, response integration and task stewardship, and higher confidence in the tool goes with less critical thinking.
  2. 2. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence Stanford Digital Economy Lab (Brynjolfsson, Chandar, Chen), 2025. digitaleconomy.stanford.edu Employment for 22-to-25-year-olds in the most AI-exposed occupations sits 19% below where it would be had it kept pace with less-exposed peers, driven by reduced hiring rather than separations, with no comparable gap for experienced workers.
  3. 3. Stack Overflow Developer Survey 2025: AI Stack Overflow, 2025. survey.stackoverflow.co The top-reported frustration with AI tools is "AI solutions that are almost right, but not quite" at 66%; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  4. 4. AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries Stanford HAI / Stanford RegLab (Magesh, Surani, Dahl, Suzgun, Manning, Ho), 2024. hai.stanford.edu Purpose-built legal research systems still return incorrect information: Lexis+ AI and Ask Practical Law AI more than 17% of the time, Westlaw's AI-Assisted Research more than 34%.
  5. 5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers on 246 real issues took 19% longer with AI tools allowed, and still believed afterwards that the tools had sped them up by 20%.
  6. 6. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Content validity holds to the extent the procedure is a representative sample of the content of the job; a procedure resting on inferences about mental processes cannot be supported by content validity alone.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.