Assessment design
How Do You Tell Which Finalist Did the Thinking?
Two finalists with the same quality of work can't be told apart by the deliverables, because both used a model and the output converged. The thinking that separates them is upstream. Send both the identical request within a day: how they framed the problem, which single claim they checked and what the source said, what they ran outside the chat and what changed. Rate the traces against a key you wrote first, cross-check each against the artifact, follow up once on a specific, and never break the tie on speed.
The takeA tie at the finalist stage is a verdict on the exercise, not on the two people in it. A brief the model can satisfy on its own graded the model, and the written trace is a patch applied after the fact. Worth applying. Also worth noticing. On what is public so far, nothing about that convergence loosens as the models get better, so teams still asking finalists for a finished artifact will keep arriving here with two submissions they cannot tell apart. The tie is the brief working exactly as written.
Where Olive fits
Open a role and see what the work shows
If the tiebreak keeps landing on an account of the work rather than a record of it, that gap is what Olive assesses: a 40-to-60-minute occupational assignment done with an AI assistant available, returned as six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session. A human reviewer writes every finding, and the candidate is granted the identical report.
Rank your shortlistWhy the Output Stopped Separating Them
Because the deliverable is now the part a model produces, and both finalists used a model. What used to differ (structure, coverage, polish, tone) converges when two people prompt the same system on the same brief. What still differs is upstream: what each of them decided to ask for, what they refused to accept, and what they went and checked against the world.
The instinct is to ask each finalist how much of it was theirs. That returns an opinion, and the opinion is unreliable in a way that has been measured. METR randomized 246 real issues across 16 experienced open-source developers' own repositories, allowing or forbidding AI tools issue by issue. With the tools allowed the developers took 19% longer, and after finishing they still believed the tools had sped them up by 20% 1. Self-report about AI-assisted work describes an experience rather than an outcome.
The defect you are hunting is not visible in the artifact either, because it does not look like a defect. In Stack Overflow's 2025 developer survey the most common frustration with AI tools, at 66%, was output that is almost right but not quite; 45% said debugging AI-generated code is more time-consuming 2. Almost-right compiles, reads well, and renders. It fails on the one number, the one edge case, or the one claim the recommendation rests on, and it fails after you have hired.
So two finalists can hand in work that is almost-right in different places, or one of them can hand in work whose thin spots were found and closed before submission. Nothing on the page tells you which. The same problem in a different format is worked through in How to Tell Who Did the Work in an AI-Built Portfolio.
Ask for Three Process Traces, in Writing
Three, sent to both finalists as the identical message, answered in under 300 words each, within a day. The framing: what were you deciding, and what would have made an answer wrong? The evidence: which single claim did the work rest on, and what did you open to check it? The outside check: what did you run, recompute or read that the assistant could not have supplied?
- The framing. A candidate who framed the problem first can name the decision and the failure condition in two sentences, because they wrote them down before generating anything. A candidate who asked the model for the deliverable will describe the deliverable back to you. The tell is a constraint or a criterion named specifically enough to be wrong: "the recommendation flips if the churn figure includes trials, and the definition note says it does not."
- The claim they demanded evidence for. Not the reference list: the one claim the whole thing rests on, and what happened when they went to check it. A generated draft arrives already carrying plausible citations, so "I cited five sources" separates nobody. "I opened the one that mattered and it did not say what the draft claimed" separates immediately. Ask which source, what it actually said, and whether it changed anything.
- The outside check. This is the trace hardest to invent, because it has a result, and the result could have come back wrong. A survey of 319 knowledge workers describing 936 real tasks found the effort does not disappear when a tool is used; it moves toward verifying information, integrating the response and stewarding the task. Higher confidence in the tool tracked with less critical thinking, higher confidence in one's own ability with more 3.
Rate the three separately and never blend them into an impression. A strong framing answer should not carry a missing check across the line. The anchoring problem is the same one covered in How to Score an Interview Answer Produced With AI.
What Counts as a Real Outside Check in Your Field
It varies, and that is the point: a check that means something for an engineer means nothing for a marketer. The common shape is an act with a result the assistant did not supply, that could have come back wrong. A test that fails. A source that says something else. A number that does not reconcile. Ask for the result, including the case where the check found nothing.
- Software. A test that could fail, run against the actual repository rather than the snippet. Ask which case they wrote a test for that the assistant had not covered, and what happened the first time it ran. Happy-path tests written alongside the code only prove the code agrees with itself.
- Financial analysis or accounting. One number recomputed from the source document rather than from the summary. Ask for the figure, the page it came from, and whether the recomputation matched. A model reproduces a number fluently and never opens the workbook.
- Consulting or market sizing. A second method that should land in the same order of magnitude. Ask for the back-of-envelope run beside the main estimate, and what they did when the two came out a factor of three apart.
- Anything citing authority. The cited source opened and read. Purpose-built legal research systems still return incorrect information. Stanford researchers measured Lexis+ AI and Ask Practical Law AI at more than 17% of benchmark queries and Westlaw's AI-Assisted Research at more than 34% 4. A citation that exists is not a citation that says what the draft claims.
- Marketing or research. The most on-message statistic in the brief, traced to its source, sample and year. It is the one least likely to have been opened, precisely because it fits.
- Product or design. One assumption tested against something that answers back (a real ticket queue, a real analytics number, a real user) rather than against a persona the assistant wrote.
The question underneath is identical everywhere and only the instrument moves: what did this person do that could have proved them wrong? How to Hire for Verification When AI Does the Production works through the same shift at the level of the role itself.
How to Run the Tiebreak Without Adding a Round
Send both finalists the identical message on the same day, cap the reply at 300 words, and write your answer key before either reply lands. Rate the three traces item by item, across both candidates, before you re-open either deliverable. That is fifteen minutes of your time and twenty of theirs, and it is one documented, consistent step rather than a fourth interview nobody has calendar space for.
Consistency is not housekeeping here. A step that decides who gets hired is a selection procedure, and federal guidance on employment tests is explicit that a procedure must be job-related, consistent with business necessity, and applied the same way to everyone it touches 5. Two finalists asked two different questions produce two answers you cannot compare and one decision you cannot defend.
Structure also buys accuracy. The most recent re-estimation of selection-method validity puts structured interviews at .42 and work samples at .33 against job performance, with job-specific measures outperforming general ones 6. A side-by-side read of two finished artifacts is neither structured nor job-specific; it is a preference, formed in the first two minutes and defended afterwards.
- Write the key first. Two or three lines per trace describing what a weak, an adequate and a strong answer contains for this specific brief. A key written after the first reply is a description of that reply.
- Rate across, not down. Both framing answers, then both evidence answers, then both checks. Reading one reply end to end is how a strong opening carries a missing check through the whole rubric.
- Follow up once, on a specific. One question each, about a detail inside their own answer. Specificity that survives a follow-up is not borrowable, and the mechanics are in Follow-Up Questions That Expose Real Understanding.
Tell both finalists the request is rated and roughly how long it should take. Undisclosed criteria produce noise: half of the replies get treated as paperwork, and you end up rating that guess instead of their judgment.
What a Written Trace Can't Tell You
Whether it happened. A written trace is an account of the work, produced after the fact, by a candidate who knows it is being rated, with the same assistant still open. It is better evidence than the deliverable (it can be checked against the deliverable, which can be checked against nothing), but it is a report of a process rather than a record of one.
Three things narrow the gap. Ask within a day, while the specifics are still recoverable rather than reconstructed. Cross-check every claim against the artifact: an assumption the deliverable contradicts, or a source check the numbers do not reflect, is a finding in itself. And ask that one follow-up, because a borrowed answer holds at the first level of detail and comes apart at the second.
If that still leaves you guessing between two people you are about to bet a year of payroll on, the honest conclusion is that the exercise has to happen somewhere the decisions are visible while they are being made. Take-Home or Live Working Session? covers that trade, and Should You Run a Paid Trial Instead of an Assessment? covers the version where you pay for a week of real work and watch it happen.
Speed will offer itself as the tiebreak. Whoever produced the deliverable faster with an assistant has told you about their throughput, not their judgment, and throughput is the part that got cheap. Measure How Fast They Work With AI, or How Well They Judge It? has the argument in full.
Common questions
Can I just ask which finalist used AI more?
No. Both used it, and the amount tells you nothing about the judgment applied. A candidate who decided the model was the wrong instrument for one step and did that step by hand made a better call than one who generated everything twice as fast. Ask what they kept and why, not how much they handed over. The answer to "how much" is also the answer least likely to be accurate, since people misjudge their own AI-assisted performance in a measurable direction.
What if one finalist writes better than the other?
Rate the content of each trace, not its prose, and write the key in advance so you notice when you are drifting. Prose quality is now the cheapest thing in the reply. It is what the assistant supplies for free. Score the framing on whether a decision and a failure condition are named, the evidence trace on whether a specific source was opened and what it returned, and the check on whether a result exists that could have come back wrong.
Is it fair to ask this of finalists only?
Yes, provided both finalists get the identical request at the same stage, with the same cap and the same deadline. Selection steps are compared within a stage, not across the whole funnel, so a step applied consistently to everyone who reaches it is defensible. What is not defensible is asking one of them a harder question, extending one deadline, or improvising a second request after reading the first reply.
What if both traces come back strong?
Then you have two good hires and no tie to break with more evidence. Stop adding rounds. Decide on the thing you were always going to decide on: which of them fits the work the role actually needs in the next six months. Say so plainly to both. Manufacturing a fourth signal to justify a choice you have already made costs you a finalist and teaches you nothing.
How long should the request take a candidate?
Twenty minutes, and say so in the message. Three traces at under 300 words each is a recall exercise, not new work, and any longer means they are writing rather than remembering. Cap it explicitly. An uncapped request quietly asks finalists to spend an evening on an unpaid task, and the ones who comply are not the ones you learn most about.
References
- 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues from their own repositories, randomized to allow or forbid AI tools, took 19% longer with the tools; afterwards they still believed the tools had sped them up by 20%.
- 2. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co The biggest reported frustration with AI tools, at 66%, is output that is almost right but not quite; 45% say debugging AI-generated code is more time-consuming.
- 3. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org 319 knowledge workers describing 936 real tasks: effort shifts toward information verification, response integration and task stewardship; higher confidence in the tool is associated with less critical thinking, higher confidence in one's own ability with more.
- 4. AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries ✓ hai.stanford.edu Purpose-built legal research systems still return incorrect information: Lexis+ AI and Ask Practical Law AI on more than 17% of benchmark queries, Westlaw's AI-Assisted Research on more than 34%.
- 5. Employment Tests and Selection Procedures ✓ eeoc.gov Employment tests and other selection procedures must be job-related and consistent with business necessity, and applied consistently across the candidates they touch.
- 6. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors ✓ doi.org Revised operational validity estimates: structured interviews .42, job knowledge tests .40, work samples .33, general mental ability .31, with job-specific measures outperforming general ones.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.