Interviewing

First Round Collects the Claim, Final Round Watches the Work

Ask nothing the first round could have asked. Round one settles scope, motivation and whether the claimed experience fits the role, and names one specific project worth testing later. The final round has to produce evidence a conversation cannot, so it should be a piece of the real job done with the candidate's own AI assistant while someone watches the choices, then questions about the choices they just made. If the last round is still a conversation, it is round one at a slower pace.

The takeMost loops differentiate rounds by the seniority of the interviewer and the depth of the probe, which is a seating chart rather than a design. The useful split is the kind of evidence each round can produce. A conversation produces an account, and accounts are cheap to prepare now. Work produces a record: what got handed over, what came back, what the person did about it. Put the expensive round where the record is.

Where Olive fits

Open a role and see what the work shows

The final round is where a conversation runs out of things it can establish. Olive returns the record that round is looking for: an occupational assignment done on the candidate's own clock with an assistant present, read afterwards by a person who writes six findings, each carrying the moment it rests on.

Rank your shortlist

What can the first round actually establish?

Everything a conversation can settle, and nothing beyond it. Round one is for scope and fit: what the role is, what the person has actually done, what they want next, and whether the claimed experience matches the job. It should also produce one thing the rest of the loop uses, which is a specific project worth testing later, named by the candidate rather than picked off a resume.

That is a real job and it is worth doing well. Structured interviews estimate at .42 against .19 for unstructured ones in the 2022 re-analysis of the selection literature, so the distance between two interviews is larger than the distance between an interview and most other methods 1. Same questions, same order, a rating scale everyone shares. A first round run that way is the cheapest reliable evidence a hiring process can buy.

What round one cannot do is settle whether somebody works well with a model, because every answer to that question is now preparable. A candidate can draft, refine and rehearse a story about catching a hallucination the night before, and the interviewer across the table has no way to separate the person who does it from the person who prepared to describe it. That is not a candidate integrity problem. It is what happens when a question has a knowable good answer and the answer is free to obtain.

So use round one to pick the target. Ask what they built or shipped in the last six months that they would be happy to walk through, get the specifics on paper, and stop. Whether the loop needs an AI-fluency round at all is the question this decision sits inside.

What should the final round do instead of asking again?

Produce evidence a conversation cannot. That means work: a slice of the real job, done with the candidate's own assistant, while somebody watches what they choose. The questions in that round are about decisions made in the last forty minutes rather than decisions made two jobs ago, which is the one class of answer nobody can prepare in advance and nobody has to remember accurately.

Watch where the follow-ups can point. Instead of asking "how do you verify AI output", the reviewer asks why the candidate accepted the third paragraph without checking it and rewrote the second one twice. Instead of asking "what would you refuse to delegate", the reviewer asks about the moment the assistant offered to do the whole thing and the candidate either took it or did not. Those questions have no rehearsed answer available, because the event happened in front of both people.

The evidence base supports the format without overselling it. Work samples estimate at .33 in the same 2022 re-analysis, which replaced a .54 figure that had circulated since a 1974 review and that its own lineage abandoned; 53 of the 54 underlying studies tested people already doing the job rather than applicants 1. Read .33 against .42 carefully. The credibility intervals overlap heavily and the paper argues the pattern is coherent, so this is not a demotion of work samples, and none of that literature is about AI-open assessment, because none exists yet.

What the format buys is a different kind of record, which is why redesigning the interview so AI assistance becomes signal changes what the last round can be asked to decide.

Build the final round out of one task and a clock

One task, forty to sixty minutes, the tools the job uses, and a reviewer who watches without taking part. Pick material the role touches this quarter. Put one thing in it that is confidently wrong. Tell the candidate the assistant is allowed and that the decisions matter more than the polish. Then leave them to it and read what happened afterwards.

1. Material from the actual job. A generic puzzle measures puzzle skill. Give an analyst a real filing, a recruiter a real requisition, a revenue-cycle candidate a real denial. Judging an output requires knowing what a normal output looks like here, and generic material removes exactly that. 2. The assistant open. Banning it tests work without AI, which is a different job from the one being filled, and it converts the exercise into a compliance test. 3. A planted error. Something the field would catch and the model will happily build on. This is the part the whole round turns on, and it is the part most take-homes leave out. 4. A reviewer who stays quiet. The reviewer's job is to note what got delegated, what got checked, and where the candidate stopped. Prompting them defeats the measurement.

Run it live or asynchronously depending on what the role is like, and be honest that the two produce different evidence: take-home or live session is a real trade rather than a preference. For engineering roles specifically, running a coding interview with the assistant open covers the mechanics that differ from a written task.

What does this cost, and what does it replace?

Less than adding a round, because it replaces one. Median time-to-fill for nonexecutive roles sits at 44 days, and the median organisation spends 7 of those conducting interviews against 5 screening and 4 deciding 2. Swapping a second conversation for a working session moves no days at all. What changes is what the last hour produces.

The number to keep in view is the spread rather than the median. The interquartile range runs 28 to 73 days and the mean is 54, so a claim that the average hire takes 44 days is wrong twice over. Those are 2021 figures from SHRM member organisations, collected before generative AI reached hiring at all, and they measure calendar days from requisition to offer acceptance.

The same benchmarking data explains why this reads as a bigger change than it is. Work-sample interviews were used by 12%, 11% and 9% of organisations to assess executive, middle management and individual contributor candidates, and simulation exercises by 6%, 6% and 7%, while in-person interviews ran at 79%, 78% and 76% 2. Those are self-reports against SHRM's own definitions, so treat them as an upper bound on structure. A job-shaped final round is still a minority practice, which means most loops are replacing a conversation they already run instead of adding a stage.

There is a limit on how far this format travels. A working session is one piece of evidence rather than the whole decision, and stacking three of them repeats the mistake of stacking three conversations. Two or three genuinely different kinds of evidence is the shape that holds up. Whether that means three rounds or four is a separate question, worked through in how many interview rounds a loop actually needs, and the bar the final round is graded against should already be written down by level, as in the three bars at junior, mid and senior.

See how it works

Common questions

Does the final round still need a hiring manager in the room?

Yes, and their job changes. In a working session the manager is a reader rather than an interviewer: they watch what the candidate delegates, what they check, and where they stop, then ask about those moments in the last fifteen minutes. That is a different skill from running a conversation, and it is worth a short briefing before the first one. A manager who talks during the task turns the exercise into pair work and loses the measurement.

What if the candidate refuses to use AI in the session?

That is allowed, and what they say next is worth hearing. Ask what they did instead, what they checked, and how they would work if the tool were required. Somebody who works in a regulated environment, on air-gapped systems or under a company ban may have excellent verification habits and no model in the story. A refusal only becomes a problem when the role genuinely requires the tool daily, and that is a conversation rather than a mark against them.

Should the final round be paid?

Pay for anything that runs past about an hour or produces work you could use. Forty to sixty minutes is a reasonable unpaid request, and short enough that it selects for skill rather than for availability. Past that, the exercise starts filtering on who has free evenings, which correlates with things no employer wants to select on. Say the length and the payment position in the invitation, not after the candidate accepts.

Can the working session replace the reference check too?

No, because they answer different questions. A session shows how somebody works for an hour on material they have never seen. A reference shows how they worked over months with people who depended on them, including how they behaved when the work went badly. Two or three kinds of evidence that do different work is the combination worth building; substituting one for another because it is newer is how a loop ends up measuring the same thing twice.

How do I keep the final round comparable between candidates?

Same task, same material, same time box, same planted error, and notes recorded against the same four things for everyone: what got delegated, what got checked, what happened at the error, and where the candidate drew a line. Rotate the material only when you have reason to think it has leaked, and rewrite the note template before the swap, never afterwards. Comparability comes from the record being identical in shape, not from the candidates being asked identical follow-ups.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .42 versus .19 structured-unstructured gap, the .33 work-sample estimate, the superseded .54 figure, and the point that 53 of the 54 underlying work-sample studies were concurrent.
  2. 2. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports both the time-to-fill figures (44-day median, 28-to-73 interquartile range, 7 days interviewing) and the adoption figures for work-sample interviews and simulation exercises.

2 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.