Interviewing

Cut the Round That Has Never Changed a Decision

Cut the interview round that has never changed an outcome, not the longest one or the one candidates complain about. For each round, ask how many hires or rejects in the last year would have gone the other way without it. Where the answer is none, that round is confirming a decision already made. Before removing it, name the evidence it produced that no other round produces. You know the cut was right when the next four reqs reach the same decisions and the debriefs get no harder.

The takeThe reduce-your-rounds genre selects on price: drop the longest, drop the fifth, drop whatever the candidate survey called excessive. That reliably removes the wrong one. The round most likely to be dead is the one built around an artifact whose signal moved, an async video answer or a generic behavioural hour, while the round most likely to be load-bearing is the expensive live one that gets cut first precisely because it is expensive. Cost and information point in opposite directions here.

Where Olive fits

Open a role and see what the work shows

Olive returns six findings on one candidate, each written by a person and anchored to a timestamped moment in a 50-to-70-minute session. It is an input a debrief can argue with, and the candidate receives the identical report on every tier.

Rank your shortlist

Which interview round should you actually cut?

The one whose verdict has never differed from the round before it. That test beats length, cost or candidate complaint, because it asks what a round contributed rather than what it consumed. In most loops it points somewhere specific: the async video answer, the generic behavioural hour, the second conversation with a different job title sitting in the chair.

The funnel data shows the same structure from the other end of the loop. In one large applicant-tracking dataset covering more than 54 million applications and 93,000 jobs, recruiter screens pass roughly 35% of the candidates who reach them while later stages convert far higher, at 95% post-onsite and 81% at offer 1. Those are stage-to-stage rates among candidates who reached each stage, not an end-to-end funnel, stage names are configured per customer, and the heaviest filtering happens earlier still at application review, which is not one of these numbers.

The 95% is the figure worth sitting with. At that rate the onsite is doing very little independent filtering: the decision has effectively been made before it ends, and the late stages ratify a call already taken. That is a finding to act on, and it says nothing about how demanding an onsite is. The action is not automatically to delete the onsite, because it may be redundant or it may be the round everything else is quietly leaning on, and those need opposite responses. Which funnel metrics still mean anything now covers what else shifted underneath these rates.

Test whether a round has ever changed a decision

Take the last year of closed reqs and, for each round, count the decisions that turned on it. A decision turned on a round when the verdict entering it and the verdict leaving it disagreed, and the later one won. Rounds with a count of zero are confirming. Rounds you cannot count at all are telling you something about your records, which is a finding in its own right.

Run the count, then check which of your rounds rests on an unsupervised document, because that is where the signal has moved. In a blind study run through a UK university's real examinations system, researchers submitted entirely AI-written GPT-4 answers under 33 fake student accounts across five undergraduate psychology modules: 94% of those submissions were never flagged, and the AI grades averaged just over half a classification boundary higher than the real students 2. That is coursework rather than a hiring round, the tasks were short-answer and essay questions a language model handles unusually well, and a submission counted as flagged if a marker raised a concern of any kind, so the 94% is the share that drew none at all. The submissions were unedited model output, which the authors call the most detectable possible use of the tool. Treat it as a floor on what passes unremarked.

The async video round has a number of its own now. A peer-reviewed study of machine-scored video interviews reports an uncorrected sample-weighted correlation of .24 with job performance across five organisational samples, against .32 for human-rated structured interviews in the same uncorrected terms 3. Four of the five authors were employed by the vendor whose algorithms were evaluated, the estimate rests on five samples against the 105 effect sizes behind the comparison, and both figures are uncorrected, which makes the comparison internally fair and externally thin. Take .24 as what those five samples could establish about the format. It is not evidence that machine scoring matches a structured interview, and whether the async video round is worth keeping is the case that has to be argued round by round.

What do you do when you hire six people a year?

Use the weaker version of the test, which works at any volume. Have every interviewer state, in writing, the verdict they would give on the evidence they personally collected, before the debrief opens. Then count how often a later round's independent verdict differs from the earlier one. That measures the same property at the round level without needing a year of completed hires behind it.

Two conditions make the count mean anything. The verdicts have to be independent, so nobody reads anyone else's notes first, and they have to be written in the same vocabulary across rounds, so a strong yes in round two is comparable to a strong yes in round three. Neither is expensive. Both are usually missing, and that absence is why the question feels unanswerable. The first round has a one-field version of the same discipline: record what you would have decided from the application alone, before the call, and whether the screen is producing any information becomes countable over fifty of them.

Candidate complaint is the other criterion the genre reaches for, and it needs its context attached. In a 2004 meta-analysis of applicant reactions, 53.5% of samples involved people who were not actual applicants and 90% of the hypothetical-context studies used college students; correlations between justice perceptions and behavioural intentions ran stronger in hypothetical settings than in authentic hiring, and the authors conclude that the role of fairness may be overestimated in studies conducted in hypothetical selection contexts 4. Candidate feedback still tells you things nothing else will. As the criterion that picks which round dies, it is weaker than a direct count of what each round changed. The withdrawal signal worth acting on is a specific one anyway: a candidate who is sharp in the remote screen and flat in the onsite is a coverage problem, not a length problem.

Name every round a gate or a sell

Label each round before you remove any of them. A gate round produces evidence and casts a vote. A sell round exists so the candidate meets the people and the work, and it should not vote at all. Most loops that feel too long contain two or three sell rounds quietly voting, which is both where the length came from and where the debrief disagreement comes from.

Relabelling is often the whole fix. A round that everybody values and that has never changed an outcome is usually a good sell round wearing a gate's clothes: keep it, shorten it, tell the candidate what it is for, and take it out of the scorecard set. The loop gets shorter in decision terms without losing the part people actually wanted. Calendar time moves less than that suggests: in SHRM's benchmarking data the median nonexecutive role took 44 days from requisition to accepted offer while the median organisation spent 7 days conducting interviews 5, so most of the wait sits outside the rooms you are arguing about.

Cutting a gate round is the heavier move, because removing it removes its evidence. That is safe only once you can point at the cell it occupied and say what will now go uncollected. Giving every round a cell no other round owns is the map that makes an empty column visible before you find it in a bad hire six months later.

On Monday, write your loop out as a list, mark every round gate or sell, and for each gate round write the one piece of evidence it produces that no other round does. Any gate round where that line comes out vague is your candidate. Cut it, run four reqs without it, and keep the per-round verdicts so the next version of this argument has numbers in it instead of opinions.

See how it works

Common questions

How many rounds should a loop have?

The count is the wrong unit. A three-round loop where all three are conversations about past work holds one kind of evidence; a two-round loop with a short exercise and a structured conversation holds two. Ask how many distinct kinds of evidence you end the loop with and whether each round can name its own contribution. A loop that fails that question does not get better by dropping to four rounds, and one that passes it is not too long at five.

Can we cut a round without any data at all?

Yes, but start collecting on the way. Pick the round with the weakest answer to the question of what evidence it produces that no other round does, remove it for the next four reqs, and record each interviewer's independent verdict throughout. If decisions look the same and the debriefs are no harder, the cut was right. If a specific kind of doubt keeps surfacing with nothing to resolve it, you have found what the round was quietly doing and you can rebuild that part deliberately.

Should we cut the take-home or the onsite first?

Ask which one still produces evidence you cannot get elsewhere. An unsupervised take-home built from a public exercise has lost most of its evidential value, while a live session where someone works in front of you has not. The reflex runs the other way because the onsite costs more staff hours, so the expensive round with signal gets cut to protect the cheap round without it. If the take-home is the only artifact in the loop, shrink and supervise it rather than deleting it outright.

What if the round we want to cut belongs to an executive?

Relabel rather than argue. An executive round that has never reversed a decision is a sell round with real value, because candidates want to meet the person whose priorities will shape the job. Propose keeping it, shortening it to twenty minutes, and removing it from the scorecard set so it stops casting a vote on evidence it was not designed to collect. That is an easier conversation than removal, and it keeps the part the executive actually wanted.

Does cutting rounds actually speed up hiring?

Less than teams expect, because the interviews are a small share of the elapsed time in a hire. Removing a round removes its scheduling gap too, which is the real saving, but a four-round loop with a nine-day wait in the middle is slower than a five-round loop with next-day scheduling. Cut rounds for evidential reasons and manage the gaps separately. Doing it the other way round produces a shorter loop that is not noticeably faster and knows less.

References

  1. 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report Ashby, 2026. ashbyhq.com Supports the stage passthrough rates and the reading that late stages ratify a decision made earlier rather than filtering independently.
  2. 2. A real-world test of artificial intelligence infiltration of a university examinations system: A Turing Test case study PLOS ONE, 19(6): e0305354 (Scarfe, Watcham, Clarke and Roesch), 2024. journals.plos.org Supports the claim that an unsupervised written submission no longer separates a candidate's work from a model's, measured by human graders rather than by any detector.
  3. 3. Psychometric Properties of Automated Video Interview Competency Assessments Journal of Applied Psychology, 109(6), 921-948 (Liff, Mondragon, Gardner, Hartwell and Bradshaw), 2024. hirevue.com Supports the published job-performance validity figure for machine-scored video interviews, quoted with its vendor authorship and its uncorrected basis.
  4. 4. Applicant Reactions to Selection Procedures: An Updated Model and Meta-Analysis Personnel Psychology, 57(3), 639-683 (Hausknecht, Day and Thomas), 2004. ecommons.cornell.edu Supports the caution that applicant-reaction effects are stronger in hypothetical studies than in real hiring, so a survey score is weak grounds for choosing which round to remove.
  5. 5. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the point that interviewing is a small share of the elapsed time in a hire, so removing a round buys less speed than teams expect.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.