Interviewing

How Many Interview Rounds a Role Actually Needs

Count the claims, not the interview rounds. Every stage should be the only place one claim about the candidate gets evidence, and a stage that cannot name its claim should come out of the loop whole. For most roles that collapses to three: a screen for the two or three facts that would end the conversation, one stage where the work is actually done, and one about how the person works with other people.

The takeRound count is the wrong argument, and it is the one every loop redesign starts with. Four conversations that each ask a candidate to describe past work are one conversation held four times, and trimming minutes off each does not turn them into different evidence. Loops grow because nobody trusted what the previous stage produced, so the fix sits upstream of the calendar: make one stage produce evidence solid enough that the next one has nothing left to prove.

Where Olive fits

Open a role and see what the work shows

A stage earns its place only where it produces evidence the other stages cannot. Olive runs async on the candidate's own clock in 50 to 70 minutes and comes back as six findings a reviewer wrote, so panel time goes on reading evidence rather than on holding another conversation.

Rank your shortlist

How many rounds is normal, and why that is the wrong question

Skip the benchmark hunt, because the number would not help. Elapsed time is what the benchmarks actually measure. In SHRM's 2021 benchmarking survey of its member organisations, the median time-to-fill for nonexecutive roles was 44 days, of which the median organisation spent 5 days screening applicants, 7 days interviewing, and 4 days deciding and making an offer 1. Interviewing is a small slice of a long calendar.

That distribution is worth sitting with, because it reframes the usual argument. The interquartile range around that median runs from 28 to 73 days and the mean is 54, so any single average misleads twice over 1. More to the point, the days are mostly not the loop. They are waiting on a requisition approval, waiting for four calendars to intersect, and waiting on a decision nobody scheduled. A loop with one fewer round and the same scheduling habits will not be much faster.

So the round count is a proxy for two different complaints that need different fixes. If the complaint is that candidates drop out, the cause is usually elapsed days and silence between stages rather than the number of conversations. If the complaint is that the decision still feels like a coin flip at the end, adding a round has never fixed it, because the new round tends to ask the same thing again in a different room.

The question worth asking instead is what each stage establishes that nothing else does. It is answerable in an afternoon, it produces a shorter loop as a side effect, and unlike a round count it survives contact with a role that genuinely needs four stages.

Name the claim each round is there to test

Write one sentence per stage saying what it establishes that no other stage does. A screen establishes the two or three facts that would end the conversation. A work stage establishes that the person can do the central task. A team stage establishes how they work with other people. Any stage without a sentence of its own is a duplicate, and duplicates are what make loops long.

The test is whether two stages could change the decision in different directions. If a candidate who impressed the panel in stage two was always going to impress it in stage three, stage three is only confirming what stage two already decided. The selection literature puts this in its own terms: a composite built from two or three different kinds of evidence reaches about .61 as a corrected correlation with job performance, and the gain comes from the kinds being different 2.

A worked version for an individual contributor role:

  • Screen, 20 to 30 minutes. Claim: the constraints match. Location, compensation band, and the two or three things about the role that are not negotiable.
  • Work stage. Claim: this person can do the central task, with evidence of how they did it rather than only what came out.
  • Team stage. Claim: how they handle disagreement, ambiguity, and a decision somebody else made.

Three stages, three claims. A fourth earns its place only where the role carries a specific risk with evidence of its own: managing people, handling money, or a domain where being confidently wrong is expensive. Where the new claim is about working with AI, the instinct is to bolt on a stage for it, and that is usually the wrong shape (test AI skills without adding an hour to the loop).

Which rounds to cut first

Cut the stage whose claim another stage already covers, then the stage that produces an impression rather than a record. In most loops those are the second behavioural interview and the informal culture chat. Neither is deleted for being useless. They go because the same information is being collected twice and the second collection is the one nobody wrote down.

The culture chat deserves a specific warning, since it is the stage most often defended as costing nothing. Adding a low-diagnostic conversation to good information makes the prediction worse. In a controlled test, students predicting a classmate's semester grades did worse after conducting an unstructured interview, correlating .31 with the outcome against .65 from prior grades alone, and in a further study most participants chose to interview someone answering at random over not interviewing at all 3. That is undergraduates predicting a classmate's grades in a laboratory, a long way from a manager predicting job performance. The mechanism is what travels; the effect size does not.

Two more cuts are usually safe. The panel round where four people ask overlapping questions in sequence, which can become one stage with a fixed question set split between them. And the final conversation with a senior leader that exists to confirm a decision already made, worth keeping only if that person can say no and occasionally does.

Whatever gets cut, keep the claim. If the culture chat was the only place anyone asked how the candidate handles being wrong, that claim moves into another stage with a real question and an anchored scale behind it (what actually makes an interview structured).

What a shorter loop costs, and what it does not

It costs the redundancy that was catching a small number of mistakes, and that is a real trade rather than a free saving. What it does not cost is evidence, as long as each remaining stage keeps its claim and produces a record. The loops that get worse when shortened are the ones where somebody took minutes off every stage and left all of them standing.

Three things belong in place before cutting, all of which cost less than the round being removed. A written claim per stage, so the gap a deletion leaves shows up the same day. An anchored rating scale on the stage doing the heaviest lifting, so its evidence stands up without a second opinion propping it. And a decision rule agreed in advance, because a shorter loop puts more weight on the debrief (collect the scores before the debrief).

The candidate side is not a separate project. A three-stage loop that says what each stage is for, and when the answer comes, reads as a process somebody designed. A five-stage loop with two unnamed stages reads as an organisation that has not decided what it is looking for, which is generally what it is.

Run this per role. A loop is a claim about a specific job, and one template applied to every role describes something other than the position being filled. Where a role genuinely needs an extra stage, the question is the same one: what does it establish that nothing else does (whether an AI-fluency round earns its place).

See how it works

Common questions

Do more interview rounds produce better hires?

Not on their own. What predicts better is combining different kinds of evidence, so three stages that each establish something different beat five that all ask about past work. Extra rounds do have one reliable effect: they add elapsed days, and elapsed days are where candidates go to other offers. When a fourth stage is proposed, ask what it establishes that the other three cannot, and accept the answer only if it names something specific.

How long should each stage take?

As long as its claim needs, which is usually shorter than the calendar invite. A screen establishing three facts does not need an hour. A stage where the candidate does something resembling the job needs enough time to contain a mistake and a recovery, because that is where the evidence is. The stage most often over-booked is the last one, and the stage most often under-booked is the one where work actually happens.

Is a take-home an extra round or a replacement for one?

A replacement, or it is not worth the imposition. A take-home sitting on top of the same number of live conversations has added unpaid hours for the candidate and no new claim for the loop. If it establishes that the person can do the central task, it should be replacing an interview that was trying and failing to establish the same thing. Say which stage it replaces when the brief goes out, and pay for it where the scope justifies pay.

What if the hiring manager wants one more conversation?

Ask what they would learn and write it down as a claim. Sometimes the answer is specific, in which case a stage was genuinely missing. More often the answer is that the manager is not confident and another conversation feels like a route to confidence. That is a signal the earlier stages produced evidence nobody trusts, and the fix is upstream: a better assignment or an anchored scale, not another hour on four calendars.

How do you shorten a loop without looking careless to candidates?

Say what each stage is for and when they will hear back, then do both. Candidates read a short, explained process as competence and a long, unexplained one as disorganisation. Send the questions or the assignment brief in advance wherever it does not spoil the exercise, and name the decision date. The complaint is almost never that a process was short. It is that it was silent.

Does a longer loop help with senior or executive roles?

Those loops are longer in practice, and the reason is usually more claims rather than more rounds per claim. A senior role adds judgment under ambiguity, hiring and developing people, and handling a decision that will be unpopular, each of which needs its own evidence. Add stages for the claims and the executive loop stays defensible at five, while a four-stage loop for an analyst still is not.

References

  1. 1. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the time-to-fill median, the interquartile spread, and the breakdown showing how few of those days are spent interviewing.
  2. 2. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300, doi 10.1017/iop.2023.24 (Cambridge University Press), 2023. cambridge.org Supports the composite figure and the claim that the gain comes from combining different kinds of evidence rather than from adding more stages of the same kind.
  3. 3. Belief in the unstructured interview: The persistence of an illusion Judgment and Decision Making, 8(5), 512-520 (Society for Judgment and Decision Making), 2013. sjdm.org Supports the claim that an extra unscored conversation can make a prediction worse, quoted as the laboratory analogue it is.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.