Interviewing

One Cell Per Round: The Coverage Map That Ends Repeat Rounds

Two interview rounds differ when they collect different kinds of evidence, not when they carry different competency labels. Build a grid with competencies down the side and evidence type across the top: produced under observation, produced unobserved, attested by a third party, judged rather than made. Give every round one cell no other round occupies. Two rounds sitting in the same cell are one round run twice, however different their question lists look.

The takeOne or two competencies per interviewer is the standard advice, and it fixes the wrong collision. Labels stop overlapping; evidence keeps overlapping. Four rounds that are all conversations about past work measure one thing four times, and what separates the candidates is how well each has rehearsed. Preparation used to be its own limiter, back when a polished account of past work had to be built by hand. Building one is cheap now. The second axis does all the work the first axis was credited with.

Where Olive fits

Open a role and see what the work shows

A conversation can hold a candidate's account of how they check a confident claim; a work session can hold them checking one. Olive runs the second kind: a 50-to-70-minute session on a role-grounded assignment in one occupation, written up by a human reviewer as six findings, each carrying the moment it rests on.

Rank your shortlist

What makes two rounds actually different?

The kind of evidence they collect. Competency labels are cheap to differentiate and change nothing: two interviewers can be assigned communication and problem solving, run the same format, and come back with the same impression of the same rehearsed narrative. What separates rounds is whether the candidate produced something in front of you, produced it out of sight, was described by a third party, or was asked to judge something instead of making it.

The combination literature arrives at the same place from the other direction. Sackett and colleagues report that a composite of predictors reaches a corrected validity of about .61, and that removing cognitive ability from it entirely costs .05, dropping it to .56 1. The gain there comes from combining different kinds of evidence, and it assumes the predictors are combined mechanically with sensible weights. It is not what a team gets from stacking four unscored conversations, which mostly measure the same thing twice. Structured interviews sit at the top of the same single-method list at .42, with an 80% credibility interval running from .18 to .66, wide enough that the point estimate should never travel without it.

So the design question is not how many rounds you run. It is how many distinct kinds of evidence you end the loop holding. Count that for your own loop before you defend it; the answer is often two.

Build the grid: competency down, evidence type across

Draw four columns and put your competencies down the left. The columns are the evidence types: produced under observation, produced unobserved, attested by someone else, and judged rather than made. Then write each round into exactly one cell. A round that cannot be placed in a single cell has no defined job, which is usually why its written feedback reads like everyone else's.

What each column actually holds:

  • Produced under observation. A live exercise, a pairing session, a whiteboard problem, a document drafted in the room. You watch the choices, including the discarded ones.
  • Produced unobserved. A take-home, a portfolio piece, a writing sample. Cheap to collect, and the weakest column it has ever been, because the finished artifact no longer shows who or what produced it.
  • Attested by someone else. References the candidate names, a verified employment record, a confirmed credential. Different failure modes entirely, which is exactly why it is a separate column.
  • Judged rather than made. The candidate is handed work of uneven quality and asked what is wrong with it. In a study where engineering students used ChatGPT during open-book take-home exams and submitted their interaction transcripts, some of the strongest evidence of reasoning appeared where students evaluated incorrect or incomplete responses, and the author concluded that correctness of the final answer alone may no longer provide sufficient evidence of comprehension 4. That is one qualitative study with no control group and no grading comparison, so treat it as a design direction, not a validated method.

The fourth column is the one most loops have never used, and it is the cheapest to add: a different prompt inside a stage you already run covers it. Running a structured interview about AI use so candidates stay comparable is where that prompt gets written properly.

How do you tell a second read from a repeat?

Declare it in advance and staff it blind. A deliberate second independent read is a legitimate use of an occupied cell, but only when the second reader has not seen the first one's notes and the round is labelled a replication before anyone runs it. Declared after the fact, agreement is not a measurement. It is an echo you paid for with a candidate's afternoon.

Independence is also the thing a debrief quietly spends. Kuncel and colleagues classify group consensus meetings as a holistic combination method, alongside individual expert judgment, because what makes a method holistic is that the data get combined by judgment, insight or intuition rather than by a rule applied the same way for each decision 2. Their own derivation puts the cost of holistic combination at as much as a 25% reduction in correct hiring decisions at a selection ratio of .30. That 25% is a Taylor-Russell calculation from meta-analytic validities. No employer measured it, and it moves with the selection ratio and base rate you assume. The meta-analysis also never ran consensus meetings against averaged independent scores head to head, so it cannot say how much of the loss belongs to the meeting. None of this argues against holding a debrief, which is where evidence gets surfaced and challenged. Keep the meeting; take the combining out of it.

The stakes show up at the decision boundary. In Ashby's benchmark of interviews run with more than one interviewer, around 38% of scorecard pairs include at least one point difference between interviewers, and nearly half of those one-point differences fall between 2 and 3, crossing the yes-or-no threshold on Ashby's four-point scale 3. The dataset carries no outcome measure, so nothing in it says which interviewer was right, and Ashby's customers skew to venture-backed technology employers. What it does say is that ratings are unstable exactly where the decision gets made, which is the argument for making the evidence behind a rating inspectable, so a debrief at the boundary has something to examine besides the number.

Run the retrospective before you redesign the loop

Pull the last ten loops and count, per round, how often its verdict differed from the round before it. That count tells you which cells are genuinely occupied and which rounds are echoing. It is far cheaper than a redesign, and it produces the only evidence a hiring manager will accept when you propose taking a round out of their loop.

Three things make the count usable. Each interviewer states a verdict on the evidence they personally collected before the debrief opens, so the numbers are independent. Verdicts are recorded in the same vocabulary across rounds, so they are comparable. And the count is kept per round, because a loop-level agreement rate hides which round is doing the confirming.

Then hand the panel the grid at kickoff. One line per interviewer naming the cell they own, plus the competency, plus what a written finding from that cell should contain. A question list tells an interviewer what to say; a cell tells them what they are responsible for bringing back, and the second is what stops two people arriving at the debrief with the same sentence. Follow-up questions that expose whether someone understands the answer they just gave is the technique that fills the observation column once the assignment is clear.

On Monday, write the grid for one open req, in a spreadsheet, in ten minutes. Name the cell each round owns in a single line. Any round left without a cell is your candidate for the round to cut, and any empty column is the evidence your loop currently never collects.

See how it works

Common questions

How many evidence types does a loop actually need?

Two or three, held well, beats four held loosely. Most loops already collect work the candidate produced out of sight, plus a great deal of conversation, so the highest-return additions are usually observation and judgment: something made in front of you, and something judged instead of made. References occupy a genuinely separate column but come late and rarely change a decision. The test is whether you can name what each column contributed when the debrief disagrees.

Is it a problem if two interviewers agree about a candidate?

Only if agreement is all the round ever produces. Two interviewers agreeing after collecting different kinds of evidence is a genuine convergence and worth having. Two agreeing after hearing the same narrative in two conversations is one observation counted twice, and it makes a loop feel more certain than the evidence supports. The distinguishing question is what each of them saw that the other did not, and the answer should be concrete enough to write down.

Where does an AI-related round fit on the grid?

It is not a separate row. Working with AI is visible in every competency, so it belongs in the columns, not in a new competency line: work produced under observation with an assistant available, and work judged instead of made. Adding an AI round on top of an existing loop usually produces a fifth conversation about tools. Change what one existing round asks the candidate to do, and the evidence arrives without an extra hour on anyone's calendar.

What if a hiring manager insists on keeping a round that owns no cell?

Keep it and relabel it. A round can be worth running for reasons that are not evidential: meeting the team, answering the candidate's questions, selling the role. Those are real and they matter to acceptance rates. What the round should lose is its vote. Name it a sell round in the kickoff note, exclude it from the scorecard set, and the loop keeps the benefit without letting an unevidenced impression into the decision.

Do we need a new template to run this?

A spreadsheet does it. Competencies down the left, four columns across the top, one round per cell, and one line per round saying what a finding from that cell should contain. The reason to keep it small is that it has to be handed to the panel at kickoff and read in under a minute. Anything longer becomes a document nobody opens, and an unopened template is indistinguishable from having no map.

References

  1. 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300 (Sackett, Zhang, Berry and Lievens), 2023. cambridge.org Supports the composite figures and the structured-interview point estimate with its credibility interval, and the condition that predictors be combined mechanically.
  2. 2. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis Journal of Applied Psychology, 98(6), 1060-1072 (Kuncel, Klieger, Connelly and Ones), 2013. gwern.net Supports the classification of a group consensus meeting as holistic combination and the derived cost of combining evidence by discussion rather than by a stated rule.
  3. 3. Recruiting Operations Benchmarks | 2026 Talent Trends Report Ashby, 2026. ashbyhq.com Supports the scorecard-pair disagreement rate and the observation that many of those differences cross the decision threshold.
  4. 4. Reimagining Assessment in the Age of Generative AI: Lessons from Open-Book Exams with ChatGPT arXiv:2605.12363 (Mahmoud), 2026. arxiv.org Supports the judged-rather-than-made column: reasoning showed most clearly where students evaluated incorrect responses, and final-answer correctness alone was insufficient evidence.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.