Interviewing

Every Interviewer Owns One Competency or Comes Off the Loop

An interview loop should hold whoever can bring back evidence on a competency no other seat can reach, and nobody else. Give each interviewer exactly one competency and name the evidence it owes the debrief; if two seats would return the same evidence, one goes. Size follows: the loop runs as long as the list of things you cannot learn anywhere else, usually three or four seats. Take someone off when their seat owns no competency, duplicates another, or returns an opinion where the debrief needed evidence.

The takePanels grow because adding a person is a favour and removing one is a conversation. That asymmetry is how loops arrive at seven rounds. The fix is administrative rather than diplomatic: give every seat a written competency at intake, and let the seat expire with the req. A seat somebody re-earns each search is one nobody has to be told they lost, and the candidate stops paying two extra hours for a political problem inside the company.

Where Olive fits

Open a role and see what the work shows

A seat can ask a candidate how they decide what is worth checking, and it cannot capture them deciding. Olive puts that in front of the person as work, a role-grounded assignment with an assistant that will overreach if nobody stops it, and returns six findings a human reviewer wrote from what happened in the session.

Rank your shortlist

Who should be on the interview loop?

Whoever can produce evidence on a competency the rest of the loop cannot reach. That is the whole test, and it is not the usual answer, which lists the roles that ought to be represented: hiring manager, HR, a peer, someone cross-functional, three to five people. Representation is a fairness and candidate-experience question worth answering on its own terms. Answering it leaves the peer seat's ability to judge the work entirely untested.

The usual answer also leaves the loop's actual output unexamined. In one large recruiting dataset, recruiter screens pass roughly 35% of the candidates who reach them, while later stages convert at 95% post-onsite and 81% at offer 1. Those are passthrough rates among candidates who reached each stage, they do not multiply into an end-to-end rate, and the heaviest filtering happens earlier still at application review, which is not one of these numbers. A 95% passthrough late in the loop is not evidence that an onsite is easy. It means the decision has effectively been made before the onsite ends, and the seats inside it are ratifying.

Ratifying is worth something, though four calendars and six candidate hours is a steep price for it. A seat earns its place when it changes what the debrief knows, and the only way to find out whether yours do is to read the last three searches and ask which write-up moved a verdict.

Give every seat one competency and one piece of evidence

Write two lines per seat before the invites go out: the competency it owns, and the evidence it will bring back. Evidence here means something a second reader can act on. If two seats would return the same evidence, one of them is a duplicate, and duplicates are where a loop's hours go without anyone having decided to spend them.

A worked example for a mid-level analyst role, four seats:

  • Hiring manager: judgment under an ambiguous brief. Evidence is the candidate narrating a real decision where the brief was wrong and saying what they changed.
  • Peer analyst: technical depth in this material. Evidence is a specific error the candidate caught in their own past work, and how.
  • Cross-functional partner: how the candidate handles being contradicted. Evidence is a described disagreement with an outcome, not a stated philosophy.
  • Recruiter: constraints and motivation. Evidence is dates, numbers and the reason they are looking, which is the only seat that has to be the same for every candidate.

Stacking more conversations does not substitute for different kinds of evidence. The applied follow-up to the 2022 revision of the selection-validity literature shows a composite of predictors reaching about .61, and removing cognitive ability from that composite costing .05, dropping it to .56 2. Composite validity like that assumes the pieces get combined mechanically with sensible weights and that they measure genuinely different things. Four unscored conversations mostly measure the same thing four times, which is why a fifth seat rarely buys what a different kind of evidence would. Adding a stage is the expensive way to do that; testing AI skills without adding an hour to the loop is the cheaper one.

When do you take someone off?

Three conditions, any one of which is enough: the seat owns no competency, it duplicates another seat's evidence, or it keeps returning a verdict with nothing a second reader can use. The third is the awkward one, because the person is usually senior and usually certain. Certainty is the tell, and a loop that cannot remove a seat grows until the candidate is paying for the org chart.

The nearest measured version of that third case comes from job tests rather than interviews. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed job tenures by just over 25%, and comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6% to 7% shorter job durations 3. That is not a randomised experiment, the quality measure is tenure in a high-turnover setting rather than a performance rating, and overrides were common rather than pathological in that data. It also says the average override was worse, never that the test was right about any particular person. What it will not support is the idea that a senior person overriding the structured signal is exercising superior private information as a rule.

Removal is rarely the right first move, though. Most seats that return opinions are held by people who could return evidence on something else. A manager who does not use an assistant on their own work cannot fairly assess output quality on AI-assisted material and can absolutely assess how a candidate decides what to check, which is the reassignment getting managers who don't use AI to judge AI-assisted work works through in detail.

Should the panel get bigger to be fairer?

No, and the extra seats carry a cost that is easy to overlook: every seat brings its own informal opening, and the informal opening feeds the formal rating. More observers of the same kind is not more evidence. It is the same measurement taken again, with more small talk attached and two more hours on the candidate's calendar.

There is a measurement of that leak. In a study of 189 accounting students put through structured mock interviews, the interviewer's overall impression formed during the rapport-building small talk, before any structured question, correlated .42 with that same interviewer's structured interview score. When a different interviewer supplied the structured score, the correlation fell to .25 and .24 4. Much of the larger figure is one interviewer's impression predicting their own later rating. The authors' own conclusion is not that first impressions are pure contamination, since the incremental effect ran through rated competence rather than liking or similarity, and none of it establishes that interviewers decide in the first three minutes. What it does show is that a seat is not a neutral observation post, so an uneven panel calls for structure inside each seat.

What a fifth seat is genuinely good for is building the bench. Rotate one shadow through the loop with no vote, writing independent notes before hearing anyone's read, which is the practice half of training someone to interview before they run a round alone. The bench grows and the candidate pays nothing for it, in hours or in extra deciders.

See how it works

Common questions

How many interviewers is too many?

One more than the number of competencies you cannot learn anywhere else. For most individual-contributor roles that lands at three or four, and the count is a symptom rather than a target: if you need six, either the role is genuinely six competencies wide or two seats are measuring the same thing. Test it by writing the evidence line for each seat. Seats that cannot be distinguished on paper will not be distinguished in the debrief either.

Does the recruiter get a seat?

The recruiter owns constraints, timeline, compensation range and the reason someone is looking. Those are facts a decision depends on, and that seat has to run identically for every candidate, because asking them inconsistently is how two candidates end up compared on different information. It is also the seat least suited to judging competence on the work, which is fine. A seat that owns one thing and returns it reliably is doing its job.

Can someone who does not use AI themselves judge a candidate's AI-assisted work?

Not on output quality, and yes on several things that matter more. Whether a candidate can say what they chose not to hand over, how they decided a claim needed checking, and what they did when the answer turned out wrong are all judgeable by anyone who knows the work, regardless of their own tooling. Give that seat one of those competencies explicitly. What fails is leaving them to assess a polished deliverable and calling the result a technical read.

What about a skip-level or an executive who wants to meet finalists?

Make it a conversation rather than a seat. If the meeting has no competency and returns no evidence, it should not carry a vote and should not be scheduled as an evaluation, because a seat with a vote and no rubric outranks the seats that did the work. Selling the role, answering hard questions and giving a finalist a read on the company are legitimate reasons to book the time. Say which one it is on the invitation.

How do we keep an interviewer sharp if they only sit on two loops a year?

Bring them into a calibration exercise even when they are not on a live loop. Scoring one recorded answer and comparing against the panel costs half an hour and is the only practice available to someone whose reps are rare. The alternative, which is what usually happens, is that a twice-a-year interviewer arrives with a bar set by the last search they remember, and nobody finds out until the debrief splits.

References

  1. 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report Ashby, 2026. ashbyhq.com Supports the stage passthrough figures used to show that late seats ratify rather than decide: roughly 35% at recruiter screen against 95% post-onsite and 81% at offer.
  2. 2. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300 (Cambridge University Press), 2023. cambridge.org Supports the composite figures (about .61, falling .05 to .56 without cognitive ability) behind the claim that different kinds of evidence, not more seats, is what a loop is buying.
  3. 3. Discretion in Hiring National Bureau of Economic Research, Working Paper 21709; published 2018 in the Quarterly Journal of Economics, 2017. nber.org Supports the claim that overriding a structured signal is not on average superior private information: tenure up just over 25% with the test, and 6% to 7% shorter durations per standard deviation of override rate.
  4. 4. Initial Evaluations in the Interview: Relationships with Subsequent Interviewer Evaluations and Employment Offers Journal of Applied Psychology, 95(6), 1163-1172 (Murray R. Barrick, Brian W. Swider and Greg L. Stewart), 2010. homepages.se.edu Supports the claim that a seat's informal opening feeds its formal rating: .42 within the same interviewer, .25 when a different interviewer supplied the structured score, across 189 students.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.