Teams
Lock the Scores Before the Debrief, and Keep the Split Visible
Interviewers should write their scores down before anyone talks. Each files a rating and the evidence behind it before seeing anyone else's, because a score written after hearing a colleague's is a vote rather than an observation, and a panel that converges through discussion sounds more confident than it was. Then check your scorecard tool: if submitted cards are visible to the panel before everyone has filed, the instruction is already defeated. Treat a genuine split as the result the exercise produced.
The takeMost debrief advice is written as etiquette: let the junior person speak first, keep the loudest voice in check, appoint somebody to argue the other side. Manners are not the problem here. Once a rater has seen a number, their own stops being separate evidence, and no seating order restores it. The habit that costs the most is discussing until the ratings match, which spends the one thing independent scoring bought.
Where Olive fits
Open a role and see what the work shows
An interview can capture a candidate describing how they would check a confident claim; a working session can capture whether they checked one. Olive hands back six findings on one candidate, each written by a human reviewer and printed beside the timestamped moment it rests on, so a panel that disagrees has something specific to disagree about.
Rank your shortlistWhy does a score written after the debrief stop counting?
Because it is no longer an independent observation of the round. A rater who has already seen a colleague's number carries it inside their own, and averaging the two counts one judgment twice. A panel that converges after discussion sounds like four people who agree. What it often is, is one person who spoke first and three who adjusted.
A 2013 meta-analysis of how selection evidence gets combined puts this precisely. It classifies group consensus meetings as a holistic combination method, sitting alongside individual expert judgment, and defines holistic as any method where data are combined using judgment, insight or intuition rather than a rule applied the same way for each decision 1. So a debrief that ends in a shared feeling about a candidate is, in that taxonomy, the same method as one person forming an impression, whatever the headcount in the room.
Two things follow, and only one of them is the obvious one. The obvious one is to file first. The less obvious one is that this is not an argument against holding the meeting: a debrief is how evidence gets surfaced, challenged and corrected, and none of that happens on a form. What the classification indicts is the final step, when the room converts several accounts into a decision by talking until nobody objects.
The authors do put a size on it, and the size needs its label attached: they calculate that the lower validity of holistic combination can mean a 25 percent reduction in correct hiring decisions at a selection ratio of .30 1. No employer measured that figure. It is a Taylor-Russell derivation from meta-analytic validities, and it moves with the selection ratio and base rate you assume. It also prices holistic combination as a whole: consensus meetings sit in that category by classification, and the paper never ran one against a set of averaged independent scores, so the 25 percent does not price the debrief on its own. Treat it as a direction and an order of magnitude.
Check whether your tool already breaks the rule
Open your applicant tracking system, pick a live requisition, and try to read a colleague's submitted scorecard before filing your own. If you can, every line about independent scoring in the interviewer guide is advisory, and the default wins, because the tool is what people actually touch on the day the interview ends. The check takes about four minutes, and it settles whether the guide or the default is in force.
If there is a setting for it, changing it is the whole fix. If there is not, the workaround is unglamorous and works: the recruiter collects ratings by form, holds them, and pastes all of them into the debrief document at once, before anyone speaks. The person who collects should not be the hiring manager, for the same reason the hiring manager does not go first.
A second leak sits upstream of the scorecard and is worth closing at the same time. Interviewers who were briefed differently are not scoring the same thing to begin with, so independence buys less than it should. Getting a panel to judge AI use the same way is the prerequisite: a shared brief, then independent ratings, then the meeting.
This takes no policy document. It needs one setting changed, one person named to collect, and an agenda that opens with the ratings instead of the hiring manager's read of the candidate.
What should you do with a genuine split?
Keep it, and make it the agenda. A split is the one thing independent scoring produces that nothing else in the loop can, and the standard advice to discuss until consensus spends it in the first ten minutes. Name the dimension people diverged on, put the two readings side by side, and go back to what each rater actually saw. The meeting is for locating the disagreement precisely, not for closing it before anyone understands where it came from.
There are only two honest reasons to change a rating in the room. The first is new evidence: another interviewer saw the thing you were looking for and you did not, and now you have it. The second is a misread rubric, where you were rating against a standard the team had already agreed meant something else. Changed my mind after hearing Dana is neither of those, and a debrief that produces three of them has measured seniority.
The most useful failure mode is the one where a rating has no excerpt behind it. When a rater cannot produce the sentence or the moment their number rests on, that is itself the finding, and it belongs in the notes: this round produced a judgment nobody can inspect. The interviewer is rarely the problem here. Usually the question did not surface anything gradeable, which is a design problem in the loop.
Where two finalists genuinely land in the same place, more discussion will not separate them, and the tie needs a different kind of evidence. That case is worked through in telling who did the thinking when two finalists produced the same quality of work.
Should you add more interviewers instead?
Not as a bias control. The intuition is that several impressions cancel each other out, and the review that gathers the structured-interview evidence reports nothing that supports it: across five studies, race and gender similarity effects between interviewer and candidate were small or absent in one-on-one structured interviews and in two, three and four-person structured panels alike 3.
Those were large field studies with carefully built structured interviews, which is where a similarity effect would be hardest to find, so read it as a null in favorable conditions and not as proof of zero. Structure appears to be doing the work, and panel size neither added nor removed an effect on top of it. What a second rater does buy is a view not already shaped by your own first impression of the candidate. In a study of 189 accounting students in structured mock interviews, the interviewer's overall impression formed during the small talk before any structured question correlated .42 with that same interviewer's later structured score, and about .25 when a different interviewer supplied the structured score 2. It is a student sample in mock interviews, and a correlation: it is not evidence that anyone makes up their mind in the first three minutes.
The practical reading is narrow. On this evidence, adding interviewers does not remove a similarity effect, and it costs real calendar time. Keeping the people you already have from reading each other first is free, and it closes a leak you can watch happen. If the loop has one interviewer and no second opinion at all, that is a different problem with a different fix, and it is worth solving before the debrief mechanics.
So the order to spend on is: brief everyone the same way, have each of them file independently, then meet. Headcount is the last lever.
Common questions
How long should interviewers get before the rating is due?
Long enough to write, short enough that nothing else has been read first. The sequence matters more than the clock: a card filed six hours later from unaided memory carries more than one filed in fifteen minutes after skimming a colleague's notes. Set the rule as an order rather than a deadline, and say plainly what counts as contamination: another card, a group chat message about the candidate, or a summary produced from the recording.
Can an interviewer change a rating after the debrief?
Yes, and the change should be visible rather than quiet. Keep the original rating, add the revised one, and record the reason in one line naming what new evidence produced it. That leaves a trail somebody can inspect later, and it makes the difference between a panel that updates on evidence and one that converges on the loudest reading. A loop where most ratings move after the meeting is telling you the ratings were never independent.
What if the panel is only two people?
Independence matters more, not less. With two raters there is no majority to hide behind, so a split has to be resolved on the evidence or escalated, and a rating that was quietly copied from the other person removes the only cross-check the loop has. File separately, read both aloud, and if the two disagree and neither can produce the moment behind the rating, add a round.
Does this apply to the recruiter screen?
It applies wherever two people rate the same thing. A recruiter screen usually has one rater, so there is nothing to keep independent, but the note it produces becomes the first account every later interviewer reads. Treat that as an input to be labelled rather than as a verdict: write what was asked and what the candidate said, keep any recommendation short, and avoid framing the candidate before anyone has met them.
Isn't a split just a sign the process is broken?
A split on the same dimension every time is a rubric problem. A split on different dimensions across candidates is normal and useful, because interviewers see different rounds and different behaviour. What should worry you is the opposite pattern: a panel that agrees on everything, every time, which usually means either the ratings are not independent or the questions are not discriminating between candidates.
References
- 1. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports the classification of group consensus meetings as holistic combination, the definition of holistic as combination by judgment rather than a rule applied the same way each time, and the 25 percent reduction figure at a selection ratio of .30 as a Taylor-Russell derivation.
- 2. Initial Evaluations in the Interview: Relationships with Subsequent Interviewer Evaluations and Employment Offers homepages.se.edu Supports the .42 within-interviewer and .25 across-interviewer correlations between a pre-question impression and the structured interview score, in a sample of 189 students in mock interviews.
- 3. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature doi.org Supports the claim that across five studies, race and gender similarity effects between interviewer and candidate were small or absent in one-on-one structured interviews and in two, three and four-person structured panels alike, with no panel-against-individual validity comparison reported in the review.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.