Interviewing
Collect the Scores Before the Debrief, Not During It
Lock every scorecard before the hiring debrief opens. Read the spread out loud first, so disagreement sets the agenda instead of getting smoothed away. Discuss only the competencies where the ratings diverged, and change a rating only against evidence somebody can quote back. Write the decision rule before you meet the candidates, because a rule set in advance is the only one you can apply the same way twice.
The takeThe debrief is the least designed stage in most hiring processes and the one with the most influence over the outcome. Three rounds of careful scored interviewing can be undone in twenty minutes by a meeting that rewards whoever is most fluent about their own impression. Treat the room as a place to surface disagreement rather than to resolve it, and let a written rule do the resolving. The person who speaks first should not be the person the rest of the panel ends up agreeing with.
Where Olive fits
Open a role and see what the work shows
A debrief can only weigh the evidence the loop produced, and a candidate describing how they would check a doubtful claim is not the same evidence as a record of them checking one. Olive returns that as work: a role-grounded assignment with an AI assistant that will do all of it if nobody stops it, written up as six findings by a human reviewer.
Rank your shortlistWhy the consensus meeting undoes the scorecards
Because a discussion aimed at agreement mixes the ratings back into a single impression, and the impression belongs to whoever argues best. Three rounds of scored interviews produce several independent readings. The standard debrief then puts them in a room and asks for consensus, and what comes out is one reading, usually the one held by the most senior or most fluent person present.
The combination step is where the evidence is either kept or lost. Combining predictors mechanically, with sensible weights, is what produces the composite figures the selection literature reports, reaching about .61 from two or three different kinds of evidence 1. That figure assumes the predictors combined are genuinely different from one another, so it is not what a team gets from stacking several unscored conversations, which mostly measure the same thing more than once.
That leaves the debrief one job it can do well and one it cannot. It can surface where the independent readings disagree and put the evidence for each side on the table, which is genuinely valuable and happens nowhere else. It cannot manufacture a better reading by averaging conviction, and a meeting that opens with what did everyone think has already started doing the second thing.
The cost is invisible from inside, which is the awkward part. A panel that talks itself into agreement feels more rigorous than one that reports a split, so the process that destroyed the most information is the one that felt best in the room.
Run the debrief in this order
Ratings in and locked before the invite opens, the spread read aloud, discussion confined to the competencies where ratings disagreed, then the rule. That order is the whole intervention. It costs nothing, needs no software beyond whatever holds the scorecards, and it is the part of structured hiring most often skipped by teams who believe they have adopted the method.
1. Locked before the meeting. Every interviewer submits ratings and the evidence behind them by a stated deadline. Nobody sees another sheet before submitting, and nobody edits after. 2. Spread first. Read out where the panel agreed and where it split, competency by competency, before any argument. The split is the agenda. 3. Only the divergences. Skip the competencies everyone rated the same. They are settled, and rehearsing them is how a meeting talks itself into a mood. 4. Evidence, not seniority. A rating changes when someone quotes something from the interview the other person had not weighed. Edits get recorded with the reason. 5. The rule, last. Apply the decision rule written before the loop opened, and note anywhere the rule did not fit. That note is next quarter's fix.
Two logistics matter more than they look. Set the submission deadline hours before the meeting rather than minutes, because a scorecard written in the corridor is a memory of a conversation the panel is about to have. And give the least senior interviewer the first word on each divergence, which costs nothing and removes the anchor everyone else would otherwise be arguing against.
Where the panel is judging something it never agreed on in advance, the debrief will split on definitions rather than on the candidate. That is worth heading off before the loop starts rather than in the room (get the panel judging AI use the same way).
What counts as a reason to change a rating?
Something quoted from the interview that the rater did not have or did not weigh. That is the entire list. Not seniority, not enthusiasm, not the fact that four other people rated it differently, and not a reconstruction of what the candidate probably meant. If the reason cannot be traced to a specific thing the candidate said or did, the rating stands and the disagreement gets recorded as a disagreement.
The rule is that narrow because panels are good at building coherent accounts out of thin material. In a controlled test of unstructured interviewing, students predicting a classmate's semester grades did worse after conducting the interview, correlating .31 with the outcome against .65 from prior grades alone, and in a further study most participants preferred interviewing someone answering questions at random to not interviewing at all 2. Undergraduates predicting grades is a laboratory analogue rather than a selection study, and the sample is small. What travels is the mechanism: people make sense of whatever is put in front of them, including nothing.
The practical form in the room is one question, asked every time somebody offers a conclusion: what did they say. It is not a challenge and it stops being awkward after two debriefs. Where the answer is a quote, the panel now has evidence it can weigh. Where the answer is a restatement of the conclusion, everyone hears that at once and nobody has to be the person who says so.
Interviewers who do not do the work themselves have the hardest time here, because the evidence they are being asked to quote is unfamiliar (judging AI-assisted work without using AI). The fix is the same as everywhere else in this process: an anchored scale written before the loop, with a worked example of what each level sounds like.
Write the decision rule before the loop opens
Decide in advance what pattern of ratings produces an offer, and write it down while nobody knows who the candidates will be. Clears the bar on all four competencies, with no dissent on the one where being wrong is expensive, is a rule. The strongest overall impression is not a rule. It is the debrief deciding what it wants and then finding a reason for it.
A workable rule names three things: which competencies are mandatory rather than tradeable, what a single low rating on a mandatory one does regardless of the rest, and who decides when the rule does not fit the case in front of them. That last one is not a loophole while it is written down and used rarely. It becomes one when it goes unnamed, because then it defaults to whoever is most senior in the room.
The standard definition of a structured interview already contains the seed of this: the same questions in the same order, a common rating scale, and interviewers agreeing in advance on what an acceptable answer looks like 3. Agreeing in advance is the habit. The decision rule is that habit applied one level up, to the pattern of ratings rather than to a single answer.
Write the rule at the same time as the scorecard, because a rule invented after the ratings exist is fitted to them. The version to avoid is a numeric average across competencies, which lets a strong showing on something easy compensate for a gap on something the job requires, and which recreates exactly the problem scorecards were there to solve (set a bar rather than an ordering). None of this holds if the interviews were not comparable in the first place, so if only some of the parts of a structured interview are in place, fix that before redesigning the meeting (what actually makes an interview structured).
Common questions
What if someone cannot submit their scorecard before the debrief?
Move the meeting or drop their rating, and say which. An interviewer who writes up after hearing the panel is not adding a fourth reading, they are adding an echo. If the schedule genuinely will not allow it, take their ratings verbally before the discussion opens and have someone else record them, then proceed. The rule that matters is that no rating is formed after exposure to another rating, and the cheapest way to keep it is a deadline set hours before the meeting rather than minutes.
Does this replace consensus with a formula?
It replaces consensus with a rule, which is a different thing. A formula would combine ratings into one number and read the answer off it. A rule names the pattern that produces an offer and leaves the judging inside each competency to people. The discussion still happens and it still changes ratings. It changes them against quoted evidence, and it happens after the independent readings exist rather than instead of them.
How should one strong dissent be handled?
Treat it as the most valuable thing in the room and ask it to explain itself in evidence. A single low rating on a mandatory competency should be enough to stop an offer where the rule says so, and the rule should say so for the competencies where being wrong is expensive. Where the dissent sits on a tradeable competency, record it, apply the rule, and keep the note in the file, so a first-quarter problem is not a surprise to anyone who was in the meeting.
Who should run the debrief?
Someone who is not making the hire, wherever that is possible. A recruiter or coordinator running to the agenda keeps the order intact, and it removes the awkwardness of a hiring manager asking people to quote evidence for ratings that disagree with their own. If the hiring manager has to run it, they speak last on every competency without exception, and somebody else is asked in advance to hold them to it.
How long should a debrief take?
Less time than it takes now, because most of the meeting currently goes on competencies everyone already agrees about. Thirty minutes is generous for a loop of four interviewers when the ratings are in beforehand and the agenda is only the divergences. If it runs long, the usual cause is a scorecard carrying competencies nobody could rate from the interview they actually conducted, which is a loop design problem rather than a meeting problem.
References
- 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors cambridge.org Supports the composite figure of about .61 and the condition the pool's hedge attaches to it, that the predictors combined are genuinely different kinds of evidence rather than several unscored conversations measuring the same thing.
- 2. Belief in the unstructured interview: The persistence of an illusion sjdm.org Supports the claim that low-diagnostic conversation dilutes better information, quoted as the laboratory analogue it is: undergraduates predicting a classmate's grades.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the claim that agreeing in advance on what an acceptable answer looks like is part of the standard definition, which the decision rule extends to the pattern of ratings.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.