Interviewing
Calibrate on Excerpts and a Disagreement Log, Not on What a 3 Means
A calibration session puts one recorded interview answer in front of every interviewer: each writes down independently which sentence carried the rating and what it evidences, then everyone reveals at once. Disagreement about the number is normal. Disagreement about which sentence mattered is the defect, and chasing it is what the session is for. Log the competency that split the room. Run one at the start of a req, again after four or five candidates, and after any change to the role or the assignment.
The takeQuarterly is a cadence chosen because quarters exist. Calibration is not maintenance on a machine that drifts at a constant rate: it is a response to two specific people reading the same answer differently, and that is discoverable from the scorecards you already have. Read the last twenty, find the competency where pairs disagree most, and spend thirty minutes on that one. A standing quarterly hour on a rubric nobody disputes teaches a panel to treat the meeting as theatre, which is worse than not holding it.
Where Olive fits
Open a role and see what the work shows
An interview can capture a candidate describing how they would check a confident claim, and a session where they have to check one captures whether they did. Olive returns six findings, each written by a human reviewer and anchored to the timestamped moment it rests on, which is the same unit a calibration session asks a panel to argue about.
Rank your shortlistWhat actually happens in a calibration session?
Four things, in this order: everyone reads or watches the same real answer, everyone writes independently, everyone reveals at once, and the competency that split the room goes into a log. The order carries more weight than the content. Reveal before writing and you have measured deference. Skip the log and next quarter's session starts from nothing, which is how a panel ends up having the same argument twice.
The reason to bother is that panels disagree exactly where it costs something. Across interviews with multiple interviewers in Ashby's benchmark, around 38% of scorecard pairs include at least one point difference between interviewers, and nearly half of those one-point differences fall between 2 and 3, crossing the yes and no threshold on a 1 to 4 scale 1. A pair is two scorecards compared against each other, so 38% is a rate over pairs, and the report gives no equivalent rate over interviews or candidates. The dataset carries no outcome measure, so nothing in it says which interviewer was right, and Ashby's customers skew to venture-backed technology employers. What it establishes is that instability concentrates at the decision boundary.
Confidence, meanwhile, does not track agreement. A review of why employers keep trusting intuition in selection reports that the interrater reliability of the traditional unstructured interview is so low that, even with a perfectly reliable and valid criterion, interview-based judgments could never account for more than 10% of the variance in job performance 2. That is a ceiling implied by how little two interviewers agree, not a measured validity, and the meta-analysis behind it dates from 1995. The same review reports a survey of 201 HR executives who rated the unstructured interview more effective than any paper-and-pencil assessment procedure. That survey was published in 1996, so it records what practitioners believed then.
So the meeting has a specific subject. Two interviewers who split 4 against 3 on a candidate they both liked have cost you nothing. Two who split 3 against 2 have landed on opposite sides of the yes and no line, and neither number carries any record of what put it there.
Make the excerpt the unit of agreement
Ask for the sentence, not the number. Each interviewer writes the specific thing the candidate said or did and what it evidences, before anyone says a rating out loud. Two write-ups that both land on 4 can rest on completely different moments, and a rating on its own hides that. The excerpt is inspectable by a third person; the number is a summary of a judgment nobody else can reopen.
Agreeing in advance belongs to the definition of a structured interview. The US Office of Personnel Management's practical guide defines one by three properties: every candidate is asked the same questions in the same order, every candidate is evaluated on a common rating scale, and the interviewers agree in advance on what an acceptable answer looks like 3. Teams ship the first two in a template and leave the third to a meeting. That guide is US federal hiring practice from 2008 and binds no private employer, and it predates every AI question by more than a decade, but the third property is the one that needs real material to settle. An acceptable answer described in the abstract is a sentence everyone nods at and nobody can apply.
Written anchors on the scale are worth having, and they solve a different problem than the one in front of you. An anchor tells an interviewer what a 3 is supposed to mean. It does not tell a second reader which sentence the 3 rests on, so a session spent entirely on scale definitions can end with everyone using the numbers the same way and still produce two write-ups nobody can reconcile. Settle the scale once in a document and spend the room's hour on evidence. Where the material is AI-assisted work and two reviewers keep landing in different places, writing a rubric two reviewers score the same way is the narrower problem underneath this one.
How often should you run one?
Three triggers, not a calendar: at the start of a req when the bar is first written, after the first four or five candidates when the pool's real shape is visible, and immediately after any change to the role, the assignment, or the tools candidates are allowed to use. A model release can move the pass rate on a work sample without anyone touching the rubric.
That third trigger is the one no incumbent agenda contains. If the take-home suddenly produces a run of strong submissions, the first hypothesis worth testing is that the exercise got easier, and the panel needs half an hour to re-agree what a strong submission now demonstrates. Pass rates jumping after a model release is the situation version of that, and it arrives with no notice. A change in the tools also moves what the panel argues about: when answers arrive fluent and tidy, two interviewers can accept every fact in a response and still split on how much of the polish to credit, which is a disagreement no scale definition settles.
The first two triggers are cheap because they sit on the calendar already. The kickoff is where the competencies get written, so it is the natural place to score one example against them while the definitions are still soft enough to change. Four or five candidates in, the disagreements are about people the panel has actually met, which is the difference between a productive argument and a philosophical one.
Keep a disagreement log between sessions
One row per split: the competency, the two readings, and what the panel decided the evidence had to show. The log is the whole difference between a ritual and a program, because it turns the next session into thirty targeted minutes on the two competencies that actually split the room. It also survives the people. An interviewer who joins in March inherits arguments the panel already settled.
The log also settles how a debrief ends. A meta-analysis of how selection evidence gets combined classifies group consensus meetings as a holistic method, alongside individual expert judgment: what makes a method holistic is that data are combined by judgment or intuition rather than by a rule applied the same way for each decision 4. A debrief that ends in a shared feeling is that method, whatever the scorecards said on the way in. The authors derive a 25% reduction in correct hiring decisions at one selection ratio, which is arithmetic on the meta-analytic validities and was measured at no employer, so treat it as the shape of a cost. The analysis never ran consensus debriefs head to head against averaged independent scores, so how much of that gap belongs to the meeting itself is unmeasured 4.
None of that argues for cancelling the debrief. A debrief is how evidence gets surfaced, challenged and corrected, and a panel that never meets loses more than it saves. What the finding indicts is the last step, where a conversation stands in for a stated rule about how the pieces combine. Training someone to interview before they run a round alone gets one person to the bar; calibration is where the rule for combining what a panel found gets written, and the log is where it stays written between searches.
Common questions
What material should we calibrate on?
Real material from your own loop, not a vendor's sample answer. A recorded round, a take-home submission with the name removed, or a transcript excerpt all work, kept to whatever the candidate was told the material would be used for. Constructed examples fail in a specific way: they are written to have a right answer, so the panel converges on it and leaves believing they agree. The value of real material is that it is genuinely ambiguous, which is the condition your interviewers actually face and the condition the session needs to reproduce.
Do we need everyone in the same room?
No, and asynchronous often works better. The requirement is independence before reveal, which a shared document enforces more reliably than a meeting does, because in a room the first person to speak anchors everyone else. Send the material, collect written responses with a deadline, then meet only to discuss the splits. That also trims the meeting to the part that needs a conversation and lets people who could not attend still contribute a reading.
How do we calibrate when nothing is recorded?
Use the write-ups. Take three completed scorecards from the last search, strip the ratings, and ask the panel to rate from the evidence the interviewer wrote down. Two things surface at once: which write-ups contain enough to rate from at all, and where the panel diverges reading the same evidence. It is weaker than a recording, because you are calibrating on one interviewer's notes and not on the candidate, and it needs no new tooling.
Should the hiring manager run the session?
The hiring manager should set the bar and should not run the room. The bar is what a strong answer has to demonstrate for this role, and it is the input the session works from. Having the manager facilitate the reveal collapses the independence the exercise depends on, because interviewers read the manager's reaction and adjust. A recruiter or a neutral facilitator collects the responses, reveals them together, and keeps the discussion on evidence rather than on who called it right.
How do we tell rater drift from a genuinely different candidate pool?
Re-score old material, and use several pieces of it. Drift means the same answer now gets a different rating, so the clean test is putting answers the panel scored six months ago back in front of them and comparing. A single re-scored example settles nothing, because one archived answer cannot separate drift from ordinary rating noise. If ratings across the archived set hold while current candidates score higher, the pool or the exercise changed rather than the raters. That is another reason to keep the log: with no archived scored material there is nothing to re-run, and every explanation stays equally available.
References
- 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the claim that interviewer ratings disagree at the decision boundary: around 38% of scorecard pairs differ by at least a point, and nearly half of those cross the 2/3 line on a 1 to 4 scale.
- 2. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the reliability ceiling on unstructured interviewing (never more than 10% of variance in job performance) and the survey of 201 HR executives who rated it more effective than any paper-and-pencil procedure.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the claim that agreeing in advance on what an acceptable answer looks like is one of the three defining properties of a structured interview, alongside common questions and a common scale.
- 4. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports the claim that a consensus meeting ending in a shared feeling counts as holistic combination, and that the 25% figure is a derivation from meta-analytic validities rather than a measured employer outcome.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.