Teams

Calibration Is a Standing Session, Not an Onboarding Slide

Calibration means two interviewers giving the same candidate the same rating, and it is one of the few hiring practices you can measure directly. Hand four raters one past submission, collect independent ratings, and look at the spread. Then repeat it, because the useful number is not today's spread but whether it is widening. How fast it wears off has no established interval, so re-measure when the pool changes shape, when someone new starts reading the rubric, or when the tools candidates may use change.

The takeThe drift worth hunting this year has a name nobody says out loud: a private markdown applied to work that looks AI-assisted. It never enters the rubric, so it never gets argued about, and two raters can be applying discounts of very different sizes while both write down a 3. It leaves a distinctive trace, which is disagreement concentrated on one dimension while the rubric text has not moved. Write the rule down and the discount becomes a decision the team can defend or drop.

Where Olive fits

Open a role and see what the work shows

Olive's report gives a reviewer's finding and the timestamped excerpt behind it on each of six dimensions, so two people comparing notes are comparing moments in a session rather than two numbers. The candidate is given the same report on every tier.

Rank your shortlist

How do you calibrate interviewers?

Hand several raters the same real submission, have them rate it without conferring, and record how far apart they land. The spread is the measurement, and it is the only one this practice produces on its own. Everything else in the session is discussion, which is worth having, but discussion with no recorded spread leaves nobody able to say whether the last session changed anything.

Use material you already own: a past take-home with the name removed, a recorded answer, a transcript from a candidate who has since been hired or declined. Hypotheticals generate agreement that evaporates the first time a messy real submission arrives, because the disagreements live in exactly the details a clean example does not have.

Four raters is enough to see a shape and few enough to schedule. Collect the ratings before anyone speaks, reveal all of them at once, and open on the widest pair. The mechanics of the meeting itself, who runs it and what goes on the agenda, are worked through in what actually happens in a calibration session; this article is about the number it should leave behind.

One more rule that people skip: calibrate on the rubric you are actually using, not a cleaned-up version written for the session. Calibrating against an idealized rubric measures agreement with a document nobody applies on a Thursday afternoon.

What can you actually measure?

Three numbers, none of which needs a statistician. The range between the highest and lowest rating on each dimension. The share of rater pairs landing within one point of each other. And the share of raters who could name the specific sentence behind their own rating when asked. Record all three every time, because one session's spread describes that submission, and only a series describes your panel.

The reason to bother is that unaided interview judgment has a low ceiling to begin with. Highhouse's review of why employers resist selection aids states that the interrater reliability of the traditional unstructured interview is so low that even with a perfectly reliable and valid criterion, interview-based judgments could never account for more than 10 percent of the variance in job performance 1. That 10 percent is an upper bound implied by how little two interviewers agree, so it says what such judgments could at most explain, and it was derived for unstructured interviewing.

The same review reports the gap that makes this hard to fund. In a survey of 201 human resources executives, the unstructured interview was rated more effective than any paper-and-pencil assessment procedure 1. That survey was published in 1996, so treat it as a snapshot of belief rather than a current reading, but the shape of it will be familiar to anyone who has proposed adding structure to a loop.

What you are buying with the three numbers is an argument that does not rely on anyone's confidence. A panel that can see its own spread widen has a reason to meet that survives a budget conversation.

Name the unstated discount on assisted work

Write down what assisted work is rated on before a rater has to decide alone. The drift showing up now is a private markdown on submissions that look AI-assisted, at sizes nobody agreed and most raters will not state out loud. Split the spread by dimension: when one carries most of the disagreement and the standard has not been reworded all year, an unwritten rule is doing the work.

The discount rests on a judgment people cannot reliably make. In a 2021 study, evaluators asked to tell GPT-3 text from human writing across stories, news articles and recipes performed at random chance, and three quick training methods lifted accuracy only to about 55 percent 2. That was an earlier model generation and crowdworkers rather than domain experts reading work in their own field, and the model improvements since are a reason to expect unaided judgment to have got harder.

The errors are also not evenly distributed, which is what turns a bad habit into a fairness problem. Seven widely used detectors run over 91 human-written TOEFL essays by non-native English speakers produced an average false-positive rate of 61.3 percent, while the same tools were near-perfect on essays by US eighth-graders 3. Those are academic essays and 2023 tool versions, so the rate is not a constant. The direction of the error is the durable part, and it lands on people writing in a second language.

So the question a calibration session can actually settle is what the team rates once assistance is assumed. Two adjacent situations are worth reading before you write that rule: stopping managers from rejecting candidates for sounding like AI and getting managers who do not use AI themselves to judge AI-assisted work.

How fast does calibration wear off?

Calibration has no known half-life, so treat any interval you are handed as a convention rather than a finding. What is observable is the trigger, and there are three: the applicant pool changes shape, the rubric gets reinterpreted by somebody new, or the tools candidates are allowed to use change underneath the same assignment. Re-measure after each of those, and put a session on the calendar only as a floor.

The third trigger is the one that catches teams out, because nothing in the process announces it. A model release can move pass rates on an unchanged work sample, and a team with no calibration series will read that as a stronger cohort arriving. That case has its own diagnosis in pass rates jumping after a model release.

Do not expect a strong rater to hold the line between sessions. Highhouse's review reports that although it is commonly accepted that some employment interviewers are better than others, research on variance in interviewer validity suggests the differences are due entirely to sampling error 1. That summary rests on one 1996 study, so read it as an absence of evidence for stable differences between interviewers, and note that it does not say interviewers add nothing. What it undercuts is the habit of treating one person's judgment as the standard, and a series of measured spreads is the thing that replaces it.

That is the practical case for making this recurring work with a named owner. A calibration session that happens when somebody remembers is a session that stops happening in the second quarter, and the spread it was holding down goes back to whatever it was before anyone measured it.

See the benchmarks

Common questions

How many raters does a calibration session need?

Four raters is a workable default for a calibration session, and two is enough to start. What matters is that they rate the same material independently before anyone speaks, so the number is bounded by scheduling rather than by statistics. With four you can see whether disagreement is one outlier or a genuine split down the middle, which changes what you do next. With two you learn only the size of one gap, which is still more than most panels have ever measured.

What material should we calibrate on?

Real submissions from your own pipeline, with identifying details removed. Past take-homes, recorded answers and transcripts all work. Avoid vendor sample answers and invented examples: they are written to be clearly good or clearly bad, so raters agree, and the session produces a comfortable number that predicts nothing about the ambiguous submission arriving next week. Keep a small library of past cases, including two that split the panel badly, and reuse them across sessions so the spread is comparable over time.

How do we tell rater drift from a genuinely different candidate pool?

Re-rate an old submission. Drift means the same work draws a different rating than it did six months ago, which is only visible if you kept the work and the original ratings. A pool that has genuinely changed will show up as a change in the distribution of new ratings while the archived case still gets the number it always got. This is the reason to keep a fixed set of calibration cases rather than pulling a fresh one each time.

Does calibration make a hiring process fair?

Calibration makes a process consistent, which is a different claim and a smaller one. Two raters can agree perfectly and both be applying a criterion that has nothing to do with the job, and a spread of zero would not reveal it. Agreement is worth measuring because inconsistency hides everything else, but it is a diagnostic rather than a fairness result. Whether the criteria themselves are job-related is a separate question, answered by looking at the rubric and at stage-level outcomes.

Who should own calibration?

Calibration needs one named owner, usually in people operations or recruiting operations, who holds the case library, schedules the session and records the spread. The hiring manager should not own it, because the practice exists partly to check the hiring manager's own reading against everyone else's. Ownership matters more than seniority here: the failure mode is not a bad session, it is no session, and a recurring task with nobody's name against it is the one that quietly stops.

References

  1. 1. Stubborn Reliance on Intuition and Subjectivity in Employee Selection Industrial and Organizational Psychology, 1(3), 333-342 (Scott Highhouse), reporting Conway, Jako and Goodman (1995), Terpstra (1996) and Pulakos, Schmitt, Whitney and Smith (1996), 2008. edbatista.com Supports the 10 percent variance ceiling implied by the unstructured interview's interrater reliability, the survey of 201 HR executives who rated the unstructured interview more effective than any paper-and-pencil procedure, and the statement that variance in interviewer validity is attributed entirely to sampling error.
  2. 2. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, 2021. aclanthology.org Supports the claim that untrained evaluators distinguished generated from human text at chance, and that brief training lifted accuracy only to about 55 percent.
  3. 3. GPT detectors are biased against non-native English writers Patterns (Cell Press), 2023. pmc.ncbi.nlm.nih.gov Supports the 61.3 percent average false-positive rate across seven detectors on TOEFL essays by non-native English writers, against near-perfect accuracy on US eighth-graders' essays.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.