Teams
Interviewer Training Ends in a Rated Practice Round, Not a Deck
A new interviewer is ready to run a round alone when someone who did not watch the interview can act on their write-up. Get there with three steps and one bar: they read two completed write-ups from the role, score a recorded round independently, then sit with the person who ran it and compare where the evidence diverged. The bar is a quote or a described action behind every rating. Attendance certifies nobody.
The takeMost interviewer training is procurement with a slide deck attached. It gets bought after a loop produces two opposite verdicts and somebody has to be seen doing something, and it ends without anyone finding out whether those two verdicts would happen again. A program with no failure path is a scheduling exercise. If you are not willing to keep someone off a live loop this quarter, skip the program and run the practice round anyway, then read what comes back.
Where Olive fits
Open a role and see what the work shows
The same six dimensions Olive reports on describe what a good write-up captures: how a person framed the problem, what they demanded a source for, what they refused to hand over, and what they checked against something outside the conversation. Every finding is written by a human reviewer and carries the timestamped excerpt it rests on.
Rank your shortlistWhat does interviewer training have to produce?
An artifact, not an attendance record: a write-up a second reader can act on without having been in the room. Everything else in a training program is means. If the trainee finishes and the only evidence they learned anything is that they were present for four hours, the program has measured a calendar, and the first live loop is where you find out what it did not teach.
There is evidence that what an interviewer writes down carries signal a rating alone does not. Text-mining post-interview notes on 7,650 candidates hired at a large Chinese internet technology company, researchers found that the number of job-related capabilities an interviewer named in their notes was positively related to later job performance and promotions, and negatively related to turnover 1. The effect is small, roughly 2 percent of a performance measure per standard deviation, and it is correlational, at one firm, among people who passed the interview, so it says nothing about anyone who was rejected. What the study measured is whether the notes name the capabilities the job analysis called for, which is a different skill from settling on a rating. Only the first is teachable in an afternoon.
That also reframes who should run the training. Programs get built around the team's best interviewer, and the category is weaker than it sounds: a review of why employers keep trusting intuition in selection reports that research on variance in interviewer validity suggests the differences between interviewers are due entirely to sampling error 2. Highhouse is summarizing a study he cites there, so read it as evidence that fails to support stable interviewer-level differences. It does not rule out a gifted interviewer. The same review reproduces a 1943 admissions result in which high school rank plus an aptitude test correlated .45 with academic achievement, while the same two predictors plus counselors' intuitive judgment correlated .35. A practiced human judgment laid on top of structured evidence made the prediction worse.
Run the practice round before the live one
Shadow first, reverse-shadow second, and make the shadow write independent notes before hearing the interviewer's verdict. A shadow who hears the verdict first is measuring agreement, which they will supply for free. The cheap way to size the problem across a team is one recorded round, every interviewer writing it up cold, and the write-ups read side by side. The spread is the size of the problem.
1. Read two completed write-ups from the role the trainee will interview for: one strong, one a debrief had to send back. Reading a good artifact is faster than describing one. 2. Score a recorded or role-played round cold. No verdict, no scorecard, no hallway summary beforehand. 3. Compare against the person who ran it, and spend the hour on where the evidence diverged rather than where the ratings did.
The independence in step two is the step teams skip and the one carrying the exercise. An interviewer who has heard a verdict will find support for it, and the finding underneath that is how readily people build a coherent read out of very little. In one test, 96 of 169 participants chose to conduct an interview in which the other person answered questions at random over conducting no interview at all. In a separate test in the same paper, 76 students' predictions made after an unstructured interview correlated .31 with a classmate's actual grades, against .65 for prior grades alone 3. These are small lab studies, a long way from a manager predicting job performance. The mechanism is what travels: thin conversational information dilutes the good information already on the table.
A shadow seat is the cheapest version of all three steps and it works on that one condition. It is also where someone who does not personally use an AI assistant gets current on what the loop is now reading, which is its own problem: getting managers who don't use AI to judge AI-assisted work covers the seat they should hold in the meantime.
What is the bar, and what happens to someone who misses it?
The bar is that a second reader can act on the write-up without having watched the interview, and someone who misses it does not get a slot on a live loop yet. In practice that means a quote or a described action behind every rating, and you grade against exactly that. The commercial curricula that rank for interviewer training specify the delivery half and none of them says what a trainer does with a trainee who falls short.
Three responses, in ascending cost. Another practice round on new material, which handles most cases, because the usual miss is a rating with nothing underneath it and one round of being asked *what did she say that made you write that* fixes it. A seat on a different competency, for someone who judges reasoning well and product quality poorly. Or a shadow seat with no vote for a quarter, which costs the candidate nothing and keeps the person in the rotation.
The second most common miss is subtler and worth naming to every trainee: a rating whose evidence is the candidate's manner. Fluency, polish and pace are now partly a product of preparation tools rather than of the person, so a trainee who has learned to read confidence has learned to read the wrong variable. Stopping a hiring manager rejecting someone for sounding like AI is the situation this creates downstream, and it is cheaper to prevent in training than to argue about in a debrief.
When does training expire?
Retrain when the work the role does changes, not on an annual cycle. A trainer who last interviewed for this job two years ago is teaching a job that no longer exists, and the questions they hand over were built against tasks that have since moved. Budget one practice round per competency the person will own, and re-run that round when the assignment changes.
What the training is actually buying is structure, and that is the part with numbers behind it. A 2022 re-analysis of the selection literature estimates structured interviews at .42 and unstructured interviews at .19, from raw observed validities of .32 and .13 before any correction 4. Structured there is a research coding of interview format, meaning the same questions rated on a common scale, and no script or product sits behind the number. None of the underlying studies involve interviews about how a candidate works with AI. That gap sits between two interview formats, so it sizes what structure is worth. Training is what gets the structured version actually run.
One round of training gets a single person to the bar. Holding a group at the same bar is a different meeting with a different unit of agreement, and what actually happens in an interviewer calibration session is where it belongs. Schedule the first one for the week the trainee's first live loop runs, while their write-up is still fresh enough to argue about.
Common questions
How long should interviewer training take?
Less classroom time than most programs use and more practice time. Two hours to read completed write-ups and the role's competency definitions, one round scored cold, and one hour comparing that round against the interviewer who ran it will move someone further than a full day of curriculum. The variable that matters is not total hours but how many rounds the trainee scores independently before their first live one. One is the floor. One per competency they will own is the target.
Is shadowing on its own enough?
Only if the shadow writes independent notes before hearing the interviewer's verdict. A shadow who is told the read first will produce agreement rather than skill, and both parties will leave believing the session worked. Sequence it properly and shadowing becomes the cheapest form of the whole program: same room, same candidate, two write-ups, one comparison. Sequence it wrong and it costs the candidate an observer and teaches nothing.
Who should run the training?
Someone currently interviewing for that role, paired with whoever owns the competency definitions. The instinct to hand it to the team's most admired interviewer is worth resisting: the evidence does not support stable differences in interviewer validity, and a program built on one person's instincts transfers instincts rather than method. What you want from the trainer is completed write-ups worth imitating and the willingness to say where a trainee's evidence was missing.
The trainee's write-ups are good but their ratings run generous. Is that a training problem?
That is a calibration problem, and it is the easier of the two to fix. A generous rater whose evidence is specific has given the debrief everything it needs, because a second reader can look at the excerpt and disagree with the number. The reverse case, a well-calibrated rating with nothing underneath it, is the expensive one. Train for the evidence, calibrate the scale afterwards with the rest of the panel.
Does someone need retraining to interview for a different role?
Not the whole program, but yes to one practice round on the new role's material. What transfers is the method: independent notes, evidence behind every rating, a write-up someone else can act on. What does not transfer is knowing what a strong answer must contain in a job you have never done, and that is exactly the part that decides whether a follow-up question lands. Budget one round against the new competency definitions before the first live loop.
References
- 1. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company frontiersin.org Supports the claim that notes naming job-related capabilities carry predictive signal, with the 7,650-candidate sample and the roughly 2 percent per standard deviation effect stated.
- 2. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the claim that the best interviewer is not an evidenced category, and the Sarbin 1943 result in which adding intuitive judgment to two mechanical predictors lowered prediction from .45 to .35.
- 3. Belief in the unstructured interview: The persistence of an illusion sjdm.org Supports the claim that interviewers construct a read from almost nothing: 96 of 169 chose an interview with random answers over none, and post-interview predictions correlated .31 against .65 for prior record alone.
- 4. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the structured-versus-unstructured interview FORMAT gap the article uses to size what structure is worth: .42 against .19 corrected, .32 against .13 observed. Not a claim about trained versus untrained interviewers.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.