Interviewing

One Interviewer, Two Passes: Split Evidence From Verdict

When you are the only interviewer, the way to keep your first impression from deciding the hire is to split evidence from verdict: during the interview write only what the candidate said and did, with no rating and no adjectives, then score that record the next morning against a bar you wrote before you met anyone. Ask every candidate the same questions in the same order, and score each one before you put any of them side by side.

The takeGive the write-up to the candidate. They were in the room, they can dispute what you recorded, and they are the only reader a small team gets for free. Most managers will not do it, which is roughly why it works: a note written in the knowledge that its subject will read it comes out about the work rather than about the impression. What comes back that matters is the part of your record the person it describes does not recognise.

Where Olive fits

Open a role and see what the work shows

A single interviewer can record a candidate describing how they would check a confident claim, and not what they do when one is in front of them. Olive puts the checking itself in front of the candidate as work, and a human reviewer writes six findings, each carrying the timestamped excerpt it rests on.

Rank your shortlist

Why doesn't adding people fix this?

Because structure was doing the work, not headcount. In the 2022 re-analysis of the selection literature, structured interviews correlate .42 with rated job performance and unstructured ones .19, which makes the gap between two ways of running an interview larger than the gap between most pairs of methods 1. Nothing in either number is about how many people were in the room.

A panel contributes three things, and only one of them needs a second person. It asks everyone the same questions. It forces evidence into writing before verdicts get traded. And it pools independent readings at the end. The first two are procedure, they are free, and they are where the measured difference sits. Note the limits on that .42 as well: structured there is a research coding of interview format, the underlying studies mostly scored people already doing the job, and both figures replaced higher ones that circulated for two decades.

The half of the panel advice that small teams reach for first is the half that can hurt you, which is adding one more opinion on top of the process you already ran, the same trade behind spending a hiring manager's round on the role rather than on a second read of the candidate. Layering an intuitive judgment over mechanical predictors has a long history of making prediction worse: in a 1943 study of University of Minnesota undergraduates, high school rank plus a college aptitude test correlated .45 with academic achievement, while the same two predictors plus counselors' intuitive judgment correlated .35 2. The same review notes that the belief that some interviewers are simply better than others is not supported by the evidence on variance in interviewer validity. An admissions study from the 1940s is not a hiring study. Take the ordering from it and leave the numbers where they are.

Candidates can now rehearse with a model, and a fluent, well-structured, confidently delivered answer is the exact impression an unchecked interviewer trusts most. A bigger bench of unchecked interviewers does not help with that. Whether a structured round still separates people once everyone is coached is a question worth answering before you add rounds.

Take notes with no ratings in them

Write down what was said and done, in the candidate's words and yours, and nothing else. No adjectives, no impressions, no provisional rating in the margin. The moment a rating exists it starts recruiting evidence toward itself, and the notes stop being a record and become a case. A rating-free column is the cheapest structural change available to a single interviewer.

Two mechanics do most of the remaining work, and both are free:

  • Same questions, same order, every candidate. Otherwise you are comparing conversations, and the candidate who steered theirs best wins.
  • One follow-up per answer, prepared in advance. Decide before the search which claim in each answer you will push on, so the push does not depend on how interested you already are.

Start the call with the questions rather than with small talk, or at least do not write during the small talk. In a study of 189 accounting students in structured mock interviews, the interviewer's overall impression formed during the rapport-building period before any structured question correlated .42 with that same interviewer's later structured score, falling to .25 and .24 when a different interviewer supplied the structured score 3. That is a student sample in mock interviews, and much of the .42 is one interviewer's impression predicting their own later rating. The paper's own reading is not that first impressions are pure contamination: the effect ran through rated competence, and it survived when the structured rating came from someone who had done no rapport building. It is not evidence that anyone decides in three minutes, and it should never be cited for that. What it does show is a leak between the unscored part of the call and the scored part, which is a thing you can plumb.

Rehearsed answers are the other reason to write verbatim. A polished narrative reads as strong evidence and summarizes into almost nothing, which is what makes the identical STAR answer from every candidate so hard to grade from memory.

Score the record the next morning, not in the room

Put a night between the evidence and the verdict, then score the written record against your bar rather than against your memory of the person. An explicit rule beats an overall impression by a wide margin: combining candidate data mechanically correlated .44 with job performance against .28 when experts combined the same kinds of data by judgment 4. The rule can be a checklist you add up.

Two caveats before that number does too much work. The job performance comparison rests on nine studies, other criteria in the same paper show much smaller gaps, and mechanical here means something as plain as unit-weighted scores on a written scorecard, which the authors say outright. Nothing in it licenses a number standing for a person, and none of it is an argument for handing the decision to software. The finding is that a rule applied identically every time beats an impression formed freshly every time.

The scoring sitting itself has one rule that costs nothing and changes a lot: score every candidate against the bar before you put any two of them next to each other. Comparison is where a single interviewer's drift compounds, because the second candidate gets measured against the first, and by the fifth you are scoring a tournament. Finish the scorecard, close it, then open all of them together. The same discipline in a team setting is scoring before the debrief, and it exists for exactly this reason.

When the record does not settle it, the fix is another question. Go back with one specific follow-up on the claim the decision turns on, and write the answer down verbatim. Follow-up questions that expose whether someone understands their own answer are the highest-yield thing you can add to a second call.

Send the candidate your write-up

The cheapest second reader is the candidate, and they were in the room. Send the factual record and keep the score to yourself: here is what I wrote down that you said, tell me what I got wrong. Somebody who can dispute the record is a real check on it, and a note that cannot survive being read back was never evidence in the first place.

Expect three things back. Corrections, which are the point. Additions, which you weigh against the bar like anything else. And occasionally silence, which is information about how much the record mattered to them. None of this obliges you to change a decision. It is a factual check on a factual document, and it should not be framed as an appeal.

Small teams have the most to gain because they write the least down. In one large recruiting dataset, scorecard completion runs near 49% at organizations under 25 employees against roughly 72% at organizations of 500 or more, and around 38% of scorecard pairs from the same interview differ by at least one point, with nearly half of those one-point gaps falling between 2 and 3, which is the yes-or-no boundary on a one-to-four scale 5. That is a customer base skewed toward venture-backed technology employers, and the pairs are compared against each other with no outcome behind them, so nothing in it says who was right. It supports one claim only, which happens to be the one that matters here: ratings are least stable exactly where the decision gets made, so the evidence underneath a rating has to be inspectable.

And the limit, stated plainly, because the panel advice never states its own: none of this removes your bias. You will still like the candidate who reminds you of a good hire. What the two passes buy is that your reasons exist in writing, dated, before the verdict, where you and the candidate can both look at them. Judgment made arguable is not judgment made objective, and the difference is worth being clear about with yourself. If you have never done the work you are hiring for, borrow the bar before you borrow anything else.

See how it works

Common questions

How long should the write-up take?

Ten minutes, immediately after the call, while the phrasing is still exact. Type quotes during the interview if you can do it without going quiet, and fill the gaps straight afterwards rather than at the end of the day. What you are protecting is the candidate's actual words: a paraphrase written three hours later has already been shaped by the impression you formed, which is the thing the whole procedure is trying to keep separate from the evidence.

Does interviewing alone create legal exposure?

The exposure does not come from the headcount. It comes from candidates being asked different things, from a decision with no written basis, and from questions that stray onto protected ground, all of which one interviewer can do and a panel can do too. A written record of the same questions asked of everyone is what helps if a decision is ever challenged. Rules on notice, recording and automated tools vary by state and city, so run anything you add to the call past counsel before it goes live.

Can I use an AI notetaker as the second reader?

It is a transcription tool, not a reviewer. A transcript is genuinely useful because it holds the candidate's exact words and frees you to listen, but it only ever sees what you gathered, so asking a model to assess the candidate from your notes returns the frame those notes were written in. Machine transcripts are also uneven: measured accuracy on commercial systems is worse for Black speakers than for white speakers, and worse again for d/Deaf and hard-of-hearing speakers, so the transcript is not neutral evidence either. Consent and recording rules differ by state, and the candidate should be told before the call, not during it.

What if I already made up my mind in the first four minutes?

Finish the questions anyway and write the answers down anyway, because that record is what you will score in the morning. Then run the specific test: find the two strongest pieces of evidence against your conclusion in your own notes. If neither exists, the interview did not test your first impression, which is a fault in the questions rather than proof the impression was right. Book a second, shorter call built entirely around the thing you did not probe.

Should I tell candidates I am the only interviewer?

Tell them in the invitation, along with what the call covers and how the decision gets made. Candidates prepare differently for a single decision-maker than for a screening round, and the ones you want will ask better questions when they know who they are talking to. It also sets up the write-up: telling somebody in advance that they will see the record of what they said makes the offer land as procedure rather than as an unusual gesture.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the claim that structure rather than headcount is the measured lever in interviewing: .42 for structured interviews against .19 for unstructured ones.
  2. 2. Stubborn Reliance on Intuition and Subjectivity in Employee Selection Industrial and Organizational Psychology, 1(3), 333-342, Table 1 (Scott Highhouse), 2008. edbatista.com Supports the claim that layering an intuitive judgment on top of mechanical predictors can lower accuracy (Sarbin's .45 falling to .35), and that better interviewers are not an evidenced category.
  3. 3. Initial Evaluations in the Interview: Relationships with Subsequent Interviewer Evaluations and Employment Offers Journal of Applied Psychology, 95(6), 1163-1172 (Murray R. Barrick, Brian W. Swider and Greg L. Stewart), 2010. homepages.se.edu Supports the claim that the unstructured rapport period leaks into the structured scores, with the paper's own limits attached.
  4. 4. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis Journal of Applied Psychology (American Psychological Association), 98(6), 1060-1072, doi 10.1037/a0034156; full-text copy opened at gwern.net, 2013. gwern.net Supports scoring the written record against an explicit rule rather than forming an overall impression: .44 mechanical against .28 clinical for job performance.
  5. 5. Recruiting Operations Benchmarks | 2026 Talent Trends Report Ashby, 2026. ashbyhq.com Supports the scorecard completion gap by employer size and the finding that interviewer ratings disagree most often right at the yes-or-no boundary.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.