Interviewing

Plan Three Questions and Nine Follow-Ups, Not Twelve Questions

Budget two or three follow-ups per interview question, and three or four planned questions in an hour rather than a dozen. Go down instead of sideways: ask for the specific instance, then the decision inside it, then what went wrong and what was done about it. Stop when the answer turns general again, because that is where the next question stops adding evidence. Write the ladder into the guide so every candidate gets the same rungs.

The takeThe failure worth worrying about is not asking too few follow-ups. It is asking three of them to the candidate you already like and accepting the first answer from the one you do not. That asymmetry turns a loop that looks structured into unequal treatment, and no rating scale catches it, because the scale only ever sees the answers that were allowed to develop. Count probes per candidate. It is one number, it costs nothing, and it turns depth into something every candidate can reach.

Where Olive fits

Open a role and see what the work shows

Rungs two and three are where an interview runs out of room, because describing a check is not the same act as running one. Olive puts the work itself on the candidate's own clock as a 50-to-70-minute session, and a human reviewer writes six findings with the timestamped excerpt behind each one.

Rank your shortlist

How many follow-ups is the right number?

Two or three per question, planned in advance, with three or four questions in an hour instead of a dozen. The number is a budget rather than a rule: an hour holds a fixed quantity of exchanges once you subtract the opening, the role pitch and the candidate's own questions, and spending them across a dozen separate topics buys a dozen opening answers with nothing underneath any of them.

The incumbent advice has this backwards. Probing is taught as a repair tool, something you reach for when the first answer was vague or incomplete, which makes it contingent and optional. The first answer is the part that was prepared. Treating the probe as a fix for a bad answer means you probe least exactly where the preparation was best, and the candidate who rehearsed hardest gets the shallowest read.

Depth is also where the payoff sits in the evidence. The 2022 re-analysis of the selection literature puts structured interviews at .42 and unstructured ones at .19 1. Those are corrected correlations with supervisor-rated performance, and both carry wide credibility intervals. Structure there means every candidate facing the same questions and being judged against the same written scale. It says nothing about covering more ground, and written scales are only possible when there are few enough questions to write them for.

What almost nobody is trained on is the half that happens after the answer, and leaving it to whoever is naturally good at interviewing does not survive contact with the evidence. Highhouse's review reports that although it is commonly accepted that some interviewers are better than others, research on variance in interviewer validity suggests the differences are due entirely to sampling error, and it reproduces a 1943 admissions result in which high school rank plus an aptitude test correlated .45 with academic achievement while the same two predictors plus counselors' intuitive judgment correlated .35 2. That 1943 study is admissions rather than hiring, its sample size is not given on the page reporting it, and none of it says interviewers add nothing. The narrower point is the one that matters here: write the probes into the guide, because there is no evidenced category of interviewer whose own judgment reliably stands in for one.

Take the answer down three rungs before you move on

Ask for the instance, then the decision inside it, then the repair. Rung one gets the occasion and its facts. Rung two gets the choice the person made and what they were weighing. Rung three gets what went wrong afterwards and what they did about it. Each rung is harder to answer from a rehearsal than the one above it, because each depends on more of what actually happened.

Worked, for a question about shipping something under a deadline:

  • Rung one. "What was the last thing you shipped later than you wanted to?" You are after an occasion with a date, a name and a scope, not a category of experience.
  • Rung two. "What did you cut, and who decided?" The decision is the part a generic answer skips, because a generic answer has no constraints in it.
  • Rung three. "What broke afterwards, and what did you change?" This is the rung where an account either has consequences in it or does not. Consequences are specific and unglamorous, and a rehearsed answer tends to end at the success.

Going sideways instead, to a new topic, resets the candidate to their prepared material every time. That is the whole cost of a twelve-question loop: twelve fresh starts, none of them past the first rung.

The rungs compete with the warm-up for the same hour. In a study of structured mock interviews with 189 accounting students, the interviewer's overall impression formed during the rapport-building small talk before any structured question correlated .42 with that same interviewer's later structured score, falling to .25 and .24 when two other interviewers supplied the score 3. It is a student sample with mock interviews, and the authors are careful that the effect ran through rated competence rather than liking, so it is not evidence that minds are made up in the first three minutes. It is a reason to keep the unscored opening short.

When should you stop probing?

Stop when the answer turns general, when you have enough for the rating anchor, or when you have spent the minutes you budgeted. The first is the useful signal, and it reads the return: while the specifics keep arriving there is more to get, and once they stop, another rung adds nothing. Two rungs of generality in a row means the return has gone, and a third question will only produce a better-phrased version of the second.

Saying it out loud helps, because a stopping rule nobody can state is a stopping rule that varies by how much the interviewer is enjoying themselves. The version worth writing into the guide: *go down until the answer stops naming things, then move on.* Names, numbers, dates, constraints and other people are what a real account keeps producing. Adjectives are what replaces them when it runs out.

The second stop matters more than it sounds. Stop when you have enough to write the anchor's language into the scorecard, even if the answer is still going. The point of the probe is evidence for a rating, not a satisfying conversation, and an interviewer who keeps digging after the evidence has arrived is spending a later candidate's minutes.

Only one exception really holds. If the candidate has no comparable experience at all, further rungs will only measure how gracefully they handle not knowing, which is worth something but is not what the question was for. Say so, switch to the hypothetical version, and note in the scorecard that you switched. Which follow-ups expose whether someone actually understands their own answer covers the harder case, where the account is fluent and the understanding under it is thin.

Track how many rungs each candidate reached

Add one column to the scorecard: how far down each planned ladder this candidate actually went. It costs an interviewer four seconds and it is the only record showing whether depth was distributed evenly. The ratings cannot carry that, because an answer cut off at rung one leaves no trace in the score it produced.

This is where an interview that looks structured quietly stops being one. Uneven probing is unequal treatment produced by a process everyone believes is fair, and the ratings will look defensible afterwards because the shallow answers really were thinner. The cause is not in the ratings. It is in who was given the chance to say more.

The reason to take this seriously is what the measured gaps track. A resume audit of 36,880 applications to 9,220 job advertisements for new US college graduates reported callbacks 28 to 43 percent lower for Black men, Black women, White women and Hispanic men than for otherwise identical White men in management occupations, with the widest gaps in roles combining high analytical and interpersonal demands with low routine content 4. That is a preprint under review, it measures callbacks rather than hires, it covers new graduates only, and the discretion mechanism is the authors' proposed explanation rather than something the experiment manipulated. Read at its own weight, it still points at the same place: the differences are widest where the evaluation is least specified, and an unplanned probe is exactly that.

Two fixes, both cheap. Put the planned ladder in the guide beside each question, so every interviewer has the same rungs available whether or not they thought of them. Then require the probe count in the scorecard, and read it in the debrief before you read the ratings. If the candidates who present well are consistently getting three rungs and the others one, whether every candidate has to face exactly the same questions is the conversation your loop needs next.

See how it works

Common questions

Is three follow-ups a hard rule?

No, it is a budget you set before the interview so the decision is not made under time pressure in the room. Some questions are exhausted at rung two and some are still producing at rung four. What matters is that the ladder was planned, that every candidate has access to the same rungs, and that the interviewer stops for a stated reason rather than because the clock ran out on the last candidate of the day. Write the planned rungs down; treat the count as a floor for depth rather than a cap.

Does probing break the consistency a structured interview needs?

Not if the ladder is planned in advance and available to everyone. Consistency lives in the question set, the rating anchors and the opportunity to be probed, not in identical wording. Two interviewers can phrase the same probe differently and still be running the same interview. What breaks consistency is one candidate getting three rungs and another getting one, because the rating then reflects how much room each person was given rather than what each of them did.

What do I do when the candidate keeps answering in generalities?

Ask once for a specific occasion, plainly: "Give me the most recent time this happened." If the second answer is still general, note it and move on rather than escalating. Persistent generality is itself information, but it can also mean the question missed their experience, a nervous candidate is reaching for safe ground, or the role they held did not include the thing you asked about. Record what they said, not your inference about why, and let the pattern across questions carry the weight.

How do I stop an interviewer from probing only the candidates they like?

Make the probe visible. Put the planned rungs in the guide so nobody has to invent them, add a probe-count column to the scorecard, and read the counts in the debrief before the ratings. Interviewers rarely do this deliberately, which is why training alone does not fix it: the fix is a record that makes the asymmetry countable. If one interviewer's counts are consistently uneven across candidates for the same role, that is a calibration conversation, not a disciplinary one.

Do I need to write down what the candidate said, or is the rating enough?

Write down what they said. A rating with no excerpt under it cannot be reconstructed a week later, cannot be argued with in a debrief and cannot be defended if the decision is ever questioned. It also changes the interviewer's behavior: someone who knows they have to record the evidence asks the question that produces evidence.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .42 and .19 corrected estimates for structured and unstructured interviews, used here to argue that common questions and common written scales, not question coverage, are what the estimate rests on.
  2. 2. Stubborn Reliance on Intuition and Subjectivity in Employee Selection Industrial and Organizational Psychology, 1(3), 333-342, Table 1 (Scott Highhouse), 2008. edbatista.com Supports the claim that stable differences in interviewer validity are not evidenced, and the Sarbin (1943) result in which adding intuitive judgment to two mechanical predictors lowered prediction from .45 to .35.
  3. 3. Initial Evaluations in the Interview: Relationships with Subsequent Interviewer Evaluations and Employment Offers Journal of Applied Psychology, 95(6), 1163-1172 (Murray R. Barrick, Brian W. Swider and Greg L. Stewart), 2010. homepages.se.edu Supports the claim that the unstructured rapport-building opening leaks into the structured score that follows it (.42 same interviewer, .25 and .24 across interviewers), in a student mock-interview sample.
  4. 4. Hiring Discrimination and the Task Content of Jobs: Evidence from a Large-Scale Resume Audit arXiv preprint arXiv:2604.01933 (Braun and co-authors), version 2, July 2026, 2026. arxiv.org Supports the claim that measured callback gaps are widest where evaluation is most discretionary, cited as an unreviewed preprint measuring callbacks rather than hires.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.