Interviewing
Behavioral Questions Still Work If You Ask for Events, Not Stories
Behavioral interview questions still predict performance after candidates rehearse them with AI, because the story never carried the prediction. The format does: the same questions, in the same order, rated against a scale set in advance. Rehearsal reaches the narrative arc, which was always public and now arrives polished for anyone. It cannot reach the detail that only presence supplies, so ask for a date, a named artifact, the person who disagreed and what it cost, then score how much the answer narrows when you ask for the next layer.
The takeRetire the phrase gold standard and keep the method. Past behavior predicts future behavior was always a slogan sitting on top of a narrower finding about interview structure, and the slogan is what made people believe the story itself carried the weight. The current revision of the advice is worse: repurposing behavioral questions as an authorship test asks an interviewer to grade prose style, in a room, from memory, with a hiring decision attached. That is not a stronger use of the method. It is the method pointed at a question it cannot answer.
Where Olive fits
Open a role and see what the work shows
An event somebody lived through supplies detail nobody rehearsed. Olive gets that detail by making the event happen during the assessment: a role-grounded assignment with an AI assistant that will overreach, written up as six findings stated as demonstrated, partly demonstrated or not demonstrated.
Rank your shortlistDo behavioral questions still predict job performance?
Yes, and the change is smaller than it feels. Nothing about a candidate preparing with a model touches the mechanism behind the numbers, which is comparability: the same questions, asked in the same order, rated against a scale agreed in advance. What preparation removed is the accidental advantage that used to go to people who were articulate under surprise, and that was never the part doing the predicting.
The version in most training decks is stale. In the 2022 re-analysis of the selection literature, structured interviews estimate at .42 and unstructured interviews at .19, pooled by sample-size weighting from two earlier meta-analyses, McDaniel and colleagues in 1994 and Huffcutt and colleagues in 2014 1. Before any correction the raw observed validities were .32 and .13. Both figures are corrections applied to correlations between a method and supervisor ratings, pooled across many jobs. They are not accuracy rates and not a promise about any one company's round, and the older .51 figure people still quote is the 1998 Schmidt and Hunter estimate this paper revised downward, printed in the same table beside it.
The gap is the useful part. The distance between a structured and an unstructured version of the same conversation is larger than the distance between most pairs of methods on that list, which means the format decision outranks the question-selection decision, and it is the reason changing the question bank every quarter buys less than it costs. A rehearsed candidate does not touch the format.
What does a rehearsed answer actually cost you?
One weak signal, and it was already the weakest one. Before preparation was free, an unrehearsed answer gave you a rough read on composure and recall, and both were doing less work than anyone admitted. What survives untouched is everything that depends on the candidate having been present: the particulars, the sequence, the cost, and how the answer behaves when you ask for one layer more.
The tempting fix is to abandon structure and just talk to people, and that is the one move the evidence rules out. In a controlled test, 76 undergraduates predicted classmates' semester grades; predictions made after conducting an unstructured interview correlated .31 with the actual outcome, while predictions from prior cumulative grades alone correlated .65 2. In a related study, 96 of 169 participants chose to conduct an interview in which the interviewee answered at random over conducting no interview at all. Undergraduates predicting a classmate's grades are not managers predicting job performance, and the effect sizes do not transfer. The mechanism does: low-diagnostic information dilutes valid information, and interviewers construct meaning out of almost anything.
Loosening the round in response costs far more than the rehearsed answer ever did.
Ask for an event instead of a story
Six asks, and each one narrows the answer rather than extending it. When was this, to the month. What artifact exists that came out of it. Who else was in the room and where do they work now. What the number was and what it was measured against. What went wrong on the way. Who found out first. A story answers none of these without effort. An event answers all six without hesitation.
The move is subtle in the wording and large in what comes back. Instead of tell me about a time you handled a difficult stakeholder, ask who the stakeholder was, what specifically they wanted, what you sent them, and what they said next. Instead of describe a project you are proud of, ask for the week the project nearly stopped. Both versions are guessable in advance, and only the second keeps producing detail under a follow-up.
Watch what happens in the other direction too. An answer that gets vaguer as you ask for particulars is not proof of anything by itself, since people forget, misremember and get nervous. Treat it as a prompt to ask differently rather than a verdict: offer a different entry point, ask about the artifact instead of the timeline, and see whether the detail comes back. Why every candidate suddenly gives the same polished STAR answer covers the slate-wide version of this, and what follow-up questions expose whether someone actually understands the answer they just gave covers the chain.
Why is grading an answer for sounding rehearsed a mistake?
Because the error it produces is not random. The only version of this judgment anyone has measured is the software one, and it lands hardest on people composing in a second language. Seven detectors run over 91 human-written TOEFL essays by non-native English speakers produced an average false-positive rate of 61.3%, with 97.8% of those essays flagged by at least one tool, while the same detectors were near-perfect on essays by US eighth-graders 3.
That study is 91 academic essays, one language-background cohort and seven tools as they stood in 2023, so the number is not a constant and it is not about interview answers. The direction of the error is the durable part, and an interviewer running the same instinct from memory has none of the consistency a tool at least has. There is also no appeal: a candidate marked down for sounding rehearsed is never told, so the judgment is never tested against anything.
There is a fairness cost to abandoning the method too, and it runs the opposite way from what people expect. The same 2022 analysis pairs each method with its mean Black-White standardized subgroup difference, and structured interviews carry the highest validity on that table, .42, alongside a d of .23, against .79 for cognitive ability tests and .67 for work samples 1. Those d values are borrowed from other meta-analyses, several drawn from non-applicant samples, and a subgroup difference is not adverse impact, which depends on how scores are used and who applied. But the freeform chat that replaces a dropped behavioral round sits at a d of .32 with a validity of .19, worse than the structured version on both counts at once.
If a panel keeps reaching for sounding rehearsed, the underlying problem is usually a manager with no other vocabulary for a thin answer. How to stop hiring managers rejecting candidates for sounding like AI is the version of that conversation worth having before the next debrief.
Common questions
How many behavioral questions should a round include?
Three or four, each taken two or three layers deep, in a 45-minute round. Coverage is the trap: eight questions at one layer each produce eight clean arcs and nothing a debrief can compare. Depth is what a rehearsed answer cannot supply, and depth costs clock, so the number of questions has to come down as the number of follow-ups goes up.
What is the difference between a behavioral and a situational question?
A behavioral question asks what somebody did; a situational one asks what they would do. The second is entirely hypothetical, which makes it cheap to answer well and cheap to prepare, since there is no event underneath it to check. Situational questions still have a use for candidates with no relevant history, such as early-career hires, but they carry no particulars and should not be scored as though they do.
Can a candidate legitimately use notes in a behavioral interview?
Yes, and say so at the start. Notes help people with memory differences, people interviewing in a second language, and anyone nervous enough to lose a detail they actually have. What notes cannot supply is the answer to a follow-up nobody could have anticipated, which is where the round is decided anyway. Allowing them removes an unfair advantage rather than creating one.
Do behavioral questions work for early-career candidates?
Yes, if the events you accept are the events they have had. Coursework, a part-time job, a club, a family responsibility and an open-source contribution all contain decisions, disagreements and things that went wrong. The question to avoid is the one that assumes a workplace, because it tests exposure rather than judgment. Ask for an event, name a wider range of acceptable settings, and run the same follow-up chain.
Should the same interviewer ask every candidate the same questions?
The same questions, yes; the same interviewer is a bonus rather than a requirement. Comparability is what the validity evidence rests on, and it breaks when one candidate gets four questions and another gets a conversation. If several interviewers share a round, give them the same script, the same rating scale and the same worked examples of a strong and a weak answer, then compare written evidence rather than impressions in the debrief.
References
- 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the .42 and .19 validity estimates for structured and unstructured interviews, the raw observed .32 and .13, and the subgroup difference figures in Table 3.
- 2. Belief in the unstructured interview: The persistence of an illusion sjdm.org Supports the claim that unstructured interviewing dilutes valid information: predictions after an interview correlated .31 against .65 from prior grades alone.
- 3. GPT detectors are biased against non-native English writers pmc.ncbi.nlm.nih.gov Supports the claim that judging text for machine authorship misfires on second-language writers: 61.3% average false-positive rate across seven tools on 91 TOEFL essays.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.