Interviewing
What Follow-Up Questions Expose Whether Someone Understands the Answer They Gave?
Three follow-up questions, in order, on the candidate's own case, expose whether someone understood the answer they gave. Specify: which alternative did you rule out, and what number or constraint ruled it out? Invert: what would you have done if that constraint had moved the other way? Falsify: what would have to be true for this answer to be wrong, and what would you have checked? Stop when one rung gives you the evidence, and ask the same rungs of every candidate. Thin answers are not proof of AI use.
The takeThe ladder's real subject is the interviewer. Naming, before the round, the option a competent person would have rejected forces you to decide what the job actually requires, and most of the value lands before anyone is in the room. Skip it and you are grading how comfortable someone sounds, which is a separate skill and unevenly distributed. Rung one is the only one you can ask without knowing the work, which is, I would guess, why so many panels stop there and call it three. If you cannot name that option yourself, you have not written an interview. You have written a conversation.
Where Olive fits
Open a role and see what the work shows
A follow-up can get a candidate describing the alternative they discarded; it cannot show you them discarding one. Olive puts that in front of the candidate as work: an occupational assignment with an AI assistant that will overreach, read by a human reviewer who writes six findings against the moments they happened.
Rank your shortlistWhy does a fluent answer prove so little?
Because fluency and understanding come apart under one specific demand: producing the causal chain. People rate their understanding of ordinary devices highly, then rate it far lower after being asked to write a single step-by-step explanation of how the thing works. The effect is strongest for explanatory knowledge and weaker for facts and procedures 1. An interview answer is a summary. The follow-up is where the chain is either there or missing.
The second problem sits on your side of the table. When explanations carried extra detail that added nothing logically, non-experts and students rated them as better explanations, and the added detail did most of its work on the weak ones, masking problems that would otherwise have shown 2. Experts were not moved. If you are interviewing outside your own specialty, technical texture reads as depth.
That is the argument for probes whose answers you can check without being expert in the work. You are not judging whether the rejected approach deserved rejecting. You are checking whether a specific thing was rejected, for a specific reason, at a specific point: detail a summary never carries and a practiced answer rarely invents on the spot. Well-structured stories are the baseline now rather than a signal, which is why the same polished STAR answers keep arriving.
Run the ladder: specify, invert, falsify
Ask for the discarded alternative, then move a constraint, then ask what would prove the answer wrong. Each rung wants something the first answer did not contain, and each is harder to produce without having done the work. Climb only while the answers hold. Three rungs on the one claim the role depends on beats one rung on six claims.
Take a real answer: "We moved payment methods above the fold and checkout conversion went up about 12%."
- Specify: the discarded alternative. "What else was on the table that sprint, and what took it off?" Any account of finished work carries the option that won. Only someone who was in it carries the one that lost, and the reason it lost. A real answer names both: the alternative, and the number, constraint or objection that ended it. A thin answer names the alternative and stops: "we looked at a few approaches and this one fit best" is the shape to listen for.
- Invert: the moved constraint. "If mobile had been eighty percent of that traffic instead of forty, would you have made the same change?" Recall stops working here; the answer has to be derived while you watch. Someone who understood the decision re-runs it and usually lands somewhere specific and slightly untidy. A rehearsed answer returns the original argument with its polarity flipped and no new detail in it.
- Falsify: the check, not the caveat. "What would have to be true for that 12% not to be your change, and what would you have pulled to find out?" The second half is the one that matters. "Seasonality could have confounded it" is a caveat anyone can produce; "I'd have pulled the same two weeks from the prior year against the control cohort" is a check. This rung sits closest to the behavior the job now needs, because a confident wrong answer from an assistant is only caught by someone who habitually asks what would make it wrong, which is the thing worth testing directly.
A ladder without limits is just a stress test, so hold to two rules. Stop climbing the moment you have your evidence: a candidate who answers rung two cleanly has already told you what rung three would sound like. And ask the same rungs of every candidate on the same kind of claim, or you have run a different interview for each person and then compared them anyway.
What counts as a discarded alternative in your field?
Something a working practitioner would plausibly have considered and rejected, for a reason that occupation recognizes: a rejected architecture, a rejected channel, a rejected model specification, a rejected rating basis. If you cannot name one plausible rejected option before the interview starts, you cannot hear whether the answer is real, so the ladder gets written per occupation, not per company.
| Field | Ask for the option they dropped | What a credible answer names |
|---|---|---|
| Software engineering | The approach you started and abandoned | A library, schema or service by name, and the constraint that killed it (a latency budget, a migration cost, an on-call burden) |
| Data and analytics | The specification that did not survive | A feature or model form dropped on a diagnostic: leakage found, a residual pattern, a sample too thin to split |
| Marketing | The channel or message you stopped funding | A test cut at a number (cost per qualified lead, a payback period, a message that lost to control) |
| Financial analysis | The assumption you refused to carry | A driver taken out of the model, and what the sensitivity did when it moved |
| Product management | The requirement you cut from the spec | The requirement, who it cost, and what they were told |
| Legal and contracts | The fallback position you dropped | A clause position abandoned against a playbook or a signed precedent, and the exposure that made it indefensible |
| Revenue cycle | The denial you conceded instead of appealing | The payer rule that made the appeal unwinnable, and what the write-off came to |
Writing your own row takes about ten minutes. Take one decision the job actually made last quarter, write down the option the team rejected and the constraint that rejected it, and you have both the probe and the standard you grade against. Do it before the first interview rather than after the third, which is the whole argument for writing the rubric first.
One caution on transfer: the rungs work on the candidate's material, not yours. Asking a marketer to invert a constraint in your funnel tests how quickly they learn your business. Asking them to invert a constraint in a campaign they ran tests whether they understood it. Keep the case theirs.
How do you ask without running an interrogation?
Say what you are doing, ask everyone the same rungs, and let "I don't remember" be a real answer. Structured interviews put the same predetermined questions to every candidate in the same order and grade the answers against the same standards for acceptable answers 3. A probe improvised for one person and skipped for the next is a different interview wearing the same name.
Frame the probe as a check on your own question. Clinicians do this deliberately with teach-back: the patient is asked to say the plan back in their own words, and the clinician frames it as making sure the explanation was clear rather than as a test of the patient 4. The same sentence works in an interview ("I want to make sure I asked that well: which option did you rule out?") and it costs you nothing.
- One rung, then silence. The common failure is the interviewer filling the pause with a hint, which hands over the specifics the rung was asking the candidate to supply.
- Grade content, not delivery. Talking well about your own work is a separate skill from doing it, and it is not evenly distributed. Score whether the constraint was named, not how smoothly it arrived.
- Offer the same substitute to everyone. Some people answer rung two far better in writing with ten minutes. If that is fair for one candidate it is fair for all, and it changes what the round measures less than it appears to.
- Write the answer down as it happens. A follow-up that decides who advances is part of a selection procedure, and carries the same consistency expectations as the rest of the process 5.
Remote rounds raise the obvious question about assistance mid-interview. The honest read: rung two needs specifics from the candidate's own case, which a model cannot supply unless it has been fed them, so those answers thin out rather than arrive late. That is a difference in substance, not a finding. Decide on what the answer contains, and treat assistance during a live interview as its own policy question rather than an inference you draw in the moment.
Where does the ladder stop?
At description. Three rungs get a candidate reconstructing a decision from months ago, and reconstructions improve with interview practice. You learn whether the causal chain can be produced on demand. You do not learn what the person does when a confident, wrong answer arrives on a problem they have never seen, which is most of the work now.
The fix is not a fourth rung. It is putting a decision in front of the candidate while you watch: a work sample where the alternative gets discarded live instead of recalled. The same three questions then make a good debrief rubric: which option did you drop, what would have changed it, what would prove this wrong. What capable AI work actually looks like is a short list, and every item on it is an act rather than an account of one.
Keep the ladder anyway. It costs three questions, it fits inside a thirty-minute screen, and it is the cheapest thing in the round that separates a person who did the work from an answer about the work. See how Olive measures this.
Common questions
How many follow-ups should one answer get?
Two or three, on one claim. Each rung costs a minute or two of real time, so spend it on the claim the role actually depends on and let the other answers stand. Running all three rungs on every answer turns a forty-minute round into an endurance test and returns worse information, because a candidate probed six times starts hedging everything. Pick the claim before the interview rather than during it.
What if the candidate says they don't remember?
Take it as an answer and move. Memory for a decision made a year ago is genuinely patchy, and someone who says so is often being more accurate than someone who produces a tidy reason on the spot. Ask for a decision they do remember, or move the case closer: last month rather than last year, the thing they owned rather than the thing their team shipped. Grade the answer you get, and offer the same substitution to everyone.
Can a candidate prepare for these follow-ups?
Yes, and preparation is not a problem here. Preparing for rung one means going back through your own work and recovering what you rejected and why, which is the thing you were trying to find out whether they could do. What preparation does not cover is rung two on a claim the interviewer picked: moving a constraint nobody anticipated needs the actual decision, not a rehearsed account of it. A candidate who answers all three has told you what you wanted.
Is this a way to tell whether an answer was AI-assisted?
No, and using it that way produces wrong calls. A model can write a plausible discarded alternative and a plausible caveat. What it cannot supply is the specific number or constraint from a case it was never told about, which is why those answers thin rather than fail outright. Thin answers also come from nerves, from working in a second language, and from people who did the work and narrate it badly. Grade what the answer contains against a written standard, and keep authorship out of the decision.
Where does Olive fit next to a follow-up ladder?
After it. Olive is an employer-purchased assessment: a 40-to-60-minute assignment built for the occupation, worked with an AI assistant, after which a human reviewer writes six findings, each attached to a timestamped moment in the session. The report carries no composite number and no hiring recommendation, and the candidate is granted the same document the employer reads. The free tier covers ten attempts a month.
References
- 1. The misunderstood limits of folk science: an illusion of explanatory depth ✓ pmc.ncbi.nlm.nih.gov Self-rated understanding drops sharply after writing a step-by-step causal explanation; the effect is specific to explanatory knowledge.
- 2. The Seductive Allure of Neuroscience Explanations ✓ pmc.ncbi.nlm.nih.gov Logically irrelevant detail made explanations more satisfying to non-experts and masked problems in weak ones; experts were unaffected.
- 3. Structured Interviews ✓ opm.gov All candidates are asked the same predetermined questions in the same order, and all responses are evaluated on the same rating scale and standards for acceptable answers.
- 4. Use the Teach-Back Method: Tool 5 ✓ ahrq.gov Comprehension is checked by asking the person to state the plan in their own words, framed as checking the explanation rather than testing the person.
- 5. Employment Tests and Selection Procedures ✓ eeoc.gov A step that decides who advances is a selection procedure and has to be applied consistently.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.