Interviewing

A Good Question Is One a Prepared Stranger Cannot Answer

An interview question is good when the answer depends on something only the candidate did, and when the difference between a strong and a weak answer is visible to whoever is listening. Run four checks against the question bank you already use: a well-prepared stranger cannot answer it convincingly, it names a specific instance instead of inviting a policy statement, it carries a follow-up only someone who was there could survive, and it has a written anchor saying what strong, adequate and weak answers contain.

The takeMost loops are over-questioned and under-anchored, so the useful Monday move is subtraction. Score the bank you already own against the four checks, delete everything that fails, and spend the recovered minutes going deeper on what survives. Keep one deliberately preparable opener anyway, because how a candidate spends preparation time is worth seeing, and write down what you saw as context rather than as something rated. A shorter guide that everyone actually follows beats a long one nobody finishes.

Where Olive fits

Open a role and see what the work shows

An interview can capture someone describing how they would check a confident claim; it cannot capture them checking one. Olive puts that in front of the candidate as work instead: an occupational assignment done with an AI assistant that will overreach, and a human reviewer who writes six findings with the moment behind each one.

Rank your shortlist

What separates a good interview question from a bad one?

Two properties: the answer has to depend on something only this candidate did, and the difference between a strong and a weak answer has to be visible to the person listening. Neither is the question's category. A question that fails the first collects preparation. A question that fails the second collects a conversation, and a conversation is not evidence anyone can argue with in a debrief.

The second property is the one hiring teams under-build. Written anchors saying what strong, adequate and weak answers actually contain are what turn a rating into something two people can check against each other. In a meta-analysis of selection and admissions decisions, the average correlation with job performance was .44 when applicant data were combined by an explicit rule and .28 when the same kinds of data were combined by expert judgment, a relative improvement the authors put at more than 50% 1. That comparison rests on nine studies, and it covers only the last step, how the evidence gets added up once it is in. It says nothing about which evidence to gather. A plain unit-weighted written scorecard counts as such a rule, and nothing in the result licenses a number standing for a person.

The first property is why format alone settles nothing. The 2022 re-analysis of the selection literature puts structured interviews at .42 and unstructured ones at .19 2. The gap between two interviews is wider than the gap between most pairs of methods. What those studies coded as structure is narrow: every candidate faces the same set of questions and is rated on a common scale. Whether the questions are open or behavioral is not part of that coding.

Interviewer agreement in a traditional unstructured interview is low enough that, even against a perfectly reliable measure of job performance, judgments from it could never account for more than 10% of the variance 3. That is a ceiling on how much such judgments could ever explain, implied by reliability rather than measured directly, and it comes from a review reporting a 1995 meta-analysis. It still sets the price of asking whatever comes to mind.

Why has the fabrication test stopped separating good questions from bad?

Because the property that test rested on was cost, and the cost is gone. Producing a detailed, plausible, well-structured account of a past project used to take real effort, which is the reason open behavioral questions were held to be hard to prepare for. A candidate can now assemble one in minutes, and preparing answers is what candidates have always been told to do.

The nearest measurement of how that reads to a person comes from education rather than hiring. In a blind study run through the real examinations system of a UK university, unedited GPT-4 answers submitted under 33 fake student accounts across five undergraduate psychology modules drew no concern of any kind from markers in 94% of cases, and were graded on average above the real students 4. Coursework is not an interview, the tasks were essays and short answers, the markers had no instruction to look for anything, and the authors describe their own method as using AI in the most detectable way possible. What travels is the direction: a plausible written answer, produced quickly, reads as competent to a reader who is not testing it.

So the criterion the published lists are built on has stopped doing the job it was hired for. The questions themselves are still sound. What changed is that the first answer carries much less information than it used to, which is a different problem with a different fix. Every candidate suddenly giving the same polished STAR answer is the shape this takes in a real loop.

The replacement criterion is survivability. Ask what happens when you push the account one level past what a prepared answer contains: the number they were working against, who disagreed, what they would do differently, what the second version cost. A real instance has texture at that depth because the person was there. A rehearsed one gets general at that depth, because a script has nothing underneath the first answer.

Run the four checks against the bank you already have

Open your existing interview guide and put every question through four checks in order. The bank you already run is the one producing your debriefs, so it is the one worth testing. Expect the failures to cluster in the questions people enjoy asking, because a question that is pleasant to ask is usually one any prepared candidate can answer.

1. Could a well-prepared stranger answer this convincingly? If yes, it is a preparation question. "Tell me about a time you handled conflict" passes for anyone who has read a prep guide. "Walk me through the last disagreement you had with someone whose work you depended on" is harder to answer from nothing, because it points at one occasion in that person's history. 2. Does it ask for an instance or invite a policy statement? "How do you approach code review?" gets a philosophy. "What was the last change you asked someone to make in review, and what happened next?" gets an event. The words *the last*, *the most recent* and *the one you regret* are doing most of the work here. 3. Is there a follow-up only someone who was there could survive? Write it down next to the question. If you cannot think of one, the question has no floor, and the strong answer and the rehearsed answer will arrive looking identical. Which follow-ups expose whether someone understands the answer they just gave is the harder half of this check. 4. Is there a written anchor for strong, adequate and weak? Two or three sentences each, naming what the answer contains. A question with no anchor produces a conversation and a rating nobody can reconstruct a week later.

A question that fails checks one and two is usually rescuable by rewriting it toward a specific occasion. A question that fails three and four is a scoring problem rather than a wording problem, and rewriting the question will not touch it.

What do you keep after cutting the bank in half?

Three or four questions with room to go deep on each, plus one opener you expect people to prepare for. Depth is what the cut buys: the minutes freed by deleting six questions are the minutes that let a single answer be followed to the point where preparation runs out. Coverage is the thing teams overbuy, and depth is the thing they never budget.

Keep the preparable opener on purpose. How a candidate spends preparation time is readable in its own right, and the person who arrives having read your changelog, your docs and two of your customers' reviews has told you something real about how they work. Record it as context in the notes rather than as a rated criterion, because rating it would reward available time as much as judgment.

Put the recovered minutes somewhere specific. Three planned questions with three follow-ups each is the arithmetic on the other side of the cut, and the ladder belongs in the guide next to its question so every candidate has the same rungs available.

One thing to hold on to as the bank shrinks: the questions that survive should be about the work this role actually does. A generic question well-formed enough to pass all four checks still tells you less than a specific one about the thing the person will be doing on their second Tuesday. Cutting also needs somebody who is allowed to cut, so who owns the question bank, and what each surviving row has to carry is where that authority gets written down.

See how it works

Common questions

Is there a list of good interview questions I can just use?

There are hundreds, and none of them can be good or bad on its own. A question is only good relative to the role, the round it sits in and the anchor you will judge the answer against, which is exactly what a published list cannot supply. Borrowing questions is fine as raw material. Keep only the ones a well-prepared stranger could not answer convincingly, that name a specific instance, that carry a follow-up only someone who was there could survive, and that have a written anchor for strong, adequate and weak answers. Expect to rewrite most of them toward a specific occasion in your own work.

How many questions should an hour-long interview have?

Three or four planned questions, with room for two or three follow-ups on each, plus an opener and time for the candidate's own questions. A dozen questions in an hour buys a dozen opening answers, and opening answers are the part candidates prepare. The count matters less than the arithmetic behind it: decide how many minutes a single question deserves before you decide how many questions fit, because the second number falls out of the first.

Do behavioral questions still work?

Yes, but not for the reason they are usually recommended. The old argument was that a candidate cannot easily invent a plausible past experience, and that argument has expired. What still works is that a real instance holds up under follow-up and an assembled one does not, so the value has moved from the question to what you do after the first answer. A behavioral question asked once and rated on impression is now close to worthless. The same question probed three levels down is still the best evidence an interview produces.

What does a rating anchor actually look like?

Two or three sentences per level, describing content rather than impression. For a question about a disagreement over technical direction, a strong answer might name the specific decision, the constraint that made it hard, what the other person's case was, and what changed as a result. An adequate answer names the decision and the outcome but not the other side's case. A weak answer describes an approach to disagreement in general with no occasion attached. Every point on the scale needs its own description, including the middle.

Should I send the questions to candidates in advance?

Yes, and it costs less than it feels like it should. Any question generic enough to be damaged by advance notice was already answerable by a well-prepared stranger, and the questions worth keeping depend on the candidate's own history, which no amount of notice supplies. What advance notice does buy is a fairer read on people who need time to think or who are interviewing in a second language. Send the questions, keep the follow-ups unpublished, and take the evidence from the rungs below the first answer.

References

  1. 1. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis Journal of Applied Psychology (American Psychological Association), 98(6), 1060-1072 (Kuncel, Klieger, Connelly and Ones), 2013. gwern.net Supports the .44 against .28 gap between combining applicant data by an explicit rule and combining it by expert judgment, used here to argue for written anchors under every question.
  2. 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .42 against .19 gap between structured and unstructured interviews, used here to argue that the same question set can be worth twice as much depending on how it is run.
  3. 3. Stubborn Reliance on Intuition and Subjectivity in Employee Selection Industrial and Organizational Psychology, 1(3), 333-342 (Scott Highhouse), 2008. edbatista.com Supports the 10% ceiling on the variance in job performance that unstructured interview judgments could account for, stated as a reliability-implied ceiling rather than a measured validity.
  4. 4. A real-world test of artificial intelligence infiltration of a university examinations system: A “Turing Test” case study PLOS ONE (Scarfe, Watcham, Clarke and Roesch), 19(6): e0305354, 2024. journals.plos.org Supports the claim that a plausible written answer reads as competent to a human reader who is not testing it: 94% of unedited GPT-4 submissions drew no concern of any kind from markers.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.