Interviewing

Twelve Interview Questions That Survive an AI-Prepped Candidate

Ask interview questions whose answers point at something checkable in the room. Twelve that hold up: bring an artifact from the last six months and walk through the decisions inside it; name the part you deleted; describe the version the model got wrong and how you knew. None of them is unguessable, and that is fine. AI prep is universal now. What a rehearsed answer cannot do is get more specific under a second and third follow-up.

The takeThe tempting response to universal AI prep is a detection exercise wearing an interview's clothes: questions built to catch AI-written answers. It fails twice: it puts the interviewer in the business of judging prose style, and it spends the round on authorship instead of ability. A hiring decision needs evidence about whether the candidate can do the work. Nothing about who typed the first draft of a story answers that.

Where Olive fits

Open a role and see what the work shows

Every question above asks a candidate to describe work they did somewhere else. Olive puts the work in front of them instead: a 40-to-60-minute occupational assignment done with an AI assistant, returned as six findings, each carrying the timestamped moment it rests on.

Rank your shortlist

What makes a question survive preparation?

Three properties. The answer has to rest on something the candidate can be asked to open, produce or quantify. It has to have a wrong version, so there is a claim to check rather than a preference to agree with. And it has to keep going: a good question has a second and third layer that depend on the answer to the first, which is where borrowed answers run out.

Novelty is the property people reach for, and it is the one that decays fastest. A question that is hard to anticipate stops being hard to anticipate the week somebody posts it. The evidence points somewhere less exciting. Sackett and colleagues, re-analysing the selection literature in 2022, put structured interviews at .42 against .19 for unstructured ones 1, and structured there is a research coding of format rather than a clever question bank: the US Office of Personnel Management defines it as one set of questions asked in one order, rated on a scale everybody shares, with agreement reached beforehand on what an acceptable answer contains 2. Those are corrected correlations with supervisor ratings pooled across many jobs, not accuracy rates, and the credibility band around .42 runs from .18 to .66, so one company's round can land anywhere inside it. The direction is still plain. How the round is run outweighs which questions got picked.

So the twelve below are not secrets. Several sit on public question lists already, and a candidate who prepared for them will give a better first answer than one who did not, which is fine: a better first answer is easier to check than a worse one. What interview questions actually show how a candidate works with AI narrows the same idea to rounds built specifically around AI use.

Use these twelve, grouped by what makes the answer checkable

Three groups of four. Each group hangs on a different anchor: an artifact the candidate can open, a decision that had a loser, and a claim that turned out to be wrong. Pick the anchor first and the wording second. A question with no anchor collects an opinion, and an opinion is the thing preparation supplies best. Each group gets harder the deeper you take it.

Bring something you made.

1. Bring one piece of work from the last six months you would be comfortable showing a stranger, and walk me through three decisions inside it that a different person would have made differently. 2. Which part of that did you write yourself, which part did you hand off, and why did the line fall there? 3. Show me a place where you changed your mind partway through. What changed it? 4. If this had to be half the length, what goes first?

Name the option you did not take.

5. What was the second-best approach, and what would have gone wrong with it? 6. What did you delete? Not trim: delete. 7. Who disagreed with you, and what was their strongest argument? 8. What did this cost that does not show up in the result?

Describe something that turned out to be wrong.

9. Describe a confident answer, from a model, a colleague or a vendor, that turned out to be wrong. How did you find out? 10. What is the last thing you checked because you did not believe it, and what did you check it against? 11. Which claim in your field looks true and is not? How would a new hire find out? 12. Tell me about work you shipped that you later found a mistake in. Who found it?

The third group earns the most room, because that is where the failure mode of AI-assisted work lives. It is not sloppiness. It is a confident, well-formatted answer to a task sitting just outside what the model handles well. In the Boston Consulting Group field experiment, on one task deliberately chosen to sit outside the model's capability, consultants using GPT-4 were 19 percentage points less likely to reach the right answer than the control group: 84.5% correct without the tool, against 60% and 70% in the two AI conditions 3. One task, one sample, a 2023 model, so read it as a shape rather than a rate. The shape is that people could not tell which side of the line the task was on, and the group given a prompt-engineering overview did worse than the group given none.

Questions nine through twelve ask whether somebody has been on the wrong side of that line and noticed. What being good at using AI actually looks like in an interview works through what a strong answer to question nine sounds like.

How far should the follow-ups go?

Three layers, minimum, and the third is where the round gets decided. Layer one is the story. Layer two asks for a number, a date, a name or an artifact. Layer three asks about a specific thing inside layer two's answer. A prepared answer survives layer one and often layer two. It rarely narrows correctly at layer three, because layer three did not exist until the candidate spoke.

A worked chain, on question five. Layer one: what was the second-best approach? The answer names one. Layer two: who preferred it, and what did they say when the decision went the other way? Now there is a person, a date and a room. Layer three: what would you have had to be wrong about for their version to have been right? That last question cannot be prepared, because it is assembled out of the two answers before it, and it is easy for anyone who genuinely weighed the two options.

Budget accordingly. Twelve questions do not fit in a 45-minute round and were never meant to. Pick four, one from each group plus one drawn from the role, and spend ten minutes on each. Four questions taken three layers down beats twelve taken one layer down, on the same clock. What follow-up questions expose whether someone actually understands the answer they just gave sets out the mechanics.

One rule for the panel: a follow-up is not a cross-examination. Ask it the way you asked the first question, say why you are asking, and let a good answer end the chain early. An interviewer who treats depth as pressure gets a worse answer and learns less from it.

Stop scoring the things preparation now supplies

Fluency, structure and polish stopped being information. They were weak signals before and they are noise now, because a candidate can rehearse a clean answer in twenty minutes for free. Scoring them punishes the people who did not know that was available and rewards the ones who did. Move the weight to specificity under follow-up, and to whether the answer contained anything that could turn out to be false.

The related instinct, marking an answer down for sounding machine-written, is worse than unhelpful. Non-expert evaluators asked to tell GPT-3 output from human writing performed at random chance, and three quick training methods, detailed instructions, annotated examples and paired examples, lifted accuracy only to about 55% 4. That was 2021, crowdworkers, short passages in three domains, so it is not a measurement of an experienced manager reading work in their own field. Nothing since has made unaided judgment easier.

Concretely, the scorecard changes in three places. Drop any line rewarding structure or delivery. Add a line for evidence quality, rated on whether the answer produced a checkable particular when asked for one. Add a line for the discarded option, rated on whether a real tradeoff appeared. Then rate every candidate on the same three lines and write one sentence of evidence under each, quoting what the candidate said rather than summarising your impression of them.

Whether the STAR frame survives all this is the narrower version of the same argument, since STAR is the container most of these questions arrive in.

See how it works

Common questions

Can these questions work in a 30-minute screen?

Two of them, taken three layers down. A short screen cannot cover four questions properly, so ask for one piece of recent work and one confident answer that turned out to be wrong, and let the rest wait for the longer round. The failure in a short screen is trying to cover ground: six questions at one layer each produces six rehearsed answers and no information. Two questions with real follow-ups produce something a debrief can argue about.

Should candidates get these questions in advance?

Yes for the behavioral ones, and it costs less than it sounds like. Advance notice does not manufacture an artifact, a real tradeoff, or the ability to keep narrowing under follow-up, which is what the questions are built to surface. Keep back the specific material the candidate will react to, since that is what makes the round live rather than recited.

What if a candidate has no artifact they can share?

Common, and not disqualifying. Client work, classified work and anything under an NDA cannot be shown, and a junior candidate may have nothing but coursework. Ask them to describe a piece of work at a level of detail that breaches nothing, then run the same follow-up chain against the description. Or hand them your own material and ask them to critique it, which puts the same judgment against something you can check directly.

Does asking what someone deleted feel like a trick?

Not if you say why you are asking. One sentence does it: what got cut usually says more about the decision than what stayed. Candidates answer it easily when they did the work and struggle when they are reciting someone else's account of it, which is the information the question exists for. A trick is a question with a hidden right answer. This one has no hidden answer, and giving the reason improves the response.

How do you compare two candidates who both answered well?

On the evidence each answer produced, not on which one sounded better. Write down the particulars each candidate gave under each question, then set the two lists beside each other in the debrief. One usually carries more checkable material: a number with a baseline, a person who disagreed, a mistake they found themselves. If both lists are equally rich, the round did its job and the decision belongs to another step in the loop.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .42 structured against .19 unstructured interview validity estimates and the width of the credibility band around .42.
  2. 2. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the three-property definition of a structured interview: same questions in the same order, a common rating scale, agreement in advance on an acceptable answer.
  3. 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that the failure mode is a confident wrong answer on a task outside model capability: 19 percentage points worse, 84.5% against 60% and 70%.
  4. 4. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL-IJCNLP 2021), 2021. aclanthology.org Supports the claim that untrained readers cannot tell generated text from human writing, and that brief training lifts accuracy only to about 55%.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.