Interviewing
Prep With AI Is Not the Thing Your Rule Is About
A rule about candidates using AI in interviews should cover the live round only. Preparation with a model is coaching, coaching has never been bannable, and no employer can check it anyway. What makes prepared answers feel worthless is that the question was answerable in advance, which is a question problem rather than a candidate problem. Ask about a decision only this candidate made and had to live with, and the preparation clause stops mattering.
The takeBanning preparation is the rare rule that manages to be unenforceable, unfair and pointless at once. Unenforceable because nothing on the employer's side can see it. Unfair because it lands hardest on the candidates who prepare out of necessity rather than on the ones who already speak the language of the room. Pointless because the complaint underneath it is that every answer sounds the same now, and that is a fact about the questions.
Where Olive fits
Open a role and see what the work shows
A rule about the night before is unverifiable, and a live round can only collect what a candidate says about their work. Olive collects the work itself: a 40-to-60-minute occupational assignment with an assistant present, read afterwards by a person who writes six findings from what happened.
Rank your shortlistWhich of the two is your rule actually about?
Two different activities are wearing one word. Preparation happens before the round: rehearsing answers, researching the company, practising with a model instead of with a friend. Live assistance happens during it, with something feeding the candidate sentences while they speak. Most published rules say AI and cover neither clearly, so candidates apply them to both and interviewers enforce whichever one they had in mind.
The collapse is easy to see once it is named. An employer announces that candidates should not use AI in interviews. The interviewer means the second thing. The candidate, reading a single line with no stage attached, applies it to the first and stops using the study tool they were relying on, or uses it and feels they are breaking a rule. Neither outcome is what anybody intended, and the company has taught its applicants that the process is arbitrary before the first question.
Preparation is also the part with the longest history of being fine. Interview coaching is an industry. Career services rehearse students. Books of common questions have existed for decades, and no employer has ever asked a candidate whether they read one. A model is a cheaper coach with more patience, and nothing about the substitution changes what preparation is.
So the first move is not writing a better rule. It is deciding which stage the rule governs, and saying that out loud. Setting the AI rule by hiring stage is the general version of this problem, of which the interview is one row.
Why can't a preparation ban be enforced?
Nothing on the employer's side can see what happened the night before. There is no artifact, no log and no witness, so enforcement has to run backwards from how the answer sounds, which is the one inference the evidence will not support. Rules that cannot be checked still do damage: honest candidates follow them and everybody else does not.
The closest available evidence comes from text, where employers have tried hardest. A study testing 12 publicly available detection tools plus two commercial systems used in academic settings concluded that the tools are neither accurate nor reliable and that obfuscation makes them significantly worse 1. In a separate study, seven detectors run over 91 human-written TOEFL essays by non-native English speakers produced an average false-positive rate of 61.3%, while the same tools were near-perfect on essays by US eighth-graders 2.
Those two point in opposite directions on their headline error because they measure different document sets. The first found tools leaning toward calling text human-written; the second found more than half of one cohort's genuine writing called machine-written. Neither refutes the other, and the useful reading is that the error rate is a property of the material rather than of the tool. Written text is also the easy case. A spoken answer leaves no artifact at all, so an interviewer inferring preparation from polish has less to work with than a detector does, and the detector is not good enough.
What that leaves is a rule enforced by feel, applied to the candidates who sound most rehearsed. Those are disproportionately people who had to rehearse: career changers, non-native speakers, anyone unfamiliar with the register of the room. Whether a structured interview still separates AI-coached candidates is the fairness question underneath this one.
Fix the question instead of policing the prep
The complaint underneath a preparation ban is that every answer now sounds the same. That is a property of the question. "Tell me about a time you handled conflict" has a knowable good answer and was always answerable in advance; a model only made the advance work cheap. Replace it with a question only this candidate can answer, about a decision they made and had to live with.
Preparability by itself is not the problem, and it is worth being precise about why. The best evidenced selection method in the published literature is the structured interview, top-ranked at .42 in the 2022 re-analysis, ahead of job knowledge tests at .40 and work samples at .33, with unstructured interviews at .19 3. A structured interview is defined by asking every candidate the same questions in the same order, evaluated on a common rating scale, with interviewers who agreed in advance what an acceptable answer contains 4. That format is deliberately predictable. Predictability is where its reliability comes from.
Two things limit that claim. None of that literature involves AI-coached candidates, and the .42 carries an 80% credibility interval running from .18 to .66, so a structured interview in any specific setting can land anywhere in that band. What the pairing establishes is narrower and still useful. Questions being knowable in advance is not what breaks an interview.
What breaks it is a question whose good answer requires no particular life. Three that no model can answer for the candidate, because the material only exists in one person's history:
1. Walk me through a decision you got wrong and had to unwind. Follow up on the unwinding, which is the part nobody rehearses. 2. What did you check before you sent that, and what would you have missed? The second half is the question. 3. What is the thing your last team believed that you disagreed with, and what did you do about it?
If the loop still produces four interchangeable answers after that, the problem has moved somewhere else, and why every candidate suddenly gives the same polished answer works through the rest of it.
What should the written rule say?
Three sentences, scoped to the round. Preparation with any tool is fine and the company does not ask about it. During a live round the candidate is the only participant, so nothing supplies answers in real time. For take-homes and work samples the assistant is allowed and the brief says which one. Then one more sentence about what happens if real-time assistance shows up.
That last sentence is the one most policies skip, and skipping it is what turns an awkward moment into an accusation. A workable version: if it looks like something is supplying answers during a live round, the interviewer says so in the moment and offers a different question, and the loop continues. No conclusion about the candidate is recorded from the interviewer's impression alone. That keeps a wrong guess cheap, which matters because the guess is unverifiable in both directions.
Put the rule in the invitation rather than in a policy document nobody opens, in the candidate's own reading order: what this round is, what is allowed, what is not, what happens next. Candidates who know the rule follow it, and the ones who were going to break it were never going to read the careers page either. Whether to let candidates use AI during the interview at all is the prior decision, and what live assistance actually looks like from the interviewer's side is worth reading before writing the clause.
If the live round is the only place the decision gets made, the rule is carrying more weight than any rule should. A loop that also contains a piece of real work has somewhere to put the question that a conversation cannot settle, which is the split worked through in first round collects the claim, final round watches the work.
Common questions
Should candidates be told they can prepare with AI?
Saying nothing is fine, and saying it explicitly is better. A line reading that how a candidate prepares is their business removes an anxiety that otherwise shows up as stiffness in the first ten minutes, and it costs nothing. What it must not do is turn into a question later in the process, because asking somebody how they prepared after telling them it did not matter is worse than never having said it. Decide once and keep the answer consistent across the loop.
What about a take-home assignment, is that preparation or the round?
It is the round, and it needs its own line. A take-home is work rather than conversation, so the sensible default is that the assistant is allowed and the brief says so, because a take-home done without the tools the job uses measures a job nobody is hiring for. State the time box, state that the tools are permitted, and ask for a short written account of what the candidate delegated and what they checked. That account is the part worth reading.
Can I ask a candidate whether they used AI to prepare?
You can, and there is little to gain from it. The answer is unverifiable, the question signals that a yes is a mark against them, and both effects fall hardest on candidates who prepared hardest. If the underlying worry is whether the person can do the work, that is answerable with work rather than with a disclosure question. Keep any disclosure question that survives scoped to the work product, asked of everyone, and never used as the basis for a decision on its own.
How do I handle a candidate who is obviously reading an answer?
Interrupt gently and change the question, in the moment, without an accusation. Ask something adjacent that depends on the specifics of the answer just given, and see what comes back. If the second answer is fluent and specific, the impression was probably wrong. If it collapses, that is a finding about the interview rather than a conclusion about the person, and the right response is a different round rather than a rejection built on an interviewer's read of a video call.
Does a rule like this need legal review?
A one-line rule about live assistance in an interview is ordinary process language and rarely needs more than the usual review of candidate-facing text. What does warrant counsel is anything that produces a decision from an inference about how an answer sounded, or any tool that analyses a recording of a candidate, because both touch consent and discrimination rules that vary by state and by country. Write the rule so it never rests on an unverifiable inference, and most of that exposure never arises.
References
- 1. Testing of Detection Tools for AI-Generated Text arxiv.org Supports the claim that the tested detection tools are neither accurate nor reliable and get worse under obfuscation, which is why a rule cannot be enforced by inference.
- 2. GPT detectors are biased against non-native English writers pmc.ncbi.nlm.nih.gov Supports the 61.3% average false-positive rate across seven detectors on 91 TOEFL essays, and the point that the error lands on non-native writers.
- 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the ranking used to argue that predictability is not what breaks an interview: structured interviews at .42, job knowledge tests at .40, work samples at .33, unstructured interviews at .19.
- 4. Structured Interviews: A Practical Guide opm.gov Supports the definition of a structured interview as the same questions in the same order, a common rating scale, and an acceptable answer agreed in advance.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.