Policy
Should You Let Candidates Use AI During the Interview, or Ban It?
Neither a blanket allowance nor a blanket ban on candidates using AI in interviews is right. Permit exactly what the job permits, decide it per role, and put the rule in the invitation so every candidate for that role gets the same one. Close a round only where the work itself has to be done unaided and signed by one person. Sensitive material is not that reason: write a synthetic case and keep the assistant open. Most jobs have one unaided step, so two rounds beat one blanket rule.
The takeBoth rules are defensible on paper, and only one of them ever gets asked to explain itself. A ban is a claim about the job, so when a rejected candidate asks which working condition it came from, the answer is either a line in your own policy or a preference nobody wrote down. A ban that cannot name that line tends to get dropped quietly rather than defended. The rule that survives is the one you would read aloud to the person it screened out.
Where Olive fits
Open a role and see what the work shows
Whichever rule you pick, what defends it later is the record of what each candidate actually did under it. Olive returns six separately-evidenced findings written by a person, each anchored to a timestamped moment in the session, and every released report exports with its rubric, scorer and bank versions attached.
Rank your shortlistWhat is the rule, and what backs it?
Permit in the interview exactly what the job permits. That is the only version of this decision with anything behind it: under the Uniform Guidelines, a selection procedure is supported by content validity to the extent that it is a representative sample of the content of the job, and its manner and setting should closely approximate the work situation 1. Your AI rule is part of that setting.
Whether AI in interviews is acceptable in general is not the question. It is whether the round matches the conditions of the work being bought. A ban and an allowance are both defensible, and each is indefensible in the wrong role. The two failures are symmetrical:
- Ban it for a role that runs on it, and the round measures unaided recall, which is a condition the hire will meet on no ordinary day. You select for preparation the job never needed.
- Allow it for a role whose real work excludes it, and the round measures a workflow the hire will never be permitted to use. You select on evidence that does not transfer.
The standard the decision gets measured against is fit, not preference. A selection procedure is expected to be job-related and consistent with business necessity, and where one screens out a protected group the employer should determine whether an equally effective alternative with less adverse impact exists and adopt it 2. "It feels like cheating" is not a fit argument. "This role has no approved assistant and every deliverable is signed by the person who wrote it" is one.
The rule does one more thing worth having: it makes the round readable. How a candidate used an assistant only means something if they knew they were allowed to, and unaided work only means something if they knew they were not.
When does banning AI test the wrong thing?
When the job runs on one. In the 2025 Stack Overflow Developer Survey, 84% of respondents said they use or plan to use AI tools in their development process, up from 76% the year before 3. Close the assistant in a round for a role like that and the round measures unaided recall under observation.
What gets lost is specific, and it is the part you wanted. With the tool closed you cannot see how a candidate frames a problem before generating, whether they demand a source for the claim that actually matters, what they keep rather than hand over, or what they refuse. Those are the behaviors that separate two people whose finished output reads identically. Close the tool and both look the same on paper.
What you get instead is a proxy: recall, typing under pressure, and how recently someone rehearsed. Those track interview practice more than job performance, and practice is unevenly distributed, which is a familiar way for a round to move adverse impact without moving accuracy.
The distrust figure in that same survey is why the skill is worth testing at all. Among developers, 46% said they distrust the accuracy of AI output against 33% who trust it 3. The job is not operating the tool. The job is knowing when the tool is wrong, and that is visible only while the tool is open. Running a coding or case interview with the assistant open is a worked format, and scoring the answer it produces is a written-key problem rather than a judgment call.
Allowed is not the same as unscored. An open round still needs a rubric, a key written before the first candidate, and a task an assistant cannot finish alone.
When does allowing AI test the wrong thing?
When the real work excludes it. Two constraints do this honestly. The first is material, meaning anything that cannot go into an outside model: client files, patient records, unreleased numbers, code under NDA. The second is authorship: a deliverable one named person must produce, sign and defend, where the review rule requires their own work rather than an edited draft.
Sort those two, because they take different fixes and most teams conflate them.
- A material constraint is not a skill constraint. If the only reason your staff cannot use an assistant on this work is what the data is, the assessment does not inherit the ban. It inherits the requirement to use fabricated material. Write a synthetic case with the same shape, keep the assistant open, and the real workflow gets tested with nothing sensitive anywhere near a model.
- An authorship constraint is a skill constraint. If the person has to read the clause, reconcile the figure or make the call with no model in the loop, a round with the assistant open tests something the job will not let them do. Close it for that round, say why in the invitation, and keep it narrow.
Narrow is the operative word. Almost no role is unaided end to end. The usual truth is that one step is and everything around it is not, so a loop that bans AI everywhere because a single step requires unaided judgment has thrown away the rest of the picture. Two rounds answer two questions; one blanket rule answers neither.
There is a third case that looks like a constraint and is not: the role where no assistant is approved yet, the work is plainly information work, and approval is in motion. Testing a candidate as though the restriction were permanent buys someone good at the version of the job about to be retired. Hire for verification rather than production when the tooling is still moving.
How do you find out what the job actually permits?
Ask two people who do the work and read your own policy. AI applicability concentrates in information work (the creation, processing and communication of information) and varies by occupation 4. Two teams hiring the same title can still work differently, so the manager's summary is a starting point, not the answer. Four facts settle it, and each already has an owner who knows.
- Which assistants are approved, on which accounts, with what logging. Security or IT owns this and it is written down somewhere.
- What material is out of bounds near those tools: client identifiers, patient data, unreleased financials, code under NDA.
- What the person must produce unaided, and whose name goes on it. Legal, compliance or the team lead owns this one.
- How much of an ordinary week is AI-assisted. Ask two individual contributors to describe last Tuesday rather than asking anyone to estimate a percentage.
Write the four answers into one paragraph at the top of the interview plan, before a single question gets drafted. That paragraph is the rule, and it is also the evidence that the rule came from the job rather than from taste. Working out what AI actually does in a role is the same exercise the job post should have run through, so start from that document if it exists.
If the four answers contradict each other, that is the finding. A role whose manager says the team lives in an assistant, and whose security policy names no approved tool, has a problem no interview format fixes, and a candidate should not be the person who discovers it.
Write the rule so it survives a challenge
State it in the invitation, apply one rule to every candidate for that role, and keep a record of how each answer was judged. A procedure that screens people out has to be job-related and consistent with business necessity, and where one screens out a protected group the employer is expected to adopt an equally effective alternative with less adverse impact if one exists 2. An unstated rule fails on the record alone: nobody can say afterward what was applied.
Five lines of bookkeeping cover the rest:
- One sentence per round, in the invitation. Name what is allowed and how it will be read. Candidates guess otherwise, and the guessing tracks coaching rather than ability.
- One rule per role, not per interviewer. A panel where two people allow the assistant and one does not has run three assessments and can compare none of them.
- A key written before the first candidate. Decide what a good answer contains, then apply it. A structured interview asks every candidate the same questions and rates the answers against the same criteria 5.
- A note of why, not a verdict. Record what the candidate did and what it demonstrated. That survives a challenge; a hunch does not.
- An accommodation path. A ban that also removes an assistive tool someone depends on is an accommodation question and gets handled as one, separately from the policy.
Two failures to steer clear of. Do not enforce a ban by watching for tells. A pause, eyes off camera or a suddenly formal register all have innocent causes more common than the guilty one, and acting on assistance you cannot see makes the interviewer the instrument. And do not run an unstated ban and then act on suspicion, which is a selection procedure with no written content at all.
What no interview does, under either rule, is show you the work. Even with the assistant open, forty-five minutes in front of an audience captures a candidate describing how they would check a confident claim far more often than it captures them checking one. A work sample with the assistant open is where that gap closes, and the record it leaves is what the rule gets judged on later. See how Olive measures this.
Common questions
Is it cheating if a candidate uses AI and you never said they couldn't?
No. A rule nobody stated is not a rule, and acting on it afterward penalizes candidates for guessing wrong rather than for working badly. If the round mattered, run it again with the rule stated. If it did not, score what you saw. The fix is the invitation: one sentence per round naming what is allowed. Stating it also makes the round readable, because how someone used an assistant only means something if they knew they were permitted to.
Can you ban AI in one round and allow it in another?
Yes, and for most roles that is the right shape. One step of a job usually needs unaided judgment while everything around it does not, so a loop with one closed round and one AI-open work sample matches the work better than a single blanket rule. Say which is which in the invitation, keep the closed round short and narrow, and check that it tests the specific unaided act the job requires rather than general recall.
How do you stop candidates using AI in a round where it is banned?
Change the question rather than watch the candidate. A ban holds only where assistance does not help: ask about decisions the candidate personally made, hand back their own submitted work with one constraint changed, or run the round in person. Detection is not a real option. The tells have innocent causes, and acting on them turns nerves and second-language processing into a rejection reason. A round that only works when nobody gets help was not measuring much to begin with.
Does allowing AI make the interview too easy?
It raises the floor on output and lowers nothing that matters. Every draft gets better, which is exactly why the score cannot live in the draft. Judge what the candidate did around the assistant: what they framed before generating, which claim they demanded a source for, what they kept, what they threw out, and what they tested against something outside the conversation. Written that way, an open round separates candidates more sharply than a closed one, because the output stops being the differentiator.
What should the invitation actually say?
One sentence per round, plainly. For an open round: this exercise is done with an AI assistant of your choice, and the review looks at how you directed and checked it. For a closed round: this conversation is unaided, because the role requires this step without a model. Then add that the same rule applies to every candidate for the role, and give a contact for anyone who relies on a tool the rule would remove.
Is banning AI legally riskier than allowing it?
Neither is risky on its own; the exposure sits in inconsistency and in a rule with no connection to the job. A procedure that screens people out has to be job-related and consistent with business necessity, so a ban that cannot be traced to a working condition is the weaker position to defend. Applying it unevenly across candidates, or enforcing it on suspicion, is worse than either policy. And a ban that removes an assistive tool someone depends on is an accommodation matter, handled separately.
References
- 1. 29 CFR 1607.14 — Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) ecfr.gov Content validity holds to the extent a selection procedure is a representative sample of the content of the job, and its manner and setting should closely approximate the work situation.
- 2. Employment Tests and Selection Procedures eeoc.gov A selection procedure must be job-related and consistent with business necessity, and an employer whose procedure screens out a protected group should adopt an equally effective alternative with less adverse impact.
- 3. 2025 Stack Overflow Developer Survey: AI survey.stackoverflow.co 84% of respondents use or plan to use AI tools in their development process, up from 76%; 46% distrust the accuracy of AI output against 33% who trust it.
- 4. Working with AI: Measuring the Applicability of Generative AI to Occupations arxiv.org Across 200k anonymized Copilot conversations, the most common and successful AI-assisted activities are information work, and applicability varies by occupation.
- 5. Assessment and Selection: Structured Interviews opm.gov A structured interview asks every candidate the same questions and rates the answers against the same criteria.
5 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.