Policy
Set the AI Rule Stage by Stage, and Let the Job Decide Each One
Whether a hiring stage should allow AI comes down to one test: with an assistant open, does the stage still measure the thing it exists to measure? A resume screen fails whichever rule you write, because you cannot read authorship off a document, so run it AI-open and say so. A live round about tradeoffs survives an open assistant. A take-home survives only if somebody reads the process next to the artifact. A final panel needs no rule. Set each stage's rule from the answer, not from what feels enforceable.
The takeThe tiered rule going around, prep allowed, take-home disclosed, live round closed, is a compromise with enforcement dressed up as a judgment about measurement. It bans the assistant in the round that most resembles the job, for roles where the hire will work with one every day. If the stage is worth running at all, it is worth running under the conditions the work happens in. The stage that genuinely cannot survive an open assistant is the resume screen, and nobody bans it there.
Where Olive fits
Open a role and see what the work shows
If you are building the AI-open stage in house, the expensive parts are the answer key and the evidence trail. Olive ships item banks grounded in twelve occupations and returns six findings, each anchored to a timestamped moment in the session rather than to a number.
Rank your shortlistWhat test decides the rule for a stage?
One question, asked stage by stage: with an AI assistant open, does this stage still read what it was built to read? A stage passes when the assistant changes how fast the work gets done and not what the work reveals. It fails when the assistant can produce the whole output, and the output was the only evidence being collected.
Two failure modes fall out of that, and they need different fixes.
The first is a stage whose only evidence is a document. A resume, a cover letter, a written screening answer: the artifact is the entire signal, and nothing you can run against the artifact tells you who produced it. A 2023 study of fourteen text-detection tools, twelve public ones plus Turnitin and PlagiarismCheck, tested in an academic-integrity setting, concluded that the available tools are neither accurate nor reliable 5. Code is no better: an ICSE 2025 study ran five general detectors and one built for code against AI-generated source code and found accuracy mostly below 0.6, which the authors call ineffective 3. That stage was already thin, and the rule you write for it changes nothing about what it measures.
The second is a stage that collects reasoning: a decision under a constraint, a defense of one option over two others, a review of somebody else's work. An assistant open during that stage changes the pace and leaves the reading intact. The evidence is the account the candidate gives of the decision, and that account has to come from them.
Federal selection law treats both as the same kind of thing. The Uniform Guidelines, the 1978 federal regulation at 29 CFR Part 1607, define a selection procedure as any measure or procedure used as a basis for an employment decision, broadly enough to name informal or casual interviews and unscored application forms 2. Swapping a scored exercise for a chat does not move a stage outside that definition, so "it was only a conversation" buys less cover than teams assume. The Guidelines attach a validation burden only where adverse impact shows up, which is the part to walk through with counsel.
Run the test on each stage in a normal loop
A four-stage loop takes about ten minutes to sort out. Application: fails the test, so no rule helps. Take-home: passes only if somebody reads the process alongside the artifact. Live exercise: passes with the assistant open, provided the interviewer asks about choices. Final panel: no AI rule applies, because nothing there is produced on the spot.
Worked through, for an analyst req:
- Application and resume screen. AI-open by default, and say so. What the stage reads is a set of claims, and no rule you write will make it read authorship.
- Take-home, 90 minutes. AI-open, with the brief asking for the two options rejected and the reason. The artifact alone stops being the deliverable, so the assistant stops being a substitute for the candidate. The narrower version of this call is worked through in whether to allow AI on the take-home.
- Live exercise, 60 minutes. AI-open, with the interviewer spending the last third on why. Running a coding or case interview with an assistant open is a different craft from running one without, and the questions have to be written for it.
- Panel and references. No rule at all. Nothing here is drafted in the room, and a rule that governs nothing is noise in the posting.
Two of those four are AI-open because the assistant improves the stage, not because it was tolerated there. That is the part the tiered template gets backwards.
Why the live round is the stage most often ruled wrong
Because the ban there is chosen for enforceability and then explained as measurement. A live round is the one stage with a person in the room, so the prohibition feels checkable, and the stage that feels checkable gets the strictest rule. That is backwards for any role whose working day includes an assistant: closing the tool makes the round less like the job, not more.
Structure is what moves an interview's predictive value. The 2022 re-analysis of selection validity puts structured interviews at .42 and unstructured ones at .19, from a sample-size-weighted combination of two earlier meta-analyses 1. Structured there is a research coding: the same questions in the same order, a common rating scale, and agreement in advance on what an acceptable answer looks like. That paper measured no AI use at all, and it did not have to: none of what it calls structure turns on whether a browser tab is open.
A live work-sample round sits in the same band on the same table. Work samples estimate at .33 in that re-analysis, down from the .54 still quoted everywhere, and the authors read the whole revised pattern as coherent 1. Almost every study behind that number tested people already doing the job, which is worth knowing before quoting the figure at a candidate. Choosing between a strong exercise and a strong structured interview is a staffing question.
There is also something a closed round cannot show you. In the BCG field experiment, on one task deliberately placed outside the model's capability, consultants using GPT-4 were 19 points less likely to reach the right answer than a control group of whom 84.5% got it right 4. The candidate who notices that moment and pushes back is the one worth hiring, and a round with the assistant closed never gets to watch it happen. Turning that into a scored round is the whole subject of redesigning the interview so AI assistance becomes signal.
Write the rule down where candidates read it
In the posting and again in the stage instructions, one line per stage, naming the posture and what follows from it. A rule that first appears in the assignment email reaches the candidate after they have already committed the evening. A rule that never appears cannot be applied without inspecting the artifact, and inspection is the part that does not work.
Three sentences cover most loops:
- "The assignment is AI-open. The call afterwards spends twenty minutes on the decisions in it."
- "The live exercise is AI-open. Bring whatever you normally work with."
- "Nothing in this process is checked for AI authorship."
That third line is the one teams flinch at, and it is the honest one. Saying it costs nothing you had, because the checking was never available, and it removes the quiet incentive to lie that an unenforceable prohibition creates. Teams that would rather keep a prohibition should read what one actually selects for in whether a blanket AI ban is worth setting.
Decide the postures before the wording. Per stage, name what the stage reads and whether an assistant changes that reading; the sentence then writes itself. If a stage cannot survive the question, the fix is the stage rather than the rule, and making the rule hold without watching anyone is where that work goes.
Common questions
Should an AI-open stage still ask candidates to disclose what they used?
Two lines, in the assignment brief rather than the application form: which tools did what, and what changed after. That is enough to open the follow-up conversation. A prompt log is not worth asking for, because nobody reads it and it converts a discussion about judgment into a documentation exercise. Keep the ask small enough that a candidate answers it in a minute.
What if one stage genuinely has to be closed?
Some are, and the reason is usually external: a licensing exam, a client confidentiality constraint, an assessment your regulator specifies. Write it as a fact about the exercise, name the reason, and keep it to that stage. A closed stage with a stated reason reads as a constraint; a closed loop with no reason reads as distrust.
Do junior and senior candidates need different stage rules?
The postures stay the same, the questions change. A senior candidate is asked what they refused to delegate and why; a junior candidate is asked to show the check they ran on something the assistant produced. Different rules by level would mean two processes inside one req, which is the condition that makes a decision hard to defend later.
How do I compare two candidates who used wildly different amounts of AI?
Do not compare the amount. It is not a quantity that predicts anything, and grading it rewards whoever guessed your preference. Compare what each of them did with a wrong or thin output: noticed it, sourced it, dropped it, shipped it anyway. That difference sorts candidates in a way you can write down and explain, which the tool count never does.
Do the stage rules have to be identical across every role?
Across every candidate in one req, yes. Across roles, no. A req for a job where nobody may open an assistant on day one earns a different set of postures from a req for a job that runs on one, and that difference is defensible because it tracks the work. What is not defensible is two candidates in the same req meeting different rules.
References
- 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the structured versus unstructured interview gap, and the revised work sample estimate used to argue that the exercise and the interview sit in the same band.
- 2. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) govinfo.gov Supports the claim that an informal conversation and a scored exercise are the same category of selection procedure, so swapping one for the other moves no stage outside the rules.
- 3. An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We? arxiv.org Supports the claim that a stage whose only evidence is a submitted artifact cannot be checked for authorship, including code.
- 4. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that the behavior worth watching in a live round is whether the candidate catches a confident wrong answer.
- 5. Testing of Detection Tools for AI-Generated Text arxiv.org Supports the claim that a stage whose only evidence is a written document cannot be checked for AI authorship, the prose counterpart to [3] on code.
5 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.