Policy
Ask What They Did With It, Not Whether They Used It
Yes, ask candidates how they used AI, but drop the yes-or-no checkbox. Attach two prompts to the work they submit instead: what did you ask the assistant to do, and what did you change or throw away in what it gave back. Two or three sentences each. Print beside them that the answer goes to the reviewer with the work and is never on its own a reason to end the process. Without that commitment the question is a trap.
The takeA disclosure checkbox is a loyalty test wearing a policy's clothes. It cannot be verified, it costs nothing to fail honestly, and the only people it reliably identifies are the ones who told the truth. Replace it with a question about acts and it stops being an integrity check and starts being part of the assessment, which is the only version worth a reviewer's time. If a hiring team will not commit in writing that an honest answer is safe, the field should not exist.
Where Olive fits
Open a role and see what the work shows
Olive puts those same two questions inside the work rather than after it: the session keeps what the candidate asked an assistant and what they turned down, and a human reviewer writes six findings from those moments, each quoting the point in the session it came from.
Rank your shortlistWhy does a yes-or-no question produce nothing?
Because both answers are unfalsifiable, and only one of them is risky to give. A yes costs a dishonest candidate nothing and hands an honest one a label. A no cannot be checked either. And the case that actually arrives, a person doing real work with an assistant helping in places, is the case nothing available to you can adjudicate.
A 2025 study ran three detectors over 72 journal abstracts, and its two halves point opposite ways. On the clean split, fully human text against fully machine text, GPTZero reached 97.22% accuracy with a 0% false-positive rate. On AI-assisted text, human writing polished by a model, its over-detection rate was 25% for non-native English authors against 11% for native authors 1. Small sample, academic abstracts, detector versions from late 2024. The shape still holds: the easy case is easy and the case that arrives in hiring is the mixed one, where the error lands unevenly by the writer's language background.
Unaided reading is not the fallback. Asked to tell GPT-3 output from human writing across stories, news and recipes, untrained evaluators performed at random chance, and three quick training methods lifted them only to about 55% 2. That was crowdworkers on a 2021 model, not an engineer reading a submission in their own field, so it is not a verdict on expert judgment. It does mean that a reviewer who says they can tell is making a claim nobody has measured.
So the binary rests on a verification step nobody has. Once you accept that, the field has a better job available to it: showing you how someone works. The clause-level version of this, and the four other clauses it sits beside, is in the five clauses a candidate AI-use policy needs.
What exactly do you ask, and where does it go?
Two prompts, attached to the submission, with a commitment printed beside them. The prompts ask what the candidate asked the assistant to do and what they changed or threw away in what it gave back. The commitment: this note goes to the reviewer alongside the work, and it is never on its own a reason to end the process.
One wording that carries all of it:
> If you used an AI assistant on this exercise, tell us two things. First, what you asked it to do. Second, what you changed, corrected or threw out in what it gave back. Two or three sentences each is plenty, and there is no right answer here. This note goes to the reviewer with your work. It is never on its own a reason to end the process.
Three details in that text matter more than they look. "If you used" rather than "did you use" removes the confession framing. "There is no right answer here" is the sentence that gets you real answers rather than performances. And the commitment is printed in the same box as the question, not linked from it, because a promise a candidate has to click through to is a promise they will not find.
Put it in the submission form, next to the file upload, and repeat one line of it in the assignment brief. A policy page nobody opens is not where a request belongs.
One thing to avoid: asking candidates to rate their own AI proficiency. In a study of 288 teachers who took both a self-report and a knowledge-based test built on the same framework, correlations between the objective and self-reported factors ran from 0.07 to 0.24 3. Teachers in Taiwan rather than job candidates, and a weak correlation means the two instruments measure different things rather than that people overrate themselves. Either way, a self-rating is not a description of work. Ask about acts. The application-form version of this decision, where the stakes and the wording both change, is in whether the application form should ask about AI at all, and what you may lawfully ask is covered in whether you can legally ask candidates how they use AI.
Read the answer for the discard, not the prompt
The second prompt is the one carrying signal. Anybody can describe what they asked for. What separates a candidate who used the tool from a candidate the tool used is what came back that they refused, and how they knew to refuse it. Look for something the model produced that was wrong, and for how that came to light. Naming the tool is not an answer.
Consultants in a large field experiment, handed one exercise that sat beyond what the model could do, got it right far less often with GPT-4 than without: 84.5% of the control group reached the answer, against 60% and 70% in the two assisted conditions 4. More coaching did not help. The group given a prompt-engineering overview finished further behind than the group given none. One exercise, one sample and a 2023 model, so carry the mechanism and leave the numbers where they were measured. Nobody in that experiment was lazy. They simply could not see where the tool stopped being reliable, and the coaching they were given did not make it visible.
That is what the discard question probes, at the cost of one sentence. Read for three things:
- A specific rejection. A named claim, number or approach that came back and did not survive. Vague virtue ("I checked everything") is not a rejection.
- A check against something outside the conversation. A source, a test run, a colleague, a document. Asking the model whether it was sure is not a check.
- A decision they kept. The part of the work they would not hand over, and why that part.
And read past three things that look like signal: how polished the prose is, which tools they named, and how much assistance they used. None of those separates a strong submission from a weak one. What to look for when a candidate hands over a full chat log is a longer job, worked through in reading a candidate's AI transcript.
Don't turn the disclosure into evidence against the person who gave it
The rule you apply to the first honest answer decides what every later cohort tells you. If the candidate who described their process gets rejected while the candidate who wrote nothing advances, the lesson travels through referral networks and forums quickly, and the field starts collecting silence and boilerplate. Print the commitment, then keep it in the debrief.
Keeping it means three specific things:
- The recorded reason names the work. "Weak on the trade-off question in part two" is a reason. "Disclosed heavy AI use" is not, and it is the sentence you least want to read back later.
- A blank answer is an answer. Fields get skipped, instructions get missed, and some people decline the question deliberately. Grade what was submitted, and if something in the work is unclear, put that question to them directly rather than reading meaning into an empty box.
- The disclosure gets read after the work, not before. Reading it first colours the assessment you were trying to protect. Score the submission, then read the note, then decide what to ask.
That last one costs nothing and prevents the complaint this whole practice invites: that the disclosure became the assessment.
None of this covers the case where a stated rule was actually broken and the evidence holds. That is a different decision with different steps, and it is worked through in a candidate who clearly used AI on the take-home. What it does cover is the ordinary case, which is nearly all of them: a competent person, an assistant, and a submission that will tell you more if you ask a better question of it.
Common questions
Does the same request work for a live interview round?
It works better as a spoken follow-up than as a form. In a live round, ask the same two questions about a piece of work they already submitted: what they asked for, and what they threw away. The advantage is that you can ask a second question, which is where a rehearsed answer separates from a real one. Keep the wording identical across candidates so the round stays comparable.
Should the reviewer see the disclosure before or after reading the work?
After. Reading it first frames everything that follows, and the effect runs both directions: a confident note makes a mediocre submission read better, and a candid one makes a strong submission read worse. Have the reviewer score the work first, then open the note, then decide what to ask in the follow-up. If your form makes that ordering impossible, that is a form problem worth fixing.
What if two candidates submit near-identical disclosures?
Usually it means the question was too generic, not that anything is wrong. Answers converge when the prompt invites a description of a tool rather than a description of a decision. Asking what they changed or discarded produces divergent answers because the discards are specific to the work. If answers still look templated, add one line asking for the single thing they were least sure about.
Does collecting this create records we have to manage?
Yes, and treat it like the rest of the submission. It is candidate-supplied material tied to a selection decision, so it belongs under the same retention rule, the same access limits and the same privacy notice as the assignment itself. Do not build a separate store for disclosures, and do not keep them longer than the work they describe. If your retention period is four years for hiring records, this is inside it.
What about candidates who say they used nothing?
Take it at face value and move on. Some people work that way, some roles attract them, and treating the answer with suspicion recreates the exact problem an honest disclosure request is meant to avoid. If the submission raises questions, ask about the work in the follow-up conversation rather than reopening the disclosure. The answer you cannot verify is not made more verifiable by doubting it harder.
Do we have to ask every candidate, or only the ones we wonder about?
Every candidate at that stage, without exception. A question asked only of the people someone privately doubts is the definition of inconsistent treatment, and the pattern of who got asked is exactly what a complaint will be built from. It is also worse evidence, since you lose the comparison across the cohort. One request, in the submission form, for everyone who reaches it.
References
- 1. The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication pmc.ncbi.nlm.nih.gov Supports the claim that the clean human-versus-machine split is the easy case and AI-assisted text is where accuracy falls apart, unevenly by language background.
- 2. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text aclanthology.org Supports the claim that untrained readers judged machine text at chance and that brief training barely moved them.
- 3. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arxiv.org Supports the claim that a self-rating of AI skill and a demonstrated measure of it come apart, so the request should ask about acts rather than ability.
- 4. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that the failure mode is not spotting a task outside the model's capability, which is what the discard question is written to probe.
4 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.