Policy
Enforce the AI Rule by Design, Not by Watching
You can hold candidates to an AI rule without proctoring anyone: design the assessment so undisclosed help buys nothing. Ask for choices and tradeoffs, not a finished artifact. Build the exercise on material no model has seen, such as your own messy data or a real decision from last quarter. Then spend fifteen minutes on that same work in a live conversation. A candidate who cannot account for their own reasoning fails on the exercise's own terms, and nobody had to be watched.
The takeEvery generation of the monitoring product runs into the same object: a phone lying face-down on the desk. Lockdown browsers, webcam proctoring and now network-level blocking of model endpoints all assume the candidate's only computer is the one you can see. The premise worth attacking is not which tool to buy. It is that the rule needs policing at all, when the same money spent on writing a better exercise removes the reason to police it.
Where Olive fits
Open a role and see what the work shows
The expensive part of building this in house is the answer key: deciding what a good rejection of a wrong suggestion looks like before any candidate hands one in. Olive ships that key with each occupational case, and the report comes back as six findings a human reviewer wrote, each tied to the moment in the session it rests on.
Rank your shortlistHow do you make undisclosed AI use worthless?
Ask for the part a model cannot supply. An assistant will produce a competent deliverable in seconds, so a brief that asks for a deliverable is asking for something it can hand over. A brief that asks which two options were rejected and why, against constraints the candidate had to read out of your material first, is asking for a decision only the candidate can defend.
Three changes do most of the work:
- Ask for the discards. "Give me the approach you tried and abandoned, and what made you abandon it." There is no correct answer to look up, and an assistant asked to invent one produces something the candidate cannot then explain.
- Put the constraint in the middle. A brief where the obvious answer breaks on a detail buried in paragraph four rewards whoever read to the end, and an assistant handed only the summary will walk straight past it.
- Ask for one thing you would do differently with more time. It is the cheapest question in hiring and the hardest to fake, because it requires having done the thing.
The research pointing this way is thin but consistent in direction. A team building an AI literacy assessment for a US Navy robotics programme reported that a scenario task simulating AI use on the job outperformed both the tests they adapted from prior research and the ones they wrote themselves 1. That comparison is the authors' own, with no published margin behind it, so treat it as a design argument and nothing stronger.
Pick material no model has already seen
Use your own data, your own constraints, and a decision your team actually made last quarter. Public case prompts and textbook datasets sit in the training data and in every prep guide, so an assistant answers them from memory. A messy internal extract with three columns that contradict each other cannot be answered from memory by anyone, which is the property you are after.
Sanitizing it takes an afternoon. Change the names, round the numbers, drop anything a competitor would want, and keep the mess: the duplicate rows, the field that means two things, the constraint nobody wrote down. Candidates notice the difference immediately, and the exercise stops looking like a test and starts looking like Tuesday.
Then seed one item the model will get confidently wrong. In a Harvard Business School field experiment with Boston Consulting Group consultants, on a task deliberately placed outside the model's capability, those using GPT-4 were 19 percentage points less likely to reach the right answer than a control group of whom 84.5% got it, and the group given a prompt-engineering overview did worse than the group given none 2. One task, one sample, a 2023 model. What travels is the mechanism: people could not tell which side of the line the task was on. A candidate who notices, checks, and says so has shown you the single most useful thing about how they work.
Run the fifteen-minute conversation after the submission
Book it into the next call and give it a script. Three questions: what did you try first and drop, where did the assistant give you something wrong, and what would you change with another two hours. The same questions for everyone, the same rating scale, and agreement in advance on what an acceptable answer sounds like. That is what makes it evidence instead of a chat.
An interview counts as structured only when all three hold. The US Office of Personnel Management's practical guide defines a structured interview as the same questions in the same order, a common rating scale, and interviewers who agreed beforehand on what a good answer looks like, and states that structured interviews have demonstrated a high degree of reliability, validity and legal defensibility 3. The guide is federal practice guidance for US federal hiring, published in 2008, so it binds no private employer and it predates every AI question by more than a decade. That last part is the point: none of this needed a new method.
The conversation is not an interrogation about tools. Nobody is asked to prove they worked alone. The questions are about the work, the answers are recorded against the scale, and the candidate who did the thinking has an easy fifteen minutes. Which follow-ups actually separate understanding from recall is worked through in what follow-up questions expose whether someone understands their own answer.
What this costs, and what it does not buy
About two hours of writing per exercise and fifteen minutes per candidate, against a proctoring subscription and the review queue that comes with it. What it does not buy is certainty about who typed what. Nothing buys that. It buys a decision you can defend on the work itself, which is the question a rejected candidate will actually ask.
The surveillance version carries a cost that rarely reaches the business case. A two-phase study of asynchronous AI interviewers, built on Reddit threads, 17 applicant interviews and a 180-participant evaluation, found that unmet expectations against the employer's framing damaged applicants' sense of agency and trust, and pushed them toward workarounds 4. It is a 2026 preprint on a research prototype, and it was not peer-reviewed when it was checked, so treat the direction as the usable part. The direction is that a format designed to prevent gaming produced more of it.
Two things are worth keeping on the honest side of the ledger. A designed exercise takes longer to write than a monitored one takes to buy, and it has to be rewritten when it leaks, which it will after enough candidates have seen it. Budget a refresh every two quarters and rotate the seeded error.
One thing the monitored version does catch, and the designed version does not: a different person sitting the assessment. Identity substitution is a real failure and none of this addresses it. If that is the risk being managed, the answer is an identity check at the start of the session, not a tool that watches the whole of it.
The comparison in full, including what proctoring genuinely does catch, is in proctoring software versus an AI-open assessment. And the prior decision, which stages need a rule at all, gets set stage by stage, since a single rule for the whole loop fits no stage well.
Common questions
Does this work for a live coding round, or only for take-homes?
It works better live, because the conversation is already happening. Give the candidate the assistant, give them a task with a buried constraint, and spend the last third asking why they accepted one suggestion and rejected another. The output matters less than the account of it. What does not work live is a puzzle with one known answer, which an assistant solves before the candidate has finished reading it.
Do I still need to state an AI rule if the exercise is designed this way?
Yes, and it gets shorter. One line in the posting saying the exercise is AI-open and the following call discusses the decisions in it. The rule now describes what happens instead of forbidding something, so it needs no enforcement clause. Candidates prepare for the right thing, and nobody arrives expecting a test of whether they can work without tools they use every day.
How do I keep the exercise from leaking to prep sites?
Assume it leaks and design for the leak. A published version of your brief is worth little if the constraint lives in data the candidate receives with it, and if the follow-up conversation is unscripted from the candidate's side. Rotate the seeded error each quarter, keep two variants in circulation, and treat a leaked brief as a signal to change the data rather than the whole exercise.
Is asking someone to explain their submission a fair thing to require?
It is one of the fairest steps available, provided everyone in the req gets it. It is the same questions, the same length, and the same scale for every candidate, and it rewards having done the work, not having presented it well. Offer it in writing as well as live for candidates who ask, since the substance of the answer is what gets read.
What if a candidate refuses to discuss how they used AI?
Separate the two things being refused. Refusing to name tools is reasonable and costs nothing, since the tools are not what is being read. Declining to discuss the reasoning in their own submission is a different matter: the exercise asked for a defensible decision, and an undefended decision is an incomplete submission. Record it that way, on the exercise, with no claim about authorship anywhere in the file.
References
- 1. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arxiv.org Supports the design argument that a realistic scenario task reads applied AI skill better than an abstract knowledge test.
- 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports seeding one item outside model capability, because people could not tell which side of the line a task was on.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the three properties that make the follow-up conversation evidence rather than a chat.
- 4. Expecting Too Much, Getting Too Little: Exploring the Challenges and Design Opportunities of Asynchronous AI Interviewers arxiv.org Supports the claim that a controlling, one-way assessment format costs trust and pushes applicants toward workarounds.
4 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.