Interviewing
Proctoring or an AI-Open Assessment: Which Stops Interview Cheating?
For most roles, an AI-open assessment stops more interview cheating than proctoring software does. Proctoring only raises the cost of hiding an assistant, and it loses that race to overlay tools funded to keep winning it. Permit the assistant instead, and put something inside the task the model cannot reach: your failing test, last quarter's real numbers. Concealment then buys nothing. Two jobs stay with proctoring: confirming who is at the keyboard, and running a licensure-style exam where the knowledge has to be shown unaided.
The takeProctoring sells something an authored task cannot: the sense that integrity has been handled without anyone rewriting a question. That is why it keeps selling. A subscription is a line item and an afternoon per question is a person's afternoon, and I'd expect the second one to be the harder approval. I have not seen anyone cost the adjudication queue against the authoring it replaces, though the queue usually lands on recruiters and the authoring on hiring managers, which tells you whose time the decision is spending. An exercise that only holds while somebody watches was never measuring much.
Where Olive fits
Open a role and see what the work shows
If you build the AI-open version yourself, the expensive parts are the case for each occupation and the answer key behind every judgment. Olive ships twelve authored cases per occupation and returns six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session, with no camera anywhere in the product and the same report granted to the candidate.
Rank your shortlistWhich one stops more cheating?
An AI-open assessment, for most roles. Proctoring is a contest against software written specifically to beat it, and that software has venture funding behind it 2. An AI-open task changes what is being tested, so concealment stops paying. The exception is identity: only proctoring tells you who sat down.
The two products answer different questions. Proctoring asks whether a candidate did something forbidden and returns a flag stream your team has to adjudicate. An AI-open assessment asks what the candidate did with a tool they were told they could use, and returns a record of the work. One produces suspicion, the other produces evidence, and only the second is readable to the candidate you turn down.
The arms-race arithmetic favors the generator, and it is not close. Cluely, which began as an undetectable tool called Interview Coder aimed at technical interviews, raised a $15 million Series A led by Andreessen Horowitz in June 2025 2. Whatever your proctoring vendor ships next quarter, someone with a product roadmap and a funded engineering team is shipping the counter. You pay a subscription to stay level; the other side pays salaries to get ahead.
An AI-open design does not win that race. It declines to enter. If the assistant is permitted, hiding it buys nothing, and the interesting question becomes what the candidate did when the model was confidently wrong. Candidates quietly using AI to answer questions in a live round stop being a category of misconduct and become a thing you can watch on purpose.
What does proctoring actually catch, and who does it flag?
It catches the obvious and flags the innocent, unevenly. In a study of automated proctoring software, students with darker skin tones had their faces detected for an average of 78% of an assessment against 92% for lighter skin tones, and were flagged as missing from frame 4.79 times per assessment against 0.83 1. Those flags are not cheating. They are work for your team.
Run that arithmetic on your own funnel. A handful of flags across 40 senior candidates is a conversation. The same rate across 4,000 campus applicants is a review queue nobody staffed, worked through by whoever is free that week, against candidates who differ systematically in how often the software loses their face.
The uneven part is the legal part. Under the Uniform Guidelines, a selection rate for any race, sex or ethnic group below four-fifths of the rate for the highest group is generally regarded by the federal enforcement agencies as evidence of adverse impact 5. If a proctoring flag routes a candidate out of your process (even informally, even for a second look), it is operating inside a selection procedure, and a selection procedure has to be job-related and consistent with business necessity 4. "The software flagged them" is not that.
The intrusion has its own record. In *Ogletree v. Cleveland State University*, a federal court concluded that the student's privacy interest in his home outweighed the university's interest in scanning it, and held the practice of conducting room scans unreasonable under the Fourth Amendment 3. A private employer is not a state actor and is not bound by that holding. It still tells you how a judge weighed the two interests you are weighing.
The flags also land hardest on the people least able to absorb them: a candidate whose condition produces movement, who shares a room, who has one device, who cannot sit still for ninety minutes. Accommodation is a live question for any AI-based assessment and it arrives sooner with a camera in the loop. Before you sign, read what an off-camera glance is actually worth as evidence, because your panel will be handed one.
Why does an AI-open task remove the incentive to cheat?
Because the thing worth concealing is permitted. Cheating in an interview is almost always hidden assistance, and hiding pays only while assistance is banned. Tell candidates the assistant is allowed, put a constraint inside the task that the model has no access to, and the exercise stops rewarding concealment. What is left to grade is judgment, which is what you wanted to buy.
The constraint is the whole design. A generic brief (size this market, design this rate limiter) can be finished from its own wording, so an AI-open version of it measures the model. A brief carrying your failing integration test, last quarter's real retention table, or the client's actual budget cannot. The assistant produces a fluent answer that is wrong in a specific, findable way, and the candidate either catches it or ships it. The rewrite is question by question, and it is more work than buying a subscription.
Score acts, not impressions. What did they frame before generating anything? Which claim did they demand a source for, and did they open it? What did they keep for themselves? What did they refuse, and on what grounds? What did they test against something outside the chat? Each is a yes or a no two interviewers will agree on, and agreement is what survives a job-relatedness challenge 4.
Two honest costs. Authoring is real work: roughly an afternoon per question, plus a rubric, plus a dry run against someone already doing the job. The task does not transfer between fields, though the rubric does. And a live AI-open round carries setup the closed version never had: screen sharing, what the interviewer watches while the model is generating, what happens when a tool stalls. Running a coding or case round with the assistant open is a different exercise, not the old one with a rule removed.
Decide by funnel shape, not by company policy
The tolerance calculation is different at 4,000 candidates than at 40, and one company-wide policy gets one of them wrong. A high-volume funnel cannot absorb an adjudication queue or the uneven flag rate that comes with it. A small senior slate can absorb the authoring cost of a real task, and cannot absorb one wrongly accused finalist.
For a high-volume analytics or campus funnel: an AI-open task, gradable against a fixed key at the first stage, with a short structured follow-up on whoever advances. Proctoring at that volume is a review queue, and the review queue is where uneven flag rates turn into uneven outcomes 1 5.
For a small senior consulting or finance slate: an AI-open case built from a real engagement, forty to sixty minutes, followed by a debrief on the decisions the case forced. Twelve people is not a proctoring problem. It is an authoring problem, and at twelve people you can afford to author.
For a remote-first role where nobody will ever meet the hire in a room: identity is the risk, not assistance. Verify identity once, at a named step, by a method you would describe out loud to a candidate, and keep it away from grading, so a document check never becomes a performance signal. Which AI-use policy you can actually defend is worth settling before the vendor call rather than after it.
One rule across all three: whatever you choose, say it in the invite, in the same words for everyone. A surprise produces two populations (the candidates who believed the rules and the ones who hedged), and you cannot compare them to each other.
What is proctoring still the right tool for?
Identity, and licensure-style exams where a body of knowledge has to be demonstrated unaided. Proctoring is the only one of the two that answers "is this the person we interviewed." If your real risk is a paid proxy sitting the assessment, an AI-open design does nothing about it: permission to use a model is not permission to be someone else.
Keep it narrow if you keep it. A one-time identity check at a named step is a different product from ninety minutes of continuous room monitoring, and it carries a fraction of the intrusion the court weighed in *Ogletree* 3. Buy the check, not the monitoring.
Ask any proctoring vendor three questions before signing. What is the flag rate, broken out by group? What share of flags survive human review? Who adjudicates the rest, and what are they told to do with an ambiguous one? A vendor that cannot answer the first is selling you the disparity measured in that study 1 without telling you its size, and you will own it.
Both have a ceiling, and it is worth saying. Neither instrument stops a determined candidate. Proctoring raises the cost of one method; an AI-open design makes that method pointless and grades what happens instead. Olive is one instrument in the second category and not the only one. Multiple-choice AI literacy tests, code-collaboration graders and unwatched take-homes are all real approaches with different trade-offs.
Common questions
Does proctoring software catch AI use during an interview?
Only the visible kind. It watches the camera, the browser and sometimes the desktop, so it catches a second window or a phone held below the frame. What it cannot watch is anything rendered outside the surfaces it captures, and the products sold into this market are marketed as undetectable 2. Treat a vendor's claim of complete coverage as a statement about today's tools rather than next quarter's.
Can we reject a candidate on a proctoring flag alone?
Don't. A flag is a machine's report of movement, noise or a lost face, and the measured rates show how unevenly those land across groups 1. If a flag routes someone out of the process, it is operating inside a selection procedure, which has to be job-related and consistent with business necessity 4. Review the recording, ask the candidate, write down what you concluded. If the flag is the only evidence, it is not evidence.
Isn't an AI-open assessment just permitting people to cheat?
Only if the task can be finished by the model. Cheating means credit for work you did not do; an AI-open task is built so the model's output is starting material rather than the answer. The candidate still has to catch what it got wrong about your data, your test suite or your budget. Someone who accepts the fluent version and ships it has failed the exercise in a way you can point at afterwards.
What about impersonation (someone else taking the assessment)?
That is the one risk an open design does not touch, and the honest reason to keep an identity step. Verify once, at a named point, using a method you would describe out loud to a candidate. Keep it separate from grading so an identity check never becomes a behavioral signal. Continuous monitoring is a far larger intrusion than a one-time check and adds very little against this specific risk.
Which one costs more to defend if a candidate challenges it?
Proctoring. Defending it means explaining an automated flag, its rate per group and who reviewed it. The four-fifths comparison is the first thing an agency runs 5. Defending an AI-open assessment means producing a rubric, the acts you graded and the artifacts behind them. Work samples are a named type of selection procedure 4; job-relatedness is still yours to prove, and the rubric is how you prove it. One is evidence about the work; the other is a suspicion you now have to justify.
References
- 1. Racial, skin tone, and sex disparities in automated proctoring software ✓ frontiersin.org Facial detection averaged 78% of an assessment for students with darker skin tones against 92% for lighter skin tones; missing-from-frame flags averaged 4.79 per assessment against 0.83.
- 2. Cluely, a startup that helps 'cheat on everything,' raises $15M from a16z ✓ techcrunch.com Cluely began as an undetectable tool called Interview Coder for technical interviews and raised a $15 million Series A led by Andreessen Horowitz in June 2025.
- 3. Ogletree v. Cleveland State University, No. 1:21-cv-00500, Amended Opinion and Order ✓ govinfo.gov The court concluded the student's privacy interest in his home outweighed the university's interests and held the practice of conducting room scans unreasonable under the Fourth Amendment.
- 4. Employment Tests and Selection Procedures ✓ eeoc.gov Selection procedures must be job-related and consistent with business necessity; work samples are a named example of a selection procedure.
- 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 ✓ ecfr.gov A selection rate for any race, sex or ethnic group below four-fifths of the rate for the highest group is generally regarded as evidence of adverse impact.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.