Assessment design
How Do You Run a Coding or Case Interview With AI Open?
With the candidate's AI assistant open, run the coding or case interview as an audit, not a build. Hand them working AI output with a real flaw and score four columns: what got tested against something outside the chat, what got refused and on what grounds, what they kept for themselves, and what changed in the final call. How much AI they used and how fast they worked stay off the sheet. Make it two flaws of different kinds, where your field gets burned, calibrated on two practitioners.
The takeThe column that decides these rounds is refusal, and it is the one nobody gets promoted for. Every incentive inside a company rewards shipping the thing; almost none reward holding back something that looked fine. So the audit tests a habit those same incentives seem to train out of people, and on a resume it tends to read as slowness. The one finding here I would stake a hire on is the smallest: the developers who trusted the assistant less and reworked their prompts wrote fewer vulnerabilities. Distrust is not a temperament problem you coach away. It is the skill.
Where Olive fits
Open a role and see what the work shows
Building the audit yourself means authoring the flaw, the answer key and the rubric again for every role you hire into. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session, and the candidate is granted the same document.
Rank your shortlistWhy does the produce-the-answer format stop working?
Because the answer is no longer scarce. A candidate with an assistant open produces a working solution to a standard exercise in the time it takes to paste the prompt, so the exercise ranks typing speed and model access. Worse, the polish reads as competence: candidates who prepared with generative AI received higher overall interview performance ratings 1, and a panel cannot tell polish from composition.
The failure is not that the candidate cheats. It is that the artifact stops carrying information about the person, and the interviewer's own reading gets worse at the same time. In a controlled study of AI-assisted decisions, adding an explanation raised the chance a person accepted the model's recommendation whether or not it was correct, and did not improve team accuracy beyond what the unexplained model already provided 4. A fluent walkthrough of a wrong answer is persuasive to the panel too.
Meanwhile the job moved. Developers now name "AI solutions that are almost right, but not quite" as their single biggest frustration with these tools, and 45% say debugging AI-generated code is more time-consuming 2. A survey of 319 knowledge workers describing 936 real tasks found the effort that remains shifts toward verifying information, integrating responses and stewarding the task, and that higher confidence in the tool predicted less critical thinking, not more 3.
So the round has to test the part that got harder rather than the part that got cheap. That is a different exercise, not a stricter version of the old one. Locking the assistant out and proctoring the room buys a clean sample of a task nobody performs that way any more.
Turn the interview into an audit: the format, step by step
Give the candidate a deliverable that already exists and is wrong in one load-bearing way, then ask for a decision about it rather than a rewrite. Forty-five minutes: ten to read, twenty-five to work with the assistant however they like, ten to say what they would ship and what they would not. The AI stays open the whole time, and so does the source material.
Four moves make it work.
- Give them the output, not the prompt. Paste in a response a competent assistant would plausibly produce (code, a memo, a four-option recommendation), and do not caricature it. If the flaw is visible in the first paragraph, you have written a proofreading test.
- Name the decision. "Would you merge this?" "Which of these four goes to the client, and what gets cut?" A decision forces a position. "Review this" produces a list of observations that two interviewers will score differently.
- Leave the assistant unrestricted. They can ask it anything, including to defend its own output. What they ask, and whether they believe the reply, is the data.
- Ask for the residue. The last ten minutes, in writing: what was checked, against what, what it showed, and what would not ship. Written beats narrated, because a second reader gets the same thing you did.
Two design constraints keep this from drifting into a personality read. The Uniform Guidelines support a selection procedure on content validity to the extent it is a representative sample of the work behavior or work product of the job, and say plainly that content validity is not an appropriate basis for a procedure claiming to measure a construct 7. "Good judgment" is a construct. "Wrote the failing case before touching the fix" is a work behavior. Score the second and the round survives a question from your legal team.
What do you actually observe, per discipline?
The observable is whatever that field re-opens before signing. For engineering it is what gets tested and what gets refused: a candidate who writes the case the model did not think of, and who declines to ship the clever refactor it volunteered. For consulting it is which of the model's four options gets killed, and on what evidence. Same rubric, different task, every time.
Software engineering. The assistant returns a function that passes the tests it also wrote. Observables: does a new case get written, does the change get run rather than read, and what gets refused. Refusal is the underrated one. Participants with a code assistant wrote significantly less secure code than the control group and were more likely to believe it was secure, while the participants who trusted the assistant less and reworked their prompts produced fewer vulnerabilities 5.
Management consulting and case rounds. The model produces four entry options, all plausible, one resting on a figure that does not survive being opened. Observables: which option dies first, what evidence killed it, and whether a criterion was named before the options appeared. A candidate who sorts four options without saying what would make one wrong has been sorted by the model.
Legal, finance and anything carrying a citation. The cite resolves and the paragraph it points at does not say what the memo claims. Purpose-built legal research systems still returned incorrect information more than 17% of the time in Stanford's benchmark, and one exceeded 34% 6. Opening the source is the job, not diligence theater.
Write the flaw where that field's practitioners actually get burned. If you cannot name yours, ask the two strongest people on the team what they re-open before they sign anything, and put the mistake there. This is also the honest argument for keeping the round live rather than sending it home: an audit compresses into a 45-minute slot far better than a build does, because the evidence is the checking, not the artifact.
How do you score it without arguing afterwards?
Four columns, written before the first candidate and filled from the record rather than from impressions: what was tested against something outside the chat, what was refused and on what grounds, what the candidate kept for themselves, and what changed in the final answer. Every column is a yes or a no with a quotation attached. Two interviewers should fill them in separately and agree.
"Tested" means an act you can point at: a test run, a page opened, a figure recomputed. "I double-checked that" with nothing opened is a no, and it is the most common thing a nervous candidate says. Grade selection as well as volume: checking everything is unavailable in real work, so which claim they picked matters more than how many. Someone who verified three peripheral facts and shipped the load-bearing one has done worse than someone who checked only the load-bearing claim. A worked set of rows is an hour of writing that outlives every task you swap in behind it: what a rubric for an AI-assisted answer actually says.
Do not score the candidate's account of their own session. Sixteen experienced developers working on real issues in repositories they knew well took 19% longer with AI tools allowed, and afterwards still believed the tools had sped them up by 20% 8. If a person inside the work cannot read their own speed, their summary of their own rigor is not evidence either. The record is.
Two things stay off the sheet. How much AI they used, because a candidate who judged the model was the wrong instrument for a step and did that step by hand has demonstrated the thing being tested. And raw speed, which the assistant now supplies to everyone: the trade between speed and judgment is the whole reason the round changed shape. Brief the panel on both before the first candidate, or one interviewer will quietly keep scoring the old exercise.
What breaks an AI-open interview?
Four things. The flaw is too obvious, so everyone finds it and the round separates nobody. The flaw is undiscoverable inside the time limit, which makes it a trick rather than a test. The task leaks after the twentieth candidate. And the panel never agreed what an audit looks like, so one interviewer scores the record and another scores confidence.
Calibrate before it decides anything. Hand the task to two people already doing the job: one should find the flaw inside ten minutes, and neither should call it unfair. If both miss it, it is undiscoverable. If both find it in thirty seconds, it is decoration. Then write a second flaw of a different kind into the same task (one arithmetic, one attribution), because a single planted flaw is a single observation, and a competent person can miss one thing on a bad Tuesday while a lucky one stumbles straight into it.
Assume it leaks. A live round leaks more slowly than a take-home and still leaks, so budget a second authored task rather than a rotation nobody will write. Publish the format in the invite, too: that any assistant is allowed, that the material may contain mistakes, and that naming what could not be settled counts toward the answer. Candidates who assume the packet is clean and candidates who assume it is a trap are not a comparable population.
The honest limit is sample size. One 45-minute audit is one observation of one flaw in one task you knew well enough to write. It says nothing about how the person frames a problem nobody has framed for them, and someone can catch your planted arithmetic error and still hand every consequential decision to the model on the job. What an AI assessment should measure covers more ground than a single round can carry, and the gap is worth naming out loud before the round becomes a gate.
Common questions
Should the candidate use their own AI account or one you supply?
Theirs, where your policy allows it. The round is about how this person works with the tool they actually use, and an unfamiliar model in an unfamiliar interface adds a confound you cannot separate from judgment afterwards. Say in the invite that any assistant is allowed and that a copy of the exchange is part of the deliverable. If candidate accounts cannot touch your material, supply one and give every candidate the same one. Comparability matters more here than realism does.
How is this different from planting a bug and asking them to find it?
A bug hunt ends when the bug is found. An audit ends in a decision, and the decision is the thing you score. A candidate can find the flaw and still merge, or miss it and still refuse to ship for a reason you had not considered. Ask "would you ship this, and what changes first" rather than "what is wrong here." The second has a right answer and a scavenger-hunt shape. The first produces a position that compares cleanly across candidates.
What if the candidate just asks the assistant to check its own work?
That is a legitimate move, and what happens next is the signal. Watch whether they accept the model's verdict on the model's own output. Asking an assistant to audit itself and taking the answer is the failure this round exists to surface; asking it, then going to the source anyway, is the pass. Write down which one you saw, verbatim. In the debrief, ask what would have changed their mind if the assistant had said the opposite.
What if the candidate barely uses the assistant?
Not a failure, and it should not be scored as one. Someone who judged the model was the wrong instrument for a step and did that step by hand has demonstrated the exact behavior the round tests. What matters is whether the choice was deliberate and whether they can say what they would have delegated and why. Score the boundary, not the volume. The version worth a follow-up is a candidate who avoids the tool and cannot say what they were avoiding.
Does the round have to be live?
No, but live is what buys you the record. Unwatched, you see the deliverable and not the checking, which is exactly where the score lives. If it has to run async, require the residue in writing (what was checked, against what, what it showed), cap it near an hour, and state the cap in the invite. Live also holds to 45 minutes more reliably, because an audit compresses into a slot in a way a build never does.
References
- 1. When Candidates Use Generative AI for the Interview ✓ sloanreview.mit.edu Candidates who prepared with generative AI received higher overall interview performance ratings.
- 2. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co "AI solutions that are almost right, but not quite" is the top-reported frustration at 66%, and 45% say debugging AI-generated code is more time-consuming.
- 3. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org 319 knowledge workers describing 936 real tasks: effort shifts toward information verification, response integration and task stewardship, and higher confidence in the tool is associated with less critical thinking.
- 4. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance ✓ arxiv.org Explanations increased acceptance of the AI's recommendation regardless of correctness, and did not increase the complementary gains beyond AI augmentation without explanations.
- 5. Do Users Write More Insecure Code with AI Assistants? ✓ arxiv.org Participants with an AI code assistant wrote significantly less secure code and were more likely to believe it secure; those who trusted the assistant less and reworked their prompts produced fewer vulnerabilities.
- 6. AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries ✓ hai.stanford.edu Purpose-built legal research systems still return incorrect information: Lexis+ AI and Ask Practical Law AI more than 17% of the time, Westlaw's AI-Assisted Research more than 34%.
- 7. 29 CFR 1607.14 — Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) ✓ ecfr.gov Content validity holds to the extent the selection procedure is a representative sample of the work behavior or work product of the job, and is not an appropriate basis for a procedure measuring a construct.
- 8. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues took 19% longer with AI tools allowed, and afterwards still believed the tools had sped them up by 20%.
8 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.