Assessment design
Should You Disqualify a Candidate for Using AI on a Take-Home?
Don't disqualify a candidate for using AI on a take-home unless the brief banned it in writing before they started and the same line holds for everyone who got that brief. The one exception: a candidate who claims the work as their own when asked directly. A rule invented after reading one submission isn't a rule; it's a mood. What's worth your time is whether they checked what the model gave them. Reread the work for what they confirmed, against what, and what shipped unexamined.
The takeWhat you're hiring for now is editorial judgment: the instinct to distrust a confident sentence until a source confirms it. People with that instinct turn AI into a multiplier; people without it turn it into a liability with good grammar. And detection-tuned hiring probably screens hardest against the first group, because careful, verified output is exactly what reads as assisted. So where the job permits the tool, stop grading for its absence. Grade for the presence of doubt.
Where Olive fits
Open a role and see what the work shows
If you rewrite the take-home so the checking is the thing being graded, the expensive parts are the authored case and the evidence behind each judgment, per occupation. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification), each anchored to a moment in the session, with the candidate granted the same report.
Rank your shortlistWhen does disqualifying actually make sense?
Three cases: the brief said not to before they started, the role genuinely forbids the tool, or they claimed work as their own when asked about it directly. Everything else is a standard you invented after reading the submission and applied to one person, and it will not survive the first candidate who asks why. Write the rule for the next round instead.
The second case is real but narrow. Classified or export-controlled work, a client contract that bars third-party processing, a data policy no approved model satisfies yet. Those are conditions of the job, and they belong in the brief in one line. Outside them, banning the tool at the assessment while the team uses it daily tests something the job never asks for.
Your own leadership is probably on the other side of this. In Microsoft and LinkedIn's 2024 Work Trend Index, a survey of 31,000 knowledge workers across 31 markets, 66% of leaders said they would not hire someone without AI skills and 71% said they would rather hire a less experienced candidate who has them than a more experienced one who does not 1. A disqualification for tool use, in that room, is a decision the hiring manager reverses on second reading.
Whatever you decide, the rule has to exist before the submission does and reach everyone who got the same brief. A standard enforced on the candidate whose writing struck you as suspicious, and not on the one whose style reads as normal to you, is a different test given to different people. It is also the version that becomes expensive when they ask for the reason behind the rejection and the honest answer is that you guessed.
How sure are you that they used AI?
Less sure than the submission feels. Unless they told you, or the file carries a receipt (a pasted chat, a stray placeholder, a leftover instruction), you are reading style and calling it evidence. Clean structure, even paragraph lengths and confident hedging are also what a careful writer produces on a third draft, and what an entire professional register produces every day.
Detectors do not close that gap. Seven of them, run over 91 TOEFL essays written by non-native English speakers, produced a 61.3% average false-positive rate: 97.8% of those human-written essays were flagged as AI-generated by at least one detector and 19.8% by all seven, while the same tools were near-perfect on essays written by US eighth-graders 2. The error is not evenly distributed. It falls on people writing in a second language, and national origin sits directly behind that, which is not a distinction you want a rejection resting on. Whether AI detectors work at all in hiring is its own question, and the answer has not improved.
So treat "clearly used AI" as a hypothesis, and test it the cheap way. The artifact will not settle it. Ask about the work: which figure the recommendation rests on, where it came from, what the second-best option was and why it lost. Someone who did the thinking answers in specifics inside a minute, and someone who did not moves to generalities and stays there.
What does unchecked AI actually cost in your field?
That is the question worth answering, and the answer changes by role. The failure mode is not an assistant writing something. It is a confident wrong thing surviving to a place where nobody catches it. In marketing copy the cost is an edit. In a valuation, a denial appeal or a merged pull request, the cost is a decision made on a number nobody opened.
The errors that get through are the near misses. In Stack Overflow's 2025 developer survey the most-cited frustration with AI tools was "AI solutions that are almost right, but not quite," named by 66% of respondents; roughly 46% distrust the accuracy of AI output against 3.1% who highly trust it 3. Almost-right is the expensive class. It survives a read-through and fails downstream, which is exactly the material a polished take-home is made of.
The person holding the output is also the worst judge of it. In a Stanford study of developers writing security-relevant functions, participants with access to an AI assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure 4. Confidence moved the wrong way. "The submission looks fine" is not a check, and neither is a candidate's own account of how carefully they read it.
What unchecked looks like, by the field you are actually hiring for:
- Marketing. The most on-message statistic in the brief comes from a vendor's own 40-person customer survey, written up as market-wide. Cheap to catch, cheap to fix, embarrassing in a deck to a client.
- Financial analysis. A growth figure the filing behind it does not support, or a multiple applied to the wrong base. It moves a recommendation, and the reader has no way to see it.
- Software engineering. Code that passes the tests it was handed. The fixture agrees with the bug.
- Legal operations. A citation that resolves to a real case and a paragraph saying something else. Retrieval is not verification.
- Data and analytics. A fluent explanation of a result the dataset does not support, restated more carefully each time you push on it.
Grade against that column, not against tool use. A candidate who generated a first draft and then re-added the segment against the filing did the job. One who wrote every word by hand and shipped an unchecked number did not.
Grade the submission again, then ask about it
Read it a second time for the checks rather than the tool. Four questions answer off the artifact: which claim is the deliverable resting on, what it was checked against, what is cited but plainly never opened, and where a range has been reported as a point estimate. Those separate a candidate who ran the check from one who typed a confident sentence about it.
Polish on its own tells you nothing, and grading a take-home when every submission comes back polished is the general version of this problem. Two graders should be able to fill those four columns separately and agree. "Showed good judgment" cannot be scored twice the same way; "reconciled the segment figure to the 10-K" is a yes or a no.
Then book twenty minutes and ask about their own submission. This is where the disqualification question dissolves, because the conversation gets you evidence and the accusation gets you an argument. The Uniform Guidelines hold that content validity runs to the extent the procedure is a representative sample of the content of the job, and that a procedure resting on inferences about mental processes cannot be supported by content validity alone 5. "Did they cheat" is an inference about a mental process. "Walk me through how you settled the discount rate" is the work.
The same three rungs that make follow-up questions expose real understanding apply here: specify, invert, falsify. Ask what the number would have to be for the recommendation to flip. Ask what they could not settle in the time. Ask the same set of every candidate who got that brief, so the round stays comparable and the record shows it. See how Olive measures this.
Write the rule before the next candidate
One line in the brief settles most of this: whether AI is allowed, what has to be disclosed, and what gets graded. Say the tool is permitted, ask for a short note naming what was generated and what was verified, and tell them the note is part of the deliverable. A rule nobody stated is not a rule anyone broke, and a candidate reading your brief has no way to guess at it.
The three lines, verbatim enough to paste:
- AI use is allowed on this task.
- Include a short note: what you generated, what you verified, and how you verified it.
- The note is graded. The polish is not.
Then change what the task asks for, because a brief that wants a finished artifact now measures whose assistant was better. Hand over source material with one thing wrong in it that only checking catches, ask which claim was checked and what changed as a result, and cap the whole thing at an hour. Whether to allow AI on the take-home at all should follow what the job permits: an agency where every draft starts in a model and a bank with a locked-down policy are answering different questions.
The hour between sending the brief and reading the deliverable is the part you still never see. A verification note is written afterwards by the person it describes. It leaks too: a dozen candidates in, the packet is posted somewhere and the task measures recall. Budget the second case at the start, keep one rubric across both, and treat the take-home as one observation from one afternoon, which is what it has always been, with or without an assistant in the room.
Common questions
Can you reject someone for using AI on a take-home?
Yes, if the brief said not to and you hold everyone who got that brief to the same line. The exposure is not the decision, it is the inconsistency: a rule enforced on the candidate whose writing struck you as suspicious and not on the one whose style reads as normal is a different test given to different people. Put it in the brief, keep the record of how it was applied, and never rest the call on a detector score.
What if the candidate denies using AI and you are sure they did?
Drop the accusation and test the work. Ask them to walk through two decisions in the submission (where a figure came from, what the discarded alternative was, what they could not settle) and let the answers decide it. Someone who did the work answers in specifics; someone who did not hedges into generalities. You get usable evidence either way and no conversation you would regret seeing in writing. An accusation you cannot support costs you the candidate and their account of your process.
Should you require candidates to disclose AI use?
Ask for something more useful than yes or no: a short note naming what was generated, what was verified, and against what. A disclosure checkbox tells you nothing and invites a lie. A verification note is gradeable, and it lets an honest candidate show the work the finished artifact hides. Say in the brief that the note is part of the deliverable, so nobody reads it as an admission against interest.
Does allowing AI make the take-home meaningless?
Only if the task still grades the output. A brief asking for a polished artifact now measures whose assistant was better. A brief that hands over source material with one thing wrong in it, then asks which claim was checked and what changed as a result, measures the candidate. Keep the take-home, change what it asks for, cap it at an hour, and pay for anything longer.
Is a live session better for this than a take-home?
For watching the checking happen, yes. Unwatched, you see the deliverable and not the decisions, and the decisions are the signal. The cheap hybrid is a short take-home plus twenty minutes on their own choices: the artifact gives the conversation something to be about, and the conversation gives you the record the artifact cannot. Run the walkthrough with every finalist, not only with the one whose submission looked suspicious to you.
What if the whole submission reads like the assistant wrote it?
Then grade it as work, not as prose. A deliverable that is fluent and hollow fails on its own terms: no claim traced to a source, no alternative considered, false precision where the data was thin. Name those in the feedback rather than naming the tool. If the candidate can defend every load-bearing decision in a twenty-minute conversation, the fluency was never the problem you thought it was.
References
- 1. AI at Work Is Here. Now Comes the Hard Part (2024 Work Trend Index Annual Report) ✓ microsoft.com Survey of 31,000 knowledge workers across 31 markets: 66% of leaders say they would not hire someone without AI skills, and 71% would rather hire a less experienced candidate with AI skills than a more experienced one without them.
- 2. GPT detectors are biased against non-native English writers ✓ pmc.ncbi.nlm.nih.gov Seven detectors over 91 human-written TOEFL essays: 61.3% average false-positive rate, 97.8% flagged by at least one detector and 19.8% by all seven, against near-perfect accuracy on US eighth-grade essays.
- 3. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co The top-reported frustration with AI tools is "AI solutions that are almost right, but not quite" at 66%; about 46% of respondents distrust the accuracy of AI output against 3.1% who highly trust it.
- 4. Do Users Write More Insecure Code with AI Assistants? ✓ arxiv.org Participants with access to an AI code assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure.
- 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 ✓ ecfr.gov Content validity holds to the extent the selection procedure is a representative sample of the content of the job; a procedure resting on inferences about mental processes cannot be supported by content validity alone.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.