Assessment design
Do Take-Home Assignments Still Tell You Anything if Candidates Use AI?
A take-home still tells you something when candidates use AI, but only about the part of the brief you left open. Keep the assignment and change what it asks for: author one real ambiguity into the material, a missing assumption, a defective dataset, two constraints that cannot both hold. Name a channel for questions and answer within a day, then grade what the candidate asked, checked outside the chat and refused. Unwatched, none of it proves who did the work; fifteen minutes on their own submission does.
The takeThe authoring cost is the half everyone budgets: a day per case, a rewrite after the first few submissions, and it lands on a calendar. The other half is answering the questions an open brief provokes, on a deadline, every round, and nothing on a hiring calendar protects that afternoon. If the pattern holds, the case survives the redesign and the channel is what quietly lapses. An open brief with nowhere to send a question is a guessing game with an answer key. The ambiguity is the part you write. Whether it measures anything depends on somebody being there to answer.
Where Olive fits
Open a role and see what the work shows
If you author this in-house, the recurring costs are the ambiguity that has to be genuine for the occupation and the answer key that says what a good question looks like. Olive ships twelve authored cases per occupation and returns six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session and granted to the candidate in the same document.
Rank your shortlistWhat does a take-home still measure once AI is in it?
It measures whatever the brief left open, and nothing else. Everything the brief fully specified is now free. The assistant produces it, and every submission converges on the same competent shape. What survives is the handling of ambiguity: which assumption got surfaced, which supplied number got refused, which pair of constraints got noticed as incompatible. A brief containing no ambiguity measures nothing at all.
That is a smaller loss than it feels like, because the take-home was never the strongest thing in your loop. The 2023 re-analysis of selection-procedure validity puts work samples at an operational validity of .33, behind structured interviews at .42 and job knowledge tests at .40 1. What AI removed is the production half, the part you were grading by accident, because it was easy to see. The judgment half was always the part worth paying for, and it was always the part nobody authored on purpose.
So the useful question is not whether candidates used AI. It is whether your case has anything in it that an assistant answers confidently and wrongly. If the answer is no, the assignment is a typing test with a two-day deadline, and no scoring rubric rescues it. If the answer is yes, you have a work sample that got harder to fake rather than easier, because the failure now happens in front of you: the candidate who never asked what the missing assumption was ships a memo built on it.
Why does a fully-specified brief measure nothing now?
Because a fully-specified brief is a prompt. Paste it into any assistant and the deliverable arrives structured, formatted and confident. Real work does not arrive that way, and the Uniform Guidelines are explicit that a selection procedure earns content validity as a representative sample of the job's content, with the manner, setting, level and complexity of the sampled behavior closely approximating the work situation 2. Under-specification is part of that situation.
Think about which analyst you would actually hire. Nobody hands a financial analyst a packet and a stated discount rate. Nobody hands a data analyst a clean file. Nobody hands a marketer a brief whose budget, timeline and positioning all agree. The moment you sand those edges off to make grading easier, you have removed the only content the job was made of and left behind a formatting exercise.
The second reason is that a finished artifact is weak evidence about the work that produced it. In a controlled study of sixteen experienced open-source developers across 246 real tasks on their own repositories, allowing AI tools made them 19% slower, and afterward they still believed the tools had sped them up by 20% 3. If people cannot read their own process off their own output, a reviewer certainly cannot read a stranger's off a submitted document. Grading polish is guessing, and it is a guess that favors whoever had the better assistant. Grading take-homes that all come back polished is a scoring problem downstream of a design problem.
How do you write a brief an assistant can't finish?
Put one real problem in the material rather than in the wording. Three devices carry almost every case: an assumption the supplied data cannot settle, an artifact with a genuine defect in it, and two constraints that cannot both hold. Each one has an answer the assistant will produce fluently and wrongly, and each one has to be authored out of the occupation's own work.
| Device | What goes in the packet | What the assistant does with it | The act that earns the score |
|---|---|---|---|
| Unsettled assumption | A model with no stated growth rate, and a footnote that contradicts the headline figure | Picks a plausible rate and never says it picked one | The assumption is named, sourced or flagged, and its effect on the answer is shown |
| Defective artifact | A dataset with duplicated records, a mid-period unit change, or a column that silently means two things | Explains the result it computes, confidently and coherently | The defect is found, and the recommendation changes because of it |
| Conflicting constraints | A launch brief whose budget, deadline and channel mix cannot all be met | Produces a plan satisfying all three on paper | The conflict is stated and a trade-off is proposed with its cost |
Two rules keep this honest. The ambiguity must be resolvable by asking rather than by guessing. Otherwise you are testing telepathy and rewarding the candidate whose guess matched yours. So name a channel in the brief and answer within a day: "Questions to this address; expect a reply the same working day." The questions that arrive are the highest-signal data you will get from the whole exercise, and they cost you nothing to collect.
And the ambiguity has to be real for the field, not a puzzle. A trick with a single clever answer measures whether the candidate has seen that trick. A discount rate nobody stated, a survey question with a leading construction, a contract clause the system record and the signed copy disagree about: those are Tuesday, and an experienced person recognizes them as Tuesday. Write the case with somebody currently doing the job, then dry-run it against two more.
Budget for it: roughly a day per case including an answer key, plus a rewrite after the first three candidates, because version one is always either guessable from the brief or undiscoverable inside the time limit. This is per role, and it recurs, because cases leak and a leaked case measures preparation. Redesigning the work sample around what AI cannot do and screening for AI judgment outside engineering both bottom out in this same authoring cost.
Grade the questions asked before anything was generated
The record matters as much as the deliverable, and four acts in it are worth scoring: the questions asked before the first generation, the assumptions written down, the claims checked against something outside the chat, and the direction the candidate turned down. A survey of 319 knowledge workers describing 936 real tasks found the thinking that remains under AI shifts toward verification, integration and oversight, and that higher confidence in the tool goes with less of it 4.
Make the record cheap to produce or you will not get it. One page, four prompts, written as the work happens rather than reconstructed afterward:
- What I asked before I generated anything: the questions sent to you, and the ones answered by reading the packet.
- What I assumed, and what it would take to be wrong: each assumption with the figure it moves.
- Two claims I checked, and how: the source opened, the number recomputed, the query run. Not "verified with the model."
- One thing the assistant proposed that I did not use, and why: the substantive refusal, not a style edit.
Score each row separately as demonstrated, partly demonstrated or not demonstrated, and keep the excerpt that made you say so. Resist blending them into a total: a candidate who framed the problem well and verified nothing is a different hire from one who checked everything and never questioned the brief, and one number hides exactly that difference. Two reviewers should land on the same three words without discussion; when they do not, the anchor sentence under the row is ambiguous, which is a rubric problem rather than a candidate problem. The same rows work on an interview answer the candidate produced with AI.
Tell candidates in the brief that the amount of assistance is not measured, because a candidate who believes volume is being scored will generate in order to be seen generating. The one who decided the model was the wrong instrument for a step and did it by hand has demonstrated the judgment being tested, and a rubric that rewards prompt count will select against them.
Should you ban AI on the take-home, or try to detect it?
Neither works. A ban you cannot enforce selects for the candidates willing to say they complied, and detection is not available: an evaluation of fourteen detection tools found them neither accurate nor reliable, biased toward classifying generated text as human-written, and degraded further by light editing 5. Accusing a candidate on that evidence is both wrong and expensive.
Say in the brief that AI is allowed, name what you will ask about it, and make honest disclosure the low-risk answer. "Use whatever tools you use at work. The submission includes a one-page record of how you used them; that record is graded, the tool choice is not." A candidate who believes disclosure will be held against them writes a fiction, and then you have neither the artifact nor the record. Whether to allow AI on the take-home at all is settled by which of those two outcomes you prefer.
Here is the honest limit, and it does not go away with a better rubric: an unwatched take-home cannot prove who did what. A decision log can be written after the fact by the same assistant that wrote the deliverable. The cheap defense is a fifteen-minute conversation on the submission: pick one assumption from their log and ask what would change if it were wrong, then pick one claim and ask how they checked it. Specific work survives that; a reconstructed log does not. That conversation is also why a take-home and a live working session answer different questions, and why a candidate who used AI heavily is not by itself a reason to reject the submission.
If you get one thing from this: the take-home did not stop working. It stopped working unattended. The parts that used to happen invisibly (framing, asking, checking, refusing) now have to be requested explicitly, authored into the material deliberately, and read by a person.
Common questions
How long should an AI-open take-home be?
Sixty to ninety minutes, stated as a cap in the brief and enforced by the size of the packet rather than by trust. The ambiguity does the work, not the volume, and a case that needs four hours is usually one where you failed to cut the production busywork the assistant now does for free. Add a fifteen-minute follow-up conversation on the submission instead of adding scope. If the case genuinely cannot fit in ninety minutes, pay for the time.
Can't a candidate just fake the decision log?
Partly, yes: anything submitted asynchronously can be generated. Two things make it expensive. Specificity: a real log names the footnote on page 14, the duplicated record IDs, the exact figure that moved, and a fabricated one stays general. And the follow-up conversation: ask what would change if one named assumption were wrong, and ask how one specific claim was checked. Someone who did the work answers in seconds. Treat the log as evidence to be tested in conversation, not as a sworn statement.
What if the candidate asks no questions at all?
Silence counts, provided the brief invited questions and named a channel with a stated response time. Shipping a memo built on an assumption nobody stated is the exact failure the case was designed to surface, and it is worth recording as such. If the brief did not invite questions, you measured your instructions instead of the candidate. Check the invitation is unmissable: in the brief, in the email, and answered fast enough to be useful inside the time cap.
Does an AI-open take-home disadvantage candidates without paid tools?
It can, so remove the dependency. Say a free tier is fine, say the tool choice is not graded, and make sure the case does not require a capability only one paid product has: file upload, code execution, a long context window. Say plainly that working without an assistant is a supported choice, because a candidate who judged the model was the wrong instrument for a step has demonstrated the thing being assessed. Offer an alternative for anyone whose employer forbids these tools.
Do you still need a live round if the take-home is well designed?
Yes, but a short one. The take-home shows what someone does with time and their own tools; the live round shows what they do when a decision is questioned in real time. Fifteen to twenty minutes on their own submission is usually enough, and it is far more informative than a fresh exercise. Structured interviews carry the highest measured validity of the common procedures 1, so the conversation is not the soft part of the loop.
References
- 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors ✓ doi.org Revised operational validity estimates place structured interviews at .42, job knowledge tests at .40 and work samples at .33, with structured interviews top-ranked among widely used selection procedures.
- 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 ✓ ecfr.gov Content validity holds to the extent a procedure is a representative sample of the content of the job, and where it samples a work behavior the manner, setting, level and complexity should closely approximate the work situation.
- 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org Sixteen experienced developers on 246 real issues in their own repositories took 19% longer to complete issues when allowed AI tools, and afterward still believed the tools had sped them up by 20%.
- 4. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in the tool is associated with less critical thinking, and the thinking that remains shifts toward information verification, response integration and task stewardship.
- 5. Testing of Detection Tools for AI-Generated Text ✓ arxiv.org Fourteen detection tools judged neither accurate nor reliable, biased toward classifying generated text as human-written, and degraded further by light editing or obfuscation.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.