Policy
Should You Allow AI on the Take-Home Assignment?
Allow AI on the take-home unless the job itself doesn't. Copy the rule off the seat: if the team drafts in a chat window every day, a ban tests obedience instead of skill. Ban assistants only where the job genuinely runs without them, meaning regulated records that can't enter a third-party model, air-gapped systems, or work whose value is unaided recall. Then write the rule in three lines (what's allowed, what must be disclosed, what gets graded) and put those lines in the brief, worded identically for every candidate.
The takeThe part that gets challenged later is never the rule you picked. It's the record of how you applied it. A ban leaves you two documents to defend it with, and neither holds: a detector output nobody should rely on, and a candidate's own account of tools people misremember using. An open rule with a graded record gives you the opposite: one written standard, the same brief for everyone, and a page in the candidate's own words to point at. Most likely the safest policy is simply the one that leaves you something to show.
Where Olive fits
Open a role and see what the work shows
A take-home nobody watched leaves no record of whether the rule was followed, which is the part a policy has to be defensible about. Olive runs the assignment with an AI assistant in the open and returns six findings a human wrote by hand, each anchored to a timestamped excerpt and exported with its rubric, scorer and bank versions; the candidate is granted the identical document.
Rank your shortlistShould you allow AI on the take-home, or ban it?
Allow it if the job allows it, and ban it if the job bans it. The take-home is a sample of the work, and a work sample earns its validity by approximating the real work situation. The Uniform Guidelines put it in exactly those terms 1. A rule that contradicts the day job stops measuring skill and starts measuring compliance with an arbitrary instruction.
That reframes the question usefully. "Should candidates use AI" has no general answer, because the thing being sampled is not general. An agency where every deck starts in a chat window and a bank whose model policy forbids pasting client data into a third-party service are hiring for different work, and a take-home that ignores the difference tests neither of them. Copy the rule off the job, not off a blog post.
It also prices a ban for a role that doesn't have one. Ban AI on a take-home for a team that runs on it, and your strongest candidates either comply and produce work no colleague would recognize, or don't comply and say nothing. You have selected for willingness to follow a rule that gets reversed on day one, and you have spent the one chance to watch someone work the way they will actually work.
The reverse costs as much. Hand an AI-open case to a candidate for a seat where the tools are locked down, and the best submission may come from whoever is best at something the job will never permit. What AI actually does inside the role is the input to this decision, and it is usually one conversation away.
How do you decide which rule this role gets?
Ask the people already doing the job three questions: which assistants they can open on a work machine, what they are forbidden to paste into one, and what a normal Tuesday's output owes to a model. The answers give you the rule directly. Don't infer it from the industry: teams inside the same bank differ, and the written security policy is usually narrower than the practice.
Expect the answer to be less AI-saturated than the discussion around it. In a Pew Research Center survey of 5,273 employed U.S. adults fielded in October 2024, 55% said they rarely or never used AI chatbots at work, and about one in ten used one every day or a few times a week 2. If the team you are hiring into sits in that 55%, an AI-required take-home is testing something the job does not ask for.
Three conditions cover almost every role:
| What the job permits | The take-home rule | What it measures |
|---|---|---|
| Assistants open and used daily, no data restriction on this material | AI allowed and expected; the record of its use is part of the submission | How the candidate frames, delegates, checks and refuses, which is the work as it will happen |
| Assistants permitted, but data cannot leave the tenant or only one tool is sanctioned | AI allowed on the candidate's own tools; the case carries nothing that would breach the real policy | The same acts, plus whether the constraint is noticed and worked inside |
| No third-party model may touch the work (air-gapped systems, regulated records, unaided recall) | AI not permitted, with the reason in one clause | Unaided competence, which is what the seat actually requires |
Write the case material to fit the rule rather than the other way round. If the role's real constraint is that client data cannot go into a model, build the case out of public or synthetic material so the rule is honestly followable, and say that it is. A brief that quietly makes compliance impossible is a trap, and candidates read it as one.
Write the policy in three lines
Three lines cover it: what's permitted, what must be recorded, and what gets graded. Keep them under forty words in total, put them in the brief rather than on a linked policy page, and use identical wording for every candidate on the role so nobody sits a different test. Anything longer than three lines is a legal document, and candidates skim legal documents.
The AI-open version:
> Tools. Use any AI assistant you'd use at work. A free tier is fine, and the choice of tool isn't graded. > Record. Include one page: what you asked before generating anything, what you assumed, what you checked outside the assistant, and one thing it produced that you didn't use. > Grading. The record and the deliverable are read together. How much AI you used isn't scored.
The restricted version changes only the first line: "Use any AI assistant, but this packet contains no client data and none may be added. Treat it the way you'd treat a live file." The closed version replaces it with a reason: "Work without an AI assistant on this one. The seat runs on systems where that isn't available, and this is the closest sample the exercise can take."
Three details do most of the work. Naming a free tier removes a spending test. Saying the tool choice isn't graded stops candidates buying a subscription to look serious. And saying volume isn't scored is what makes an honest amount safe to report, including none, which is a legitimate answer from someone who judged the model was the wrong instrument for a step.
What goes in the brief, word for word?
One paragraph, at the top of the brief, before the task. It states the rule, names what you'll ask about afterward, and says plainly that disclosure isn't held against anyone. That last clause is load-bearing: a candidate who thinks admitting AI use will cost them the job writes a fiction instead, and then you have neither an honest artifact nor an honest record.
Paste this, adjusted to the rule you picked:
> AI assistants are allowed on this exercise, and most people on the team use them. Along with the deliverable, send one page describing how you used one: the questions you asked first, the assumptions you made, what you checked against a source outside the chat, and anything the assistant suggested that you rejected. That page is graded. The amount of AI in your process is not, and using none is a valid answer. A short conversation about your submission follows.
State the time cap in the same breath and mean it. Sixty to ninety minutes is the working range for an AI-open case, and a case that needs a full day selects on free time rather than judgment. Name a channel for questions with a stated response time too, because the questions that arrive before anyone generates anything are the highest-signal data the exercise produces, and they are most of what a take-home still measures once AI is in it.
Don't add a clause promising to catch undisclosed use. You can't, the claim is checkable, and a candidate who knows it's empty reads the rest of the brief as theater. Disclosure, detection or observation is a real fork with three different costs, and only two of the three are actually available to you.
When is a ban actually the right call?
When the job genuinely runs without these tools, and only then. A regulated seat where nothing client-identifying may enter a third-party model, a task performed on an air-gapped system, a role whose value is unaided recall under time pressure. Those are real, and a ban mirrors the work. Everywhere else a ban buys nothing enforceable, because AI-generated text cannot be reliably identified 3.
On that point the evidence is not close. An evaluation of fourteen detection tools found them neither accurate nor reliable, systematically biased toward classifying generated text as human-written, and degraded further by light editing 3. Acting on one of those outputs means accusing some honest candidates and clearing some who did exactly what the brief forbade, with nothing to tell you which is which.
Self-report doesn't rescue it either, and the failure isn't dishonesty. In a controlled study of sixteen experienced developers working 246 real issues on their own repositories, being allowed AI tools made them 19% slower, and afterward they still believed the tools had sped them up by 20% 4. People are unreliable narrators of their own AI use even when they're trying hard to be accurate. Ask what they did, not how much it helped.
So if you ban, ban for a reason you can state in one clause, and accept that the rule runs on trust plus a conversation. Fifteen minutes on the submission (pick one decision, ask what would have changed if the input were different) separates work someone did from work someone received far better than any policy sentence does. That conversation is also where a strong take-home that turns into a weak first quarter tends to become visible first.
What should you grade once AI is allowed?
Grade four acts, none of which appear in the finished artifact: the question asked before the first generation, the assumption written down, the claim checked against something outside the chat, and the direction turned down. A survey of 319 knowledge workers describing 936 real tasks found the thinking that survives AI use shifts toward verification, integration and oversight 5. Score those, separately.
Keeping them separate is the point. A candidate who framed the problem well and verified nothing is a different hire from one who checked everything and never questioned the brief, and one blended rating hides exactly that difference. Three words per row (demonstrated, partly demonstrated, not demonstrated) plus the excerpt that made you say so is enough, and it survives being read back to the candidate. See how Olive measures this
The same survey carries a warning worth building into the rubric: higher confidence in the tool went with less critical thinking 5. The submission that reads most confidently is therefore not automatically the strongest, and a reviewer grading polish picks the wrong person. Grading take-homes that all come back polished is the scoring half of this problem.
No row in that rubric counts prompts. Someone who decided the model was the wrong instrument for a step and did that part by hand has demonstrated the judgment the exercise was built to surface, and scoring volume selects against them. If the policy says the amount isn't graded, the scoring sheet has to say it too, and what belongs in an AI hiring policy is mostly the work of making those two agree.
Common questions
Does allowing AI make the take-home too easy?
It makes the production half easy, and that half was never the part worth grading. What gets harder is everything an assistant answers confidently and wrongly: the assumption nobody stated, the figure that doesn't reconcile, the source that says something narrower than its summary. If the case contains none of those, allowing AI does hollow it out, but the fix is the case, not the rule. Put one genuine ambiguity in the material and the exercise gets more discriminating with AI open, not less.
What if a candidate uses AI when the brief banned it?
Start from the fact that you probably can't prove it. Detection tools are unreliable and biased toward calling generated text human-written, so a suspicion isn't evidence and an accusation on that basis is wrong sometimes and expensive always. Ask instead: fifteen minutes on the submission, one decision picked out, what would have changed if that input were different. Someone who did the work answers immediately. If the answer is empty you have a judgment about the work rather than an allegation about tooling, and that is a defensible reason to pass.
Should every role at the company use the same rule?
No. The rule follows the job, and the jobs differ. The seat drafting customer copy all day and the seat touching regulated records don't share working conditions, so one company-wide rule is wrong for at least one of them. What should be uniform is the shape: the same three lines, the same disclosure promise, the same statement that volume isn't graded. Identical wording across candidates for one role matters far more than identical policy across roles, because that is what keeps the comparison fair.
Can you require candidates to use a specific AI tool?
Only if the job does, and even then say a free tier is acceptable. Requiring a paid subscription turns the exercise into a spending test; requiring an unfamiliar tool measures onboarding speed rather than judgment. If the role really is standardized on one assistant, provide access for the exercise or accept an equivalent. Naming a tool as forbidden is a different move and often reasonable: "nothing that stores this packet outside your own machine" is a constraint the job has, written so a candidate can act on it.
Do you have to tell candidates how their AI use will be judged?
Yes, and it costs nothing. A candidate who doesn't know how the record will be read either over-explains or hides, and both distort the evidence you were trying to collect. One line does it: the record is graded, the tool choice isn't, the amount isn't. Say it in the brief rather than on a policy page nobody opens, and use the same sentence for every candidate on the role. If the wording drifts between candidates you no longer have a comparison, you have two exercises with one name.
Is banning AI the legally safer choice?
Not inherently. A selection procedure's defensibility rests on how closely it samples the job's actual content and work situation 1, and a ban contradicting the day job weakens that link rather than strengthening it. What genuinely helps is the same under either rule: one written standard applied identically to every candidate for the role, a rubric fixed before submissions arrive, and a recorded reason for each decision. Check separately what notice your jurisdiction requires if any part of the process is automated.
References
- 1. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 ✓ ecfr.gov Content validity holds to the extent a procedure is a representative sample of the content of the job, and where a work behavior is sampled the manner, setting, level and complexity should closely approximate the work situation.
- 2. U.S. Workers Are More Worried Than Hopeful About Future AI Use in the Workplace ✓ pewresearch.org Survey of 5,273 employed U.S. adults fielded October 7-13, 2024: 55% say they rarely or never use AI chatbots at work, and about one in ten use one every day or a few times a week.
- 3. Testing of Detection Tools for AI-Generated Text ✓ arxiv.org Fourteen detection tools judged neither accurate nor reliable, biased toward classifying generated text as human-written, and degraded further by light editing or obfuscation.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org Sixteen experienced developers on 246 real issues in their own repositories took 19% longer when allowed AI tools, and afterward still believed the tools had sped them up by 20%.
- 5. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org Survey of 319 knowledge workers and 936 first-hand examples: the thinking that remains shifts toward information verification, response integration and task stewardship, and higher confidence in the tool is associated with less critical thinking.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.