Policy

Should Interns Be Allowed to Use AI on the Take-Home, and How Do You Grade It?

Let interns use AI on the take-home, and ask for the record: the transcript, a decision note, the intermediate artifact. A ban only holds up where you can watch the work happen live. Sent home unwatched, it rests on detectors that flagged 61.3% of essays by non-native English writers as machine-written [3]. So grade four decisions instead: the scope set before generating, one claim checked against a source, the work kept by hand, and the output refused with a reason. Never grade how much was used.

The takeThe instinct here is that permitting the tool is the risky policy and banning it is the careful one. I read it the other way round. A ban leaves you enforcing a rule with an inference about who wrote a document, and that inference falls hardest on applicants writing in a second language, which puts national origin one step behind your rejection letter. Permission plus four anchored rows leaves a filled-in sheet for every applicant instead. When somebody asks why they were cut, one of those policies has an answer on paper and the other has a hunch.

Where Olive fits

Open a role and see what the work shows

A graded take-home is a selection procedure, and "the submission looked strong" is not a record anyone can check. Olive returns six separately-evidenced findings a human wrote (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each carrying the moment in the session it rests on, with no composite number, no hiring recommendation, and the identical report granted to the candidate.

Rank your shortlist

Should interns be allowed to use AI on the take-home?

Yes. Allow it, require disclosure, and change what the rubric scores. A ban costs you enforcement you don't have: seven widely used detectors flagged 61.3% of essays by non-native English writers as machine-written while classifying US student essays accurately 3. A prohibition you police by guessing mostly punishes the honest and the international. Permission plus a rubric is the version you can defend.

The enforcement problem is the whole argument. Nothing in a submitted file tells you who did the thinking, and the tools that claim otherwise measure predictability rather than authorship. The same study rewrote those essays with richer vocabulary and watched the average false-positive rate drop from 61.3% to 11.6% 3. What the flag tracks is polish, which is exactly what a well-resourced applicant can already buy. The general case, whether AI detectors work well enough to screen on, resolves the same way at every level of seniority.

Interns arriving this cycle have not done coursework without a model available. That is the stronger argument, and it does not depend on enforcement at all. Forbidding the tool tests a working condition that will not exist on their first day, and turns the exercise into typing under an artificial constraint. Take-homes still tell you something when AI is allowed, but only once the rubric stops rewarding the finished object.

Three things a ban actually produces:

  • Submissions that used AI and say they didn't, which is the one outcome you cannot grade.
  • A disclosure rule enforced on suspicion, which is where a discrimination claim starts.
  • A signal about compliance instead of a signal about judgment.

Score four decisions the model cannot make

The rubric has four rows, and none of them is about the deliverable's quality. Score the scope the intern set before generating, the claim they checked against a source outside the chat, the work they deliberately kept by hand, and the assistant output they rejected with a reason. All four are visible in the record. None of them improves when the model gets better.

Write each row with two anchors (what earns credit and what does not), and hand the same sheet to every grader.

1. Scope set before generating. Credit: the intern states in their own words what is being decided and what would make an answer wrong, before the first generation. No credit: the first message to the assistant is a request for the deliverable.

2. One claim checked outside the chat. Credit: a specific claim in the deliverable is traced to a source the intern opened, and the deliverable says what that source did or did not support. No credit: the assistant's summary of a source is treated as the source.

3. Work deliberately kept. Credit: something of consequence was done by hand and the intern can say why: a figure recomputed, a judgment made without asking. No credit: everything of consequence was handed over, with no act of their own anywhere in the record.

4. Output refused with a reason. Credit: a direction the assistant proposed is rejected on substance: an assumption named, a framing dropped, a section cut because it was doing no work. No credit: the deliverable is the assistant's output with the edges tidied.

The fourth row is the one graders undervalue. In a controlled study of AI-assisted programming, participants who were more skeptical of the assistant and reworked their requests produced code with fewer vulnerabilities, while the assisted group overall wrote less secure code and was more confident it was secure 2. Refusal is not friction. It is the behavior that separates two submissions that look identical on the page.

Two rows you should not add: how much AI was used, and how fast the work came back. Volume is not a virtue, and an intern who judged the model was the wrong instrument for a step and did that step by hand has demonstrated the thing you are grading.

Which decisions are field-specific?

Most of them. The four rows stay constant; what earns credit inside each row is set by the occupation. Scoping is the graded decision in consulting, because the assistant will accept whatever boundary the brief implies. Data provenance is the graded decision in analytics. Failure modes are the graded decision in engineering. Write the anchor sentences for one field rather than for all of them.

Consulting and strategy. Hand an intern a market-sizing brief and an assistant, and the assistant will size whatever market the brief names, including the wrong one. Credit lands on narrowing the question before generating: naming the segment, the geography and the time window, and saying which of the three the client's decision actually turns on. A submission that sizes three markets fluently and never says which one matters has failed row one, however good the arithmetic is.

Data and analytics. O*NET's task list for data scientists puts cleaning raw data, identifying factors that could affect a result, and testing and validating models alongside the analysis itself 5. An assistant will explain a result fluently from a dataset that will not correct it. Credit lands on asking where a column came from, checking a total against the raw file, and stating a limit the data imposes on the recommendation.

Engineering. The assistant writes the happy path immediately and convincingly, and the assisted participants in the study above were both less secure and more confident 2. Credit lands on naming what could break before accepting the change (the input that isn't there, the call that times out, the migration that runs twice), and running something that proves it.

Marketing, journalism, operations, legal. Same four rows, different graded decision: which statistic gets opened, which contract clause gets read against the system record, which quoted figure gets checked with the person who gave it. Pick one occupational decision, write its anchors, and stop. A rubric stretched across every function grades none of them, which is the one-assessment-or-per-role question in miniature.

How do you grade the record instead of the deliverable?

Ask for the record as part of the submission, and read it first. Require the chat transcript, a short decision note naming what was checked and what was refused, and the intermediate artifact: the scope, the query, the test plan. Read those before the deliverable. Don't ask an intern how much AI helped: developers in one controlled study ran 19% slower with AI tools and still believed afterwards that they had run 20% faster 1.

So collect artifacts rather than accounts:

  • The transcript. Pasted in by the candidate is fine. What you need is the sequence of moves, not a proxy for honesty.
  • A decision note under 200 words. What was checked, what was refused, what was kept by hand. This is the document rows two through four are graded against.
  • The intermediate artifact. The scope, the criteria list, the query, the test plan, whatever the deliverable was supposed to come from.
  • A 15-minute follow-up call. Two questions: walk me through the part you changed, and show me the claim you checked.

The follow-up is what makes the rest hold. An intern who wrote the decision note can walk it; one who generated it stalls on the second question, and nothing had to be detected for that to happen. It is also the cheapest fix for submissions that all come back polished, because polish stops carrying information the moment the record is the thing being graded.

Cap the exercise. Two hours is a fair outside edge for an intern take-home, and a longer one selects for free time rather than for judgment. State the cap, say you will read what arrives inside it, and hold to that. An unpaid ten-hour assignment filters your pipeline by who can afford a weekend.

Write the rules into the brief before it goes out

Four lines in the assignment, published before anyone starts. AI use is permitted and must be disclosed. Here is what the submission must include. Here is what the rubric scores. Here is the time cap. Publishing the rule after the submissions arrive turns it into a justification, and a graded take-home is a selection procedure the EEOC expects to be job-related 4.

The four lines, close enough to paste:

1. Permission. Any AI assistant is allowed on this assignment. Say which ones you used. 2. What to submit. The deliverable, the transcript, the decision note, the intermediate artifact. 3. What is graded. The four rows, named in the brief. Not the polish of the deliverable. 4. Time. The cap, and what happens to work submitted past it.

Then keep the rubric and the filled-in sheets. A graded take-home decides who advances, which makes it a selection procedure, and the EEOC's guidance expects one that screens people out disproportionately by race, sex or national origin to be job-related and consistent with business necessity 4. "The submission looked strong" is not a record. Four rows with an anchor sentence each, filled the same way for every applicant, is one. Putting every grader on that sheet is the cheapest correction for managers guessing at which parts an AI wrote instead of grading against anything.

One fairness rule worth writing down: an intern who used the assistant lightly and did most of the work by hand cannot lose points for it. If they framed the problem, checked the claim, kept the work and can say why, all four rows are earned. Grading usage would quietly convert permission into a requirement, which is a different policy with consequences an AI-first mandate has to plan for.

Read the evidence

Common questions

Do you have to allow AI on an intern take-home?

No, but a ban you cannot enforce is worse than no rule. If you forbid it, you are relying on self-restraint plus a detector, and the detector's errors fall hardest on applicants writing in a second language. A closed take-home is defensible only when it is supervised in real time: a live working session with the tools you name. If the assignment goes home unwatched, permission plus a rubric on the record is the honest version.

What if an intern uses AI and doesn't disclose it?

Make disclosure part of the submission rather than a question you ask afterwards, and the problem mostly disappears. Require the transcript and the decision note as deliverables, so an undisclosed session shows up as a missing artifact instead of a suspicion. Handle a missing artifact the way you would handle any incomplete submission: ask once, in writing, and grade what arrives. Never open an investigation into authorship you cannot resolve, which is a verdict about a person built on a guess.

How long should an intern take-home be if AI is allowed?

Two hours or less, and say the number in the brief. Allowing the assistant does not justify a bigger assignment; it means the same brief now produces a finished-looking artifact faster, so the extra time buys polish rather than signal. A long unpaid assignment also filters on who has a free weekend. If two hours is not enough to see the decisions you care about, the brief is testing production volume, not judgment, and the fix is a narrower task.

Can you grade an intern down for using AI too much?

No. Volume of usage measures nothing you want to hire on, and penalizing it converts your permission into a trap. An intern who framed the problem, checked one claim against a source, kept a decision to themselves and refused an unhelpful direction has earned every row, whether that took four prompts or forty. The reverse also holds: a submission with a long transcript and no refusal anywhere in it has demonstrated less than a short one with a stated boundary.

Does an open-AI take-home create legal exposure?

It is a selection procedure either way, so the exposure comes from how it is graded rather than from allowing the tool. The EEOC treats anything used to make an employment decision as a selection procedure, and where one screens people out disproportionately on a protected basis it has to be job-related and consistent with business necessity. That argues for the rubric, not against the assignment: named rows, anchor sentences, the same sheet for every applicant, and the filled-in sheets kept. An unwritten impression of a submission is the version with no defense.

Where does Olive fit in an open-AI take-home?

Olive is an employer-purchased assessment of how a person works with AI, so it covers the same ground as the rubric above with the authoring and the reviewing already done. An intern works a 40-to-60-minute task built for one occupation with an assistant available, and a human reviewer writes six findings, each anchored to a moment in the session. There is no composite number and no hiring recommendation, and the candidate is granted the identical report.

References

  1. 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized controlled trial: 16 experienced developers across 246 issues took 19% longer with early-2025 AI tools, and still believed afterwards that the tools had made them 20% faster.
  2. 2. Do Users Write More Insecure Code with AI Assistants? Perry, Srivastava, Kumar and Boneh (Stanford University), ACM CCS 2023, 2023. arxiv.org Participants with an AI assistant wrote significantly less secure code yet were more likely to believe their code was secure; those who were more skeptical of the assistant and reworked their requests produced fewer vulnerabilities.
  3. 3. GPT detectors are biased against non-native English writers Patterns (Cell Press), 2023. pmc.ncbi.nlm.nih.gov Seven detectors over 91 TOEFL essays and 88 US eighth-grade essays: 61.3% average false-positive rate on the TOEFL essays, falling to 11.6% after a vocabulary-enrichment rewrite.
  4. 4. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Anything used to make an employment decision is a selection procedure; where it disproportionately screens out a protected group, the employer must show it is job-related and consistent with business necessity.
  5. 5. 15-2051.00 - Data Scientists O*NET OnLine, U.S. Department of Labor, 2026. onetonline.org Occupational task list includes cleaning and manipulating raw data, identifying factors that could affect research results, and testing and validating models.

5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.