Assessment design

How Do You Grade a Take-Home When Every Submission Is Polished?

Grade a take-home on the decisions, not the artifact, because polish now costs nothing and a clean deck proves little. Ask for a short decision log with every submission: what the candidate assumed, one claim checked outside the prompt and what the check returned, and what they cut from their own draft. Rate it against a key you wrote first, and read it before opening the deliverable. A log is written after the fact, so if you need the record itself, run the exercise live.

The takePolish was never the signal, which is the part nobody wants to hear. It was a proxy for effort, one graders liked because it rated fast, and a model has now taken it away. Watch which way teams jump next. The cheap answer is a longer exercise, which mostly selects for who had spare hours. The expensive answer asks the grader to decide in advance what a good decision looks like. That is a harder ask than any candidate's two hours, and it is the part of hiring nobody gets graded on.

Where Olive fits

Open a role and see what the work shows

A decision log is a candidate's account of their own decisions, written afterward and as polishable as the artifact it accompanies. Olive runs the assignment where the record is made instead (40 to 60 minutes of occupational work with an AI assistant available), and a human reviewer writes six separately-evidenced findings, each anchored to a moment in the session, with the candidate granted the identical document.

Rank your shortlist

Why Polish Stopped Being a Signal

Because the part you were reading (structure, tone, formatting, a confident executive summary) is the part a model produces for free. Presentation used to correlate with effort, and effort with care, and that chain broke. What survives is whatever the candidate had to decide: which problem to solve, what to believe, what to throw away. None of that shows in a finished file.

The clearest measurement of how far presentation has come loose from performance is METR's 2025 trial. Sixteen experienced open-source developers worked 246 real issues from their own repositories, each issue randomly assigned to allow or forbid AI tools. With the tools allowed they took 19% longer. They had forecast a 24% speedup going in, and after finishing they still believed the tools had sped them up by 20% 1.

Two things follow for a take-home. Self-report about AI-assisted work is not evidence, so a candidate who tells you the tool saved three hours is describing an experience rather than an outcome. And the defect you are looking for is not obvious badness. In Stack Overflow's 2025 developer survey the most common frustration with AI tools, at 66%, was output that is almost right but not quite; 46% of developers said they distrust the accuracy of what the tools produce, against 33% who trust it 2.

Almost-right survives every proxy you used to grade on. It compiles, it reads well, the chart renders, the summary is coherent. It fails on the one number, the one edge case, or the one claim the whole deliverable rests on. A rubric that rewards completeness and presentation now points confidently in the wrong direction. The interview round has the same problem in a different format, worked through in Why Every Candidate Gives the Same Polished STAR Answer.

What to Grade Instead: Three Decisions

Three, and each one leaves a trace the artifact does not. What the candidate assumed and on what basis. What they went outside the prompt to verify, and what the check returned. What they rejected from a draft, and on what grounds. Those are acts with a before and an after, so they can be evidenced, disputed and rated separately. Fluency cannot.

  • What they assumed. Every real brief is underspecified, and the first act of competent work is naming what you decided to take as given. "I assumed the churn figure excludes trials, because the definition note in tab 3 says so, and the recommendation flips if it does not" is a decision with a consequence attached. A candidate who assumed nothing either had a fully specified brief, which you did not give them, or resolved the ambiguity silently.
  • What they went and got. Not what they cited, but what they opened. The distinction matters because a generated draft arrives full of plausible references. Ask for one claim the submission depends on, the source opened to check it, and what the check returned, including the case where it returned nothing useful. A survey of 319 knowledge workers across 936 real tasks found the work does not disappear when a tool is used; it moves, from producing an answer toward verifying information, integrating a response and stewarding the task. Higher confidence in the tool tracked with less critical thinking, and higher confidence in one's own ability with more 3.
  • What they rejected. The strongest single question on a take-home is what did you throw away, and why. Rejection is expensive and specific: a section cut because it was doing no work, a recommendation dropped when the number moved, an approach abandoned after ten minutes. A submission that accepted everything reads exactly like one where nothing was worth refusing, and the two separate only when you ask.

Rate the three separately and never blend them into one impression. How to Write a Rubric for AI-Assisted Answers covers the anchoring.

How to Get the Decision Trail Without Adding Hours

Ask for one page beside the deliverable and cap it. A decision log of about 300 words: three assumptions with their basis, one claim checked outside the prompt with the result, one thing cut with the reason. That is ten minutes of the candidate's time, and it turns an invisible process into gradeable text. Then say in the brief that the log is rated and the artifact is not rated alone.

The brief carries the whole design, and it needs four lines.

  • AI is allowed, and the brief says so. A ban you cannot enforce tests who believed you, and the candidates you lose are the ones who follow instructions. Should Candidates Use AI on a Take-Home? covers the policy side.
  • The log is required and rated. Name the three items and the word cap. Left unnamed, you get a process narrative; named, you get decisions.
  • One claim must be checked outside the prompt. State it as a requirement rather than a suggestion, and ask what the check returned, including when it returned nothing.
  • The deliverable is capped too. Length is not effort. Five slides, two pages, one file. A cap also stops the candidate spending their budget on formatting, which is the part you have stopped grading.

Keep the whole exercise under two hours and mean it. A take-home is a selection procedure, which puts it under the same expectation as any other test: it has to be job-related and consistent with business necessity, and the burden of showing that sits with the employer 4. Unpaid hours you cannot justify against the job are both a legal exposure and a completion problem.

What Polished-but-Empty Looks Like in Each Format

Different in every output, which is why one generic rubric misses it. A deck hides the analysis that was never run behind a clean waterfall chart. A pull request hides the untested edge case behind passing happy-path tests. A campaign brief hides an audience claim nobody opened the source for. Name the format's specific hiding place in the rubric, then look there first.

  • A deck or a memo. The hiding place is the working behind the chart. Ask for the source file rather than the export, pick one number, and trace it end to end. A candidate who cannot say where that number came from in one sentence did not put it there.
  • A pull request. The hiding place is the untested case, and the code that already existed. An analysis of 211 million changed lines found duplicated blocks rising from 8.3% to 12.3% of changes between 2021 and 2024, while refactored code fell from about 25% to under 10% 5. Read the diff for what was copied rather than reused, and for the test that covers only the path the assistant wrote.
  • A campaign or positioning brief. The hiding place is the audience claim nobody opened. Take the most on-message statistic in the document and ask for the source, the sample and the year.
  • A spreadsheet model. The hiding place is the assumption cell with no note on it. Ask which input the recommendation is most sensitive to, and what they did to find that out.
  • A written analysis or recommendation. The hiding place is the counter-argument that was never considered. Ask what would have to be true for the opposite recommendation to win.

The question underneath is identical in all five and only its location moves: which part of this did the candidate have to decide, and what would I be looking at if they had not decided it.

Score the Trail Before You Open the Artifact

Read the decision log first, rate it, then open the deliverable. Keep that order and enforce it. A strong artifact contaminates every judgment that follows it, and polish is precisely the property that now produces a strong first impression. Rate each item on the same anchored scale for every candidate, and put two readers on the first few submissions.

Federal assessment guidance is specific about what makes a work sample worth anything: trained assessors rate observed behavior or measured task outcomes against standard criteria, and the method fits roles where the competency is expected on arrival rather than trained afterward 6. Applied to a take-home, that is three requirements you can meet in an afternoon.

  • Write the key before the first submission arrives. Two or three anchors per item, describing what a weak, an adequate and a strong answer contains for this specific case. A key written after reading three submissions is a description of those three submissions.
  • Rate item by item, across candidates. Grade every candidate's assumptions, then every candidate's verification, then every candidate's rejections. Reading one full submission end to end is how a handsome artifact carries a weak decision through an entire rubric.
  • Double-read the first five. Two raters, independently, then compare. Disagreement in the first five is information about the key rather than about the candidates, and it is cheap to fix at that point.

One limit is worth stating plainly, because it decides whether this is enough for you. A decision log is written after the fact, by a candidate who knows it is being graded, with the same assistant still open. It is better evidence than the artifact, because it is harder to fake convincingly, and it can be checked against the deliverable, which the deliverable cannot be checked against anything. But it is an account of the work rather than a record of it. If you need the record, the exercise has to happen somewhere the decisions are visible while they are being made, and Take-Home or Live Session? works through that trade.

See what gets scored

Common questions

Should I ban AI on the take-home instead?

No, and a ban mostly selects for who ignored it. The candidate will use the tool in the job, so the honest exercise allows it and says so in the brief. Ban it only where the role genuinely forbids it, and then run the exercise somewhere the work is visible rather than trusting a rule you cannot check. A ban also costs you the signal: what a candidate hands to the assistant, and what they keep, is one of the more useful things a take-home can show.

How long should a take-home be now?

Under two hours, and shorter once the decision log carries the grading. Length was a proxy for effort and effort was a proxy for care; both proxies broke. A tight brief with a capped deliverable and a 300-word log produces more gradeable material than a six-hour build, and it costs candidates far less to attempt, which protects your completion rate at the same time.

What if the decision log is AI-written too?

Assume some of it is, and grade it anyway. The log is checkable against the artifact in a way the artifact is not checkable against anything. If the log claims an assumption the deliverable contradicts, or a source check the numbers do not reflect, that is a finding. Then ask two questions about it in the next conversation. Specificity that holds up under a follow-up is not something a candidate can borrow.

Do I have to tell candidates the log is graded?

Yes, in the brief, in one line. Undisclosed criteria produce noise: half the submissions treat the log as paperwork, and their ratings measure that guess rather than their judgment. Saying it is rated also redirects where candidates spend their two hours, which is the point: the time should go into decisions rather than into formatting a deck you have stopped grading.

Can two people grade the same submission consistently?

Only with anchors written in advance and item-by-item rating. Give each of the three items two or three worked descriptions of a weak, adequate and strong answer for this specific case, then have two raters score the first five submissions independently. Where they disagree, the key is ambiguous rather than the candidates. Fix it before the sixth submission, and record what you changed.

References

  1. 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers on 246 real issues from their own repositories, randomized to allow or forbid AI tools, took 19% longer with the tools; they forecast a 24% speedup and still believed afterwards that the tools had sped them up by 20%.
  2. 2. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The single biggest reported frustration with AI tools, at 66%, is "AI solutions that are almost right, but not quite"; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  3. 3. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ACM CHI Conference on Human Factors in Computing Systems (CHI '25), 2025. advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in the tool is associated with less critical thinking, higher confidence in one's own task ability with more, and the remaining effort shifts toward information verification, response integration and task stewardship.
  4. 4. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Work samples and other employment tests are selection procedures that must be job-related and consistent with business necessity, with the employer carrying that burden.
  5. 5. AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones GitClear, 2025. gitclear.com Across 211 million changed lines, duplicated code rose from 8.3% to 12.3% of changes between 2021 and 2024 while refactored code fell from about 25% to under 10%.
  6. 6. Assessment and Selection: Work Samples and Simulations U.S. Office of Personnel Management, 2024. opm.gov Work sample ratings come from trained assessors observing behavior or measuring task outcomes against standard criteria, and the method suits roles where the competency is expected at entry rather than trained after selection.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.