Assessment design

What Should a Work Sample Test Now That AI Can Produce It?

Now that AI can produce the deliverable, a work sample should grade the two acts a model cannot do for the candidate: how they stated the problem and its constraints before generating, and what they tested the answer against outside the conversation. Both are field-specific. Keep the occupational task, cut polish from the rubric. Two limits. An unwatched framing note can be written afterwards, so spend fifteen minutes on those decisions. And a task an assistant finishes correctly in a minute is not a work sample.

The takeMy read is that most versions of this redesign break in the same place. The rubric gets rewritten in an afternoon; the packet of source material the candidate works from never does. With nothing buried in the packet to find, the verification half of the rubric scores every candidate the same way, and the reviewer concludes the redesign failed. The authored error is the instrument. It is also the line item nobody funds, because it looks like content and bills like research. A team unwilling to write its own wrong number ends up where it started, reading polish and calling it judgment.

Where Olive fits

Open a role and see what the work shows

If you build this yourself, the task is the cheap half: the answer key, the second reader and the evidence trail are what take the afternoons. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session, and the candidate is granted the same document.

Rank your shortlist

Why did the deliverable stop separating candidates?

Because assistance compresses the range. In a preregistered experiment with 444 college-educated professionals on occupation-specific writing tasks, ChatGPT cut time by 0.8 standard deviations and raised graded quality by 0.4, and inequality between workers fell, because the tool helped the lower-scoring writers most 1. A rubric that scores the artifact is measuring a distribution that has been squeezed flat at the top.

One more finding from that experiment matters here: the tool "mostly substitutes for worker effort rather than complementing worker skills" 1. That is the sentence to sit with, because it says the polish in front of you is not evidence about the person who submitted it. Every stack of take-homes now arrives structured, sectioned, and free of the errors that used to sort them.

Compression does not turn the work sample into the wrong instrument. It makes your version of it out of date, which the method has always been vulnerable to. The federal guidance on work samples says as much in its own cost note: a work sample "may require periodic updating," and the example given is a task built around a typewriter at an organization that has since automated its documents 3. The task drifts away from the job, and the score keeps arriving as if nothing moved.

Whether to keep running one was never the live question. It is which part of the task still costs the candidate something. If you are still grading take-homes that all come back polished, you already have the answer in front of you: the spread you used to see is gone, and nothing in the deliverable will bring it back.

What are the two acts AI cannot hand over?

Framing and the outside check. A model does everything between them (drafting, structuring, restating, formatting) well enough that grading it tells you about the model. What it cannot do is decide what the question is in a domain where it carries no stake, and it cannot open the thing outside the conversation that settles a claim. Both are acts, and acts leave a record.

The check matters because fluency and accuracy come apart, and they come apart inside the tools built for the job. A preregistered evaluation of purpose-built legal research systems found that Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time 6. Nothing in the output announces which third that is. Somebody has to go outside the conversation.

And the candidate's own account of doing so is not evidence. METR ran a randomized trial with 16 experienced open-source developers on 246 real issues in repositories they already knew; with AI tools allowed, the work took 19% longer, and afterwards the developers still believed the tools had sped them up by 20% 4. If people inside the work cannot feel a 39-point swing, a self-assessment paragraph at the end of a take-home is worth nothing on its own.

The part AI cannot hand over is field-specific, which is why a generic "AI skills" exercise measures so little:

  • Management consulting. Framing the question. Which decision the sizing is actually for, and what would make the recommendation wrong.
  • Financial analysis. Choosing the comparable. Which set, and the reason it is these companies and not the three the model reached for first.
  • Software engineering. Picking the failure case. The input that breaks it, which the existing tests do not cover.
  • Marketing. Choosing which number to stand on. The most quotable statistic in the folder is usually a vendor's own small customer survey.
  • Legal operations. Deciding which document controls when the playbook, the system record and the signed precedent disagree.
  • Revenue cycle. Deciding which denial to concede, when the assistant will draft a persuasive appeal for the one that is simply late.

Name yours before you write a rubric. This is not a technical-roles-only problem, and screening for AI judgment in a finance or marketing role starts with the same sentence: what does someone competent here re-check before they sign anything?

Show the before-and-after of one real task

One task, two versions. Before: size the US market for a connected-fitness subscription and recommend an entry approach, five-page deck, three days. After: the same packet, the same deck, plus two artifacts submitted with it: a one-page framing note written before anything is generated, and a verification log listing the claims that were opened and what each one changed. The task is identical. The rubric is not.

The framing note asks four things, in the candidate's own words and before the first prompt: what is being decided, what would make the answer wrong, which single number the recommendation rests on, and which constraint binds. It is one page because it has to be written before the work, and a page is what someone will actually produce at that point.

The verification log asks three: what claim was opened, what it was checked against, and what changed as a result. "Checked the market figure" is not a row. "Opened the source behind the summary page, found the 2019 segment figure attributed to the whole category, re-based the estimate, TAM down 31%" is a row.

That one is not hypothetical, and it is the point of authoring the packet rather than buying one. Write the source material so it settles nothing until somebody opens it: the summary page carries a figure the underlying document does not support, and the assistant will build on the summary because the summary is what it was handed.

Graded beforeGraded now
Structure of the deckWhether the framing note named the load-bearing number before anything was generated
Defensibility of the estimateWhich claim was opened, and whether it was the one the recommendation rested on
Quality of the recommendationWhat moved after the check: a figure, a recommendation, or a stated limit
Clarity of the writingNot scored

The deck still gets read, because a candidate who framed well and checked the right number and then shipped something incoherent has not done the job. It just stops being the column that decides anything.

Write the rubric around observable acts

Score acts, not adjectives, and keep the acts inside the job's own content. The Uniform Guidelines are blunt about why: a selection procedure holds up on content validity "to the extent that it is a representative sample of the content of the job," and one "based upon inferences about mental processes cannot be supported solely or primarily on the basis of content validity" 2. Judgment is named there among the constructs a content strategy cannot carry.

So "this exercise tests judgment" is the wrong sentence to write down, and it is the one most AI-era rubrics open with. "The candidate re-derived the segment figure against the underlying source" is a sampled piece of job content. The first is a claim about someone's mind; the second is a thing that either is in the record or is not.

Four columns, filled in from the record rather than from impressions: named (did the problem or the doubtful claim appear as a problem anywhere), timing (before or after the deliverable was drafted), instrument (what was opened, run or recomputed), and consequence (what in the final answer is different because of it). Consequence is the column that separates a candidate with an instinct from one with a method.

Two reviewers should fill those in independently and agree. If they cannot, the rubric is the problem rather than the reviewers, and whether two people can score AI use the same way is a measurable property of the wording, testable on five old submissions before a candidate ever sees it. Anchor every row with a real excerpt from those five.

Three things stay off the sheet. How much AI was used: a candidate who judged the model was the wrong instrument for a step and did it by hand has demonstrated the exact behavior being tested. Prompt syntax, which is tool trivia with a six-month shelf life. And speed, unless speed is genuinely what the role is bought for. The shift underneath all three is hiring for verification rather than production, and none of them point at it.

What does this still not tell you?

Less than the rubric implies. Work-sample validity is lower than the figure most people quote: a 2022 re-analysis correcting systematic overcorrection for range restriction cut mean validity estimates across selection procedures by .10 to .20 points, and structured interviews came out on top of the revised table 5. One redesigned task is one observation of one occupation on one afternoon.

The honest limits, in the order you will hit them. Authoring costs real time (the guidance calls work samples costly to develop and says they need periodic updating 3), and the packet with the buried error is the expensive part, not the brief. It leaks, so budget a second case before the first one has been sent thirty times. And an unwatched take-home gives you the artifacts you asked for and no record of how they were made, which is the trade in a take-home against a live working session.

That last one is the load-bearing gap. A framing note written before the work and a framing note written afterwards look the same on the page, and the candidate who reverse-engineers one from a finished deck will produce a better document than the candidate who wrote a rough one honestly at the start. Nothing in an emailed submission distinguishes them. Asking for a timestamp is theater; asking for a screen recording of a three-day task is worse.

The practical fix is smaller than either: shorten the task until it fits a session someone can watch or record, or pair the take-home with fifteen minutes on the specific decisions in the framing note. Ask what the candidate rejected and why. Whether take-homes tell you anything at all once AI is in play turns almost entirely on whether some record of the work exists alongside the output. And a panel that does not use these tools daily will still read the polish as skill, which is a separate problem about the people doing the grading rather than about the rubric.

See what gets scored

Common questions

Should candidates be allowed to use AI on the work sample?

Allow it, and put that in the brief. A ban measures a version of the job that no longer exists and cannot be enforced on an unwatched task anyway, so all it produces is a quiet split between candidates who followed the instruction and candidates who did not. Allowing it openly also makes the framing note and verification log legible rather than incriminating: you are asking how the tool was used, not whether. State which tools are fine, state that the process artifacts are part of the deliverable, and grade what comes back.

How long should the redesigned task be?

Under ninety minutes for an unpaid task, and pay for anything longer. Adding a framing note and a verification log adds maybe fifteen minutes, so cut scope somewhere else rather than stacking the requirements on an existing three-hour assignment. Unpaid length filters on free time instead of skill, and the candidates it removes first are the ones with a current job or caregiving duties. If the task genuinely needs a day, make it a paid trial and say the rate in the invitation.

What stops a candidate writing the framing note after the fact?

Nothing, on an emailed take-home, and it is the main weakness of the design. A note reverse-engineered from a finished deck reads better than an honest rough one written at the start, and no timestamp you can ask for settles it. Two things help. Shorten the task until it fits a recorded or watched session. Or keep the take-home and add fifteen minutes on the specific decisions in the note: what was rejected, what would have changed the recommendation, which source was opened first. Reconstructed reasoning thins out fast under a follow-up question.

Does this work for roles with no technical artifact?

Yes, because every occupation has material that reads settled until someone opens it. For a marketer it is the most quotable statistic in the research folder, sourced from a vendor's own small customer survey. For a paralegal it is a pin cite to a paragraph that does not say what the memo claims. For a recruiter it is a compensation benchmark with the wrong geography attached. If you cannot name yours, ask the two strongest people on the team what they re-check before they sign anything, and build the packet around that.

Do you need a new task, or just a new rubric?

Usually the rubric plus one edit to the source material. Keep the task: it is the part with content validity, and rewriting it costs an afternoon you do not need to spend. Add the two process artifacts, move the grading weight off the deliverable, and plant one field-plausible error in the packet so the verification column has something to catch. The exception is a task an assistant finishes correctly in under a minute. That one is not a work sample any more and no rubric rescues it.

References

  1. 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence Noy and Zhang, MIT (working paper; later published in Science), 2023. economics.mit.edu 444 college-educated professionals on occupation-specific writing tasks: time down 0.8 SD, quality up 0.4 SD, inequality between workers decreases because the tool benefits lower-ability workers most; it mostly substitutes for worker effort rather than complementing worker skills.
  2. 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Content validity holds to the extent the procedure is a representative sample of the content of the job; a procedure based upon inferences about mental processes cannot be supported solely or primarily on content validity, and judgment is named among the constructs a content strategy cannot carry.
  3. 3. Assessment and Selection: Work Samples and Simulations U.S. Office of Personnel Management, 2024. opm.gov Work samples derive validity from tasks being very representative of tasks performed on the job; they may be costly to develop and may require periodic updating, with a technology change (the typewriter task at a now-automated organization) given as the example.
  4. 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers on 246 real issues took 19% longer with AI tools allowed, and still believed afterwards that the tools had sped them up by 20%.
  5. 5. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range Sackett, Zhang, Berry and Lievens, Journal of Applied Psychology, 2022. europepmc.org Correcting systematic overcorrection for range restriction reduced mean validity estimates by .10-.20 points across selection procedures, with structured interviews emerging as the top-ranked procedure.
  6. 6. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools Magesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford RegLab and HAI (arXiv), 2024. arxiv.org Purpose-built legal research tools from LexisNexis and Thomson Reuters each hallucinate between 17% and 33% of the time in a preregistered evaluation.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.