Assessment design

Work Sample or AI-Collaboration Exercise: Which Predicts Performance?

Neither the classic work sample nor the AI-collaboration exercise predicts a first-year hire's performance better in general: the sample predicts while the hire does the first pass unaided, the exercise once an assistant drafts that pass first. Audit last quarter's entry-level tasks: if an assistant drafts under about a third, keep the classic sample; over about two thirds, run the collaboration exercise; in between, run one task both ways and compare the records. Neither carries a validity figure measured against real first-year performance, so run either beside a structured round.

The takeThe format argument hides the thing that actually matters: your assessment now has a shelf life. Resemblance is the entire source of a work sample's predictive power, and the work it resembles is being rewritten underneath it every few quarters. So the audit is a standing cost, not a setup step, and my read is that most teams will not pay it twice. They will build one exercise, call it validated, and still be running it three years after the first-year job stopped looking like it. If the pattern holds, a work sample nobody re-audits is a validity claim about a job that no longer exists.

Where Olive fits

Open a role and see what the work shows

If the audit says the collaboration exercise is your work sample, the expensive parts are the authored case and the evidence behind each judgment, per occupation. Olive ships twelve authored cases per occupation, each grounded in one SOC code, and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification), each anchored to a moment in the session and given to the candidate as well.

Rank your shortlist

Which format predicts a first-year hire better?

Whichever one still resembles the work. Predictive validity is not a property of a format; it is a property of the match between what the exercise samples and what the person will be paid to do in month one. Ask which parts of your entry-level job an assistant now does first, and the format question mostly answers itself. The research supports that framing rather than a winner.

Start by dropping the ranking you probably carry. Sackett, Zhang, Berry and Lievens re-derived the meta-analytic validity estimates that hiring teams have quoted for decades and found the standard range-restriction corrections had substantially overstated them; the revised figures fall by roughly .10 to .20 points across widely used procedures, and structured interviews, not work samples, come out top-ranked 1. Work samples remain useful. They stopped being the settled favourite a while ago, and no coefficient in that table was measured on a job where a model writes the first draft.

The federal selection guidelines make the same point in operational language. A procedure can be supported by content validity "to the extent that it is a representative sample of the content of the job," and the closer its content and context sit to real work behaviors, the stronger that basis; as the content less resembles a work behavior, the procedure is less likely to be content valid 2. That is a sliding scale rather than a badge, and an assistant moving into the middle of a job slides the role along it.

So the honest form of the question is not which format is better. It is how much of your entry-level work the model has already absorbed, and what a work sample has to change to survive that is a design problem before it is a procurement one.

When does a classic work sample still predict?

While a person still does the first pass unaided. If your first-year hire will open a file, a case or a chart and produce something before any model touches it, the classic sample is a direct sample of that act and predicts it about as well as it ever did. The test is not whether AI exists in your industry. It is whether it has reached the specific tasks a first-year is handed.

There is now occupation-level evidence for that split. Brynjolfsson, Chandar and Chen, using ADP payroll records covering millions of US workers through June 2026, find employment for 22-to-25-year-olds in AI-exposed occupations sitting 19% below where it would be had it kept pace with less-exposed peers, with no comparable gap for experienced workers, and the declines concentrated in occupations where AI usage primarily substitutes for human tasks, while employment is flat or rising where usage primarily complements the worker 3. Substitute or complement is the same line your assessment choice sits on.

In practice the classic sample survives where generation was never the bottleneck: short deliverables, judgments taken in front of a customer, work whose inputs live in a system or a physical setting a model does not read. Keep it there and change nothing about it. Running an AI-collaboration exercise for a job with no AI in it measures enthusiasm for AI, and you will hire for exactly that.

In an AI-era funnel the classic format carries a caveat: unless the task is done in front of you, you are no longer sampling unaided work, you are sampling whatever the candidate chose to do at home. That is a different exercise with a different meaning, and it is worth deciding which one you meant.

When does an AI-collaboration exercise predict better?

Once the model does the first pass, because then the residual job is the thing worth sampling. What a first-year is paid for is framing the problem, demanding evidence for the claim that matters, deciding what to hand over and what to keep, refusing output on substance, and testing a claim against something outside the conversation. A finished artifact reports none of that, because a good one and a lucky one look identical.

Delegation is the behavior that moves most, and it is measurable. Ju and Aral randomly assigned 2,234 participants to human-human and human-AI teams producing 11,024 ads, and found participants delegated 17% more work to an AI agent than to a human partner and made 62% fewer direct text edits when working with AI; the teams also split by task, with human-AI pairs producing higher text quality and human-human pairs higher image quality 4. Collaboration skill is not one trait. It shows up differently depending on which part of the task the model happens to be good at, which is why the exercise has to be built on your occupation's material rather than on a generic prompt.

The failure such an exercise must provoke is a quiet one. In Stack Overflow's 2025 developer survey the biggest reported frustration with AI tools was "AI solutions that are almost right, but not quite" at 66%, with time-consuming debugging of AI-generated code second at 45%, and more developers distrusting the accuracy of AI output (46%) than trusting it (33%) 6. Almost-right survives a read-through and fails downstream. If nothing in your packet is almost-right, everybody passes and you have measured nothing.

Grade the acts, not the deliverable: which claim was checked, against what, when, and what changed in the answer because of it. Hiring for verification rather than production is the same rubric written for early-career work. See how Olive measures this.

How do you tell which one your role needs?

Audit last quarter instead of arguing about it. Pull ten tasks a first-year actually did, and mark each for whether an assistant now produces a usable first draft of it. Under about a third, write the classic work sample. Over about two thirds, the collaboration exercise is your work sample. In between, run one task both ways on two candidates and read the two records against each other.

Do the audit on tasks rather than on job titles. Occupational exposure is an average over a whole job, and the entry-level slice is rarely the average: a junior analyst's quarter is weighted toward the parts most exposed to a model, while the senior version of the same title is weighted toward the parts that are not. Rate what a first-year will be assigned in their first ninety days, which is also the honest input to deciding what AI actually does inside the role.

Whichever way the audit lands, write one rubric and keep it across the swap. The columns (which claim was checked, against what instrument, before or after the deliverable was drafted, and what moved as a result) transfer between the two formats unedited. The task never transfers, and neither does the answer key, which is where the real authoring cost sits.

Then calibrate before it decides anything. Two people already doing the job should score the same session separately and agree, and one of them should find the planted problem inside ten minutes without calling it a trick. An exercise that two colleagues score differently is not measuring the candidate, and what an assessment score can honestly predict about performance starts with two graders agreeing on what they saw.

What does neither format predict?

Neither predicts first-year performance the way a validity coefficient implies. Both sample one person on one afternoon against one case, scored by people who know they are scoring. Neither reaches whether the habit holds in month four with a real deadline and nobody watching. And no coefficient exists yet for an AI-collaboration exercise against first-year performance, because a criterion study needs performance data collected after the hire.

That gap has a practical consequence: when a vendor quotes a validity figure for an AI-collaboration product, ask which criterion it was measured against, over what interval, and on how many hires. A number borrowed from the older work-sample literature is a number about a different exercise.

The shortcut of simply asking candidates is also closed. METR's early-2025 randomized trial found experienced open-source developers took 19% longer on real tasks with AI tools allowed, against their own expectation of a speedup, and its February 2026 update reports that self-reported speedups "can be quite unreliable" 5. If people inside the work cannot tell whether AI made them faster, an interview answer about how they work with AI is a story rather than evidence.

The guidelines put the same limit in legal terms: a selection procedure resting on inferences about mental processes cannot be supported solely or primarily on the basis of content validity 2. Asking someone how they would check a confident claim is that kind of inference. Handing them the claim, the source and forty minutes is a sample. Either way the result is one input beside a structured round rather than a replacement for one. The older comparison between a work sample and a structured interview never resolved in favour of one format, and this one will not either.

See what gets scored

Common questions

Can one exercise be both a work sample and an AI-collaboration exercise?

Yes, and for most roles that is the cheaper build. Take the occupational task you would have set as a work sample, put an assistant beside it that will do the whole thing if nobody stops it, and grade both the deliverable and the record of how it was made. You get the classic content-validity argument from the task and the collaboration signal from the record. What you cannot do is set a generic prompting quiz and call it a work sample; it samples no job.

Should candidates be allowed to use AI on a classic work sample?

Decide it explicitly and say so in the invite, because an unstated rule produces two populations whose results are not comparable. If the job is done unaided, say tools are not permitted and run the task live, since an unwatched ban is unenforceable. If the job is done with an assistant, allow it and change what you grade: the deliverable stops carrying signal on its own, and the record of how it was made starts carrying it instead.

How long should either exercise run?

Long enough that triage has a cost, short enough that completion holds. Forty to sixty minutes, with a task slightly too large for the window, forces a real choice about what to check and what to hand over, which is the thing being measured. Past about two hours completion drops and you select for candidates with free afternoons rather than for judgment. State the cap in the invite, hold to it, and pay for anything longer.

Does an AI-collaboration exercise disadvantage a candidate who rarely uses AI?

It disadvantages one who cannot judge when the model is wrong, which is the point, and it does not disadvantage one who used it sparingly. Volume of use is not the signal. Someone who decided the assistant was the wrong instrument for a step and did that step by hand has demonstrated exactly the judgment being sampled. Grade the reasoning behind the split rather than the count of prompts, and the tool-fluency gap mostly disappears inside the first ten minutes.

What about grading first-year performance itself?

Fix that before you compare formats, because a noisy criterion makes every predictor look weak. Most first-year ratings are a manager's global impression collected once, which correlates with visibility and confidence as much as with output. Write down two or three things a first-year is expected to be doing unsupervised by month six, rate those separately, and collect them on a date rather than when someone remembers. Then a comparison between assessment formats has something honest to be measured against.

Which format should a small team pick if it can only run one?

Run the one that samples the task you are most likely to get wrong at offer time. For most entry-level knowledge roles that is now the collaboration exercise, because the classic take-home no longer separates candidates: polished output is available to everyone. For roles where the first-year deliverable is produced without a model, keep the classic sample and put the time saved into a structured round instead of adding a second exercise.

References

  1. 1. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range Sackett, Zhang, Berry & Lievens, Journal of Applied Psychology (via Europe PMC), 2022. europepmc.org Standard range-restriction corrections substantially overestimated the validity of many selection procedures; revised estimates fall by .10 to .20 points and structured interviews emerge as the top-ranked procedure.
  2. 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Content validity holds to the extent the procedure is a representative sample of job content; the closer its content and context are to work behaviors the stronger the basis, and a procedure resting on inferences about mental processes cannot be supported solely or primarily on content validity.
  3. 3. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence Stanford Digital Economy Lab (Brynjolfsson, Chandar, Chen), revised August 2026, 2026. digitaleconomy.stanford.edu ADP payroll data through June 2026: employment for 22-to-25-year-olds in AI-exposed occupations sits 19% below its counterfactual, with declines concentrated where AI usage substitutes for human tasks and employment flat or rising where it complements workers.
  4. 4. Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance Ju & Aral, MIT (arXiv 2503.18238), 2025. arxiv.org 2,234 participants in human-human and human-AI teams produced 11,024 ads; participants delegated 17% more work to AI agents than to human partners and made 62% fewer direct text edits, and team type split by task with human-AI pairs higher on text quality and human-human pairs higher on image quality.
  5. 5. We are Changing our Developer Productivity Experiment Design METR, 2026. metr.org Restates the early-2025 randomized result that AI tools made experienced developers take 19% longer, and reports that developers' self-reported speedup estimates can be quite unreliable.
  6. 6. Stack Overflow Developer Survey 2025: AI Stack Overflow, 2025. survey.stackoverflow.co The biggest reported AI frustration is "AI solutions that are almost right, but not quite" at 66%, debugging AI-generated code second at 45%, and 46% of developers distrust the accuracy of AI output against 33% who trust it.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.