Assessment design
An AI Work Sample Tests Whether You Catch the Bad Answer
An AI-collaboration exercise is a timed, job-shaped task done with an AI assistant, and the material often contains something the assistant gets wrong. Find that instead of polishing the output. What gets scored is whether you notice and correct the error, not how fast or clean the final version looks. Show your reasoning wherever the format allows: what you checked, what you rejected, and what you would still verify given more time. Speed is rarely the axis scored; an unchecked confident answer often is, and negatively.
The takeThis format is more forgiving than it looks from the invitation email. You are not being asked to outperform a machine at generating text; you are being asked to do the one thing the machine still cannot reliably do for itself: notice when its own output is wrong. Treat every polished paragraph the assistant hands you as a claim to test rather than a draft to accept, and the exercise gets considerably less mysterious.
Where Olive fits
Open a role and see what the work shows
If the assessment you're sent is Olive, using the AI assistant is the point: the assignment runs openly with one, and a person writes what you framed, what you delegated and what you verified, in words rather than a score. The candidate is granted that identical report, free, on every tier.
Rank your shortlistWhat Is an AI Work Sample, Actually?
An AI-collaboration exercise gives you a real task, an assistant that will do most of the drafting if you let it, and a time box. It differs from a plain work sample less in what you produce than in what gets watched while you produce it: which parts of the assistant's output you accepted, which you rejected, and why. The deliverable matters less than the trail behind it.
Work samples earned their reputation honestly, if not for the exact number usually quoted at you. The widely repeated claim that they predict job performance at .54 is stale; a 2022 re-analysis of the meta-analytic literature puts work sample validity at .33, a figure the original authors themselves had already revised downward once before 1. An AI-collaboration exercise is a newer variant on the same idea, built for a task that now assumes an assistant is present rather than absent.
The lineage stops short of the exercise in your invitation, though, and it is worth knowing where. That .33 comes from studies that almost all tested people already doing the job, and no comparable figure has been measured for an AI-collaboration exercise against real first-year performance, because that would need performance data collected after the hire. A number borrowed from the older work-sample literature describes a different exercise. What carries over is the design idea, that a job-shaped task tells you more than a conversation about the job. The coefficient does not carry over, and nobody should quote you one.
Expect the Prompt to Contain a Trap
Assume the material you're handed contains something wrong, because a well-built version of this exercise is designed that way. Designers who build these seed a bad number, a shaky claim, or a plausible-sounding conclusion that does not hold up, then watch who finds it. An exercise where nothing is almost-right measures nothing, since everybody passes it. Spending your first minutes hunting for that flaw beats polishing prose the assistant already wrote acceptably well.
The instinct to trust a fluent answer is documented and not a personal failing. In a field experiment with 758 Boston Consulting Group consultants, those given GPT-4 on the one task chosen to sit outside the tool's range were 19 percentage points less likely to reach the correct answer than the control group working without it 3. That result does not show the tool making people worse in general. It shows that nobody in the room could tell which side of the line the task sat on, which is the gap an exercise like yours is built to look for.
Read the material once for content, then read it again looking specifically for the sentence that sounds most confident, since a confidently stated wrong answer is easier to skim past than a hedged one. The flaw is rarely hidden in the hardest part of the task. It is usually sitting in plain sight, in the part that reads so cleanly nobody thinks to check it.
Show Your Reasoning, Not Just the Output
Where the format allows notes, a chat transcript, or a short write-up, use it to show three things: what you checked, what you rejected, and what you would still verify with more time. A rubric built around this axis is looking for visible judgment, and a clean final answer with no trace of that process reads as luck rather than skill, even when the answer happens to be correct.
Employers deciding whether a work sample or an AI-collaboration exercise predicts a first-year hire's performance better are told to look at how much of the real job an assistant now drafts first, which is the same shift this exercise is built around: the task assumed a person would produce the first pass, and now it assumes the assistant does, so the thing worth measuring moved with it.
A useful shape for that trail, when the instructions don't dictate one: name the claim you decided mattered most, say what you checked it against, and say what you did with the answer. Three lines will do it. That is the difference between a grader inferring your judgment from a finished artifact and reading it directly, and a good artifact and a lucky one look identical on the page.
Does Speed Actually Matter Here?
Rarely, and self-reported speed is an especially bad guide to trust. In a randomized trial, experienced developers using AI tools on code they already knew well finished 19% slower than they would have unaided, despite forecasting a 24% speedup beforehand and still believing afterward that they had gained about 20% 4. If professionals tracking their own output got the direction wrong, a candidate guessing at how fast to move is guessing blind too.
Employers running these exercises are advised to measure how well someone judges the output rather than how fast they produce it, letting the clock cap the exercise instead of scoring it directly. The most quoted developer speed number in circulation, a 55.8% reduction in completion time with an AI coding tool, comes from a single trial where 95 freelance developers recruited on Upwork were asked to write one HTTP server from scratch, and it describes only the ones who finished 2. Nothing about a real assessment task resembles that setup closely enough to borrow the number, and nothing about your invitation email is asking you to beat it.
Don't Polish Past the Point of Being Checkable
A submission that looks finished can still fail if there is nothing in it an evaluator can verify. Leave the seams visible: a comment noting what you didn't have time to check, a flagged assumption, a line explaining why you rejected the assistant's first draft. All of that reads as judgment. Smoothing it away in the name of a clean deliverable removes the exact evidence the rubric is looking for.
A submission with two visible corrections and an honest open question usually reads stronger than one that looks effortless, because effortless is precisely what a rubric built around this axis has learned to distrust.
If the invitation looks more like an interview than a written assignment, what an AI interview actually scores covers the adjacent format, since the two are usually graded on close to the same axis.
Common questions
Is there really always a mistake hidden in the material?
Not guaranteed, but common enough that assuming one exists is a better strategy than assuming the material is clean. Well-built exercises are frequently designed so the assistant's first answer is wrong somewhere, precisely so the exercise can tell who catches it.
Should I avoid using the AI assistant as much as possible?
No. Refusing to use it usually reads as avoiding the point of the exercise. The task is to use it and catch what it gets wrong, not to prove you can do the work without it.
What if I don't find anything wrong in the time given?
Say so, and show what you checked along the way. A candidate who visibly verified several things and found nothing off is in a stronger position than one who submitted a clean answer with no trace of having checked at all.
Does finishing early count in my favor?
Rarely by itself. Finishing early with visible checking is a fine outcome; finishing early because you accepted the first output without reading it closely usually reads as the opposite of what the exercise is measuring.
How is this different from a plain take-home assignment?
A plain take-home mostly scores the finished deliverable. An AI-collaboration exercise scores the deliverable plus the visible judgment behind it: what you accepted from the assistant, what you rejected, and why, which is why the process trail matters as much as the answer.
References
- 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Corrects the widely quoted .54 work sample validity figure to .33, and its own basis (53 of 54 studies concurrent) is what supports the article saying the coefficient does not transfer to an AI-collaboration exercise.
- 2. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (arXiv:2302.06590) arxiv.org Sources the famous 55.8% developer speed figure to one toy task, supporting that a real assessment task should not be benchmarked against it.
- 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports that a confident wrong answer, not a slow one, is the documented failure mode an AI-collaboration exercise is built to catch.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) arxiv.org Supports that self-reported speed with AI tools is unreliable, so speed is a poor thing for a candidate to optimize for in the exercise.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.