Assessment design
Should You Run a Paid Trial Instead of an Assessment?
An assessment usually beats a paid trial. The trial earns its cost on three conditions at once: a week of the real work is genuinely reachable, your finalist can take that week, and you are down to one or two people. Miss any of them and an assessment wins. The trial spends senior hours on the brief, the access and the read, spends far more of the candidate's, makes them your employee for the days you pay them, and quietly loses the employed finalists who cannot clear a week.
The takeThe trial's pull is that it feels like proof, and the feeling does most of the work. A week in your building tells you what a week in your building is like: your context, your tools, your people answering questions. Change any of that and the observation goes with it. No study I can point to has five paid days predicting a year of work better than one task written carefully and graded the same way twice. What the money really buys is a decision you can defend to yourself, and that is worth paying for at the end of a loop and nowhere earlier.
Where Olive fits
Open a role and see what the work shows
A trial reaches the real job for the roles whose work fits in a week, and reaches nothing for the rest of the pipeline. Olive covers that remainder: a 40-to-60-minute occupational assignment with an AI assistant available, returned as six separately-evidenced findings written by a human reviewer (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), with the candidate granted the identical document.
Rank your shortlistWhat Does a Paid Trial Actually Cost?
More than the day rate, and the day rate is the only line most founders price. A trial costs senior hours on your side: writing a brief that is real work but not confidential work, provisioning access, answering questions during the week, then reading the output carefully enough to be fair. It costs the candidate a working week. And it costs you the finalists who cannot clear one.
Price these at what your senior people cost per hour, not at the candidate's day rate.
- The brief. A trial task has to be real work that is not confidential work, sized to the days you are paying for, and identical for everyone at that stage. Someone senior writes it once and maintains it. Different tasks for different candidates leaves you with two selection procedures and no basis for comparing what they produced 2.
- Access. Whatever the task needs (a sandbox, a data extract, a repo, a seat in a paid tool), provisioned for a non-employee and revoked afterwards. This is where trials for production-facing and regulated roles quietly die.
- Supervision. Somebody answers questions all week. If nobody does, you are measuring how well a stranger guesses at your context, and the winner is whoever guessed closest to your house style.
- The read. A week of output takes longer to grade than a two-hour submission, and it is your most expensive person doing the grading. How to Grade a Take-Home When Every Submission Is Polished applies here without modification.
- Multiplication. Four finalists at five paid days each is twenty days of pay, four briefs to support, and four long reads.
One thing the spend does not buy is a reliable reading of speed. In METR's randomized study, experienced developers working in repositories they already knew took 19% longer with AI tools allowed, and afterwards still believed the tools had sped them up by 20% 4. A week of watching output is not a throughput measurement, and the candidate's own account of the week is not one either.
Who Declines a Paid Trial?
The employed ones. A five-day trial asks a working candidate to burn a week of leave, explain the absence to a current manager, and gamble the week on your decision. The people who can say yes are between jobs, on notice, or supported enough to absorb the risk. That is a filter on availability rather than on judgment, and it fires at the end of your loop, where losing someone costs the most.
The workarounds each have a cost of their own.
- Evenings and weekends. This converts a paid trial into unpaid overtime stacked on top of a full-time job, and selects hard against anyone with caregiving duties. It also produces tired work you will grade as if it were their best.
- Shortening it to two days. Reasonable, and it moves you back toward a work sample. At that point compare it honestly against the cheaper formats in Take-Home or Live Working Session: Which Shows AI Judgment? rather than calling it a trial.
- Running it for some finalists only. Now one decision rests on two different procedures, applied to different people. A selection procedure is expected to be job-related and consistent with business necessity, and that is a much harder case to make about a step some candidates never faced 2.
- A competing offer. A finalist with a live offer on a normal timeline will take it. Your trial adds a week to the exact stage where speed decides who accepts.
One discipline makes the decline rate survivable: write the rejection before you send the invitation. If you cannot say plainly how someone who does five paid days and does not get the job will be told, and what they keep, the offer is not ready to make.
What Does Paying Someone for a Week Make Them?
An employee for that week, in every way that matters to your finance team. The Fair Labor Standards Act defines an employee as any individual employed by an employer, and defines employ as to suffer or permit to work 1. A person doing your work, on your systems, on your direction, for money, sits inside that definition. Calling the arrangement a trial, a project or a pilot does not move it.
Four consequences to settle before day one rather than after.
- Classification. Treat them as an employee for the period unless your counsel has told you otherwise in writing about this specific arrangement. A day rate paid through accounts payable is a labelling choice, not a classification.
- The unpaid version is worse, not cheaper. Permitting someone to work is the statutory definition of employing them 1. An unpaid audition does not escape the wage question; it just removes your only defence for the hours.
- The trial is a selection procedure. The EEOC lists work samples among employment tests and selection procedures, and where a procedure has disparate impact the employer has to show it is job-related and consistent with business necessity 2. A task improvised for one candidate on a Tuesday is difficult to defend on either count.
- Confidentiality, IP and revocation. What they see, what they build, who owns it, and when access ends, all of it in writing, before the first day, for someone you may never hire.
Nothing above argues against running a trial. It makes them an employment decision you are taking twice: once for the week, once for the job. Budget the second one's paperwork into the first.
When Does a Trial Genuinely Beat an Assessment?
When the real work is reachable inside the days you are paying for, the candidate can take those days, and you are down to one or two people. Under those three conditions a trial gives you something no shorter format does, because it is not a sample of the job. It is the job, in the place the job happens, with the people it happens with.
That is the strongest content-validity argument available. Content validity holds to the extent a procedure is a representative sample of the content of the job 3, and nothing samples job content as directly as the job.
The honest qualifier: realism is not the same as measured validity. Sackett and colleagues re-ran the meta-analytic estimates behind the validity table most people carry in their heads and found the standard correction had substantially overcorrected: most procedures lost .10 to .20 in mean validity, and structured interviews came out top-ranked 5. A trial is one unstandardized observation of one person on one task in one week. Its realism does not buy the consistency that makes a procedure predict anything.
So when you run one, keep five things fixed.
- One task for everyone at that stage, written before the first candidate sees it.
- A rubric written before the first trial, with two or three worked descriptions per item. A key written after reading three submissions describes those three submissions.
- Market rate, paid on your normal terms, not a token honorarium for a week of output you intend to read seriously.
- Scoped to need no access you would withhold from a stranger. If the task only works with production credentials, the task is wrong or the role cannot be trialled.
- A hard end and a written decision date. An open-ended trial is contract-to-hire, which is a fine thing to run and a different thing to call it.
If you are considering a trial because your current round produces nothing you trust, the cheaper first move is to fix the round and check the fix. How to Pilot an Assessment Before Making It a Gate covers running a new step beside the live loop before it decides anything.
Which Roles Can a Trial Reach, and Which Can't?
Cycle time and access decide it, not seniority or budget. If a competent person in this role finishes a useful unit of work end to end in about a week, from material you could hand a stranger, a trial reaches the real job. If the shortest real unit runs a month, or needs production credentials, a security review or a release train, then a week of trial measures onboarding and nothing else.
Ask the hiring manager two questions and let the answers decide the format.
- What is the shortest unit of work in this role that a competent person completes end to end? A campaign brief and its first asset. One denial worked to a decision. A research readout from a supplied extract. If the honest answer is a quarter, no trial length fixes it.
- What access does that unit need, and how long does granting it take? Count the security review, the vendor seat, the data agreement. Access that takes eight days to grant cannot support a five-day trial.
Reachable in a week, typically: marketing and content, recruiting, support and customer success, analysis from a supplied extract, design, journalism, and most agency-shaped work where a deliverable is self-contained by construction.
Not reachable, typically: platform and infrastructure engineering, security, anything touching production systems or regulated data, revenue-cycle and clinical work that lives inside a system of record, and roles whose deliverable is a decision that plays out over two quarters.
For the second list, the choice is not trial-or-assessment. It is which shorter format you trust, graded against a key you wrote first. One reminder goes with that: a week spent watching someone ramp on unfamiliar systems tells you about ramp, which METR's result suggests you will misread as speed in either direction 4.
Common questions
How long should a paid trial be?
Long enough to contain one complete unit of the real work, and no longer. For most roles where a trial is feasible at all, that is two to five days. Past a week you are running contract-to-hire, which needs a contract, a scope and a stated end date rather than an invitation. Shorter than two days is a work sample, and a work sample is cheaper to standardise, cheaper to grade and easier to offer every finalist on the same terms.
What should I pay for a trial project?
The role's real rate for the time, on your normal payment terms. Divide the target salary by working days and pay that, or pay the contractor rate you would pay a specialist for the same days. Underpaying does not reduce the obligation you took on by asking someone to work, and it changes who accepts. Pay the same amount to everyone who takes the trial, regardless of outcome, and pay it whether or not you use the output.
Is an unpaid trial ever acceptable?
Treat it as off the table. The FLSA defines employ as to suffer or permit to work, so permitting someone to do your work is the statutory shape of employing them, and not paying only removes your defence for the hours. Unpaid also selects hardest against the candidates who most need the job. If the exercise is short enough that pay feels unnecessary, run it as a work sample under an hour or two, which is a different thing with its own rules.
Can I use the work the candidate produced during the trial?
Only if you paid for it and the ownership is settled in writing before the trial starts. If you intend to ship it, say so in the brief and price the days accordingly. Deciding afterwards that a rejected candidate's week of output is useful is the fastest way to turn a hiring process into a dispute. The cleaner design is a task that is real but parallel to live work, which also removes most of the access problem.
Does a trial replace the rest of the interview loop?
No. A trial is one observation of one person on one task, and a re-analysis of selection validity estimates put structured interviews at the top of the ranked procedures. Run the structured interview, keep the trial for the last one or two candidates, and use the trial to answer questions the loop left open rather than to re-run it. If a trial contradicts everything else you saw, that is a reason to look again, not a verdict.
How do I compare two candidates who did different trial tasks?
You cannot, and the fix is upstream. Fix one task for everyone at that stage before the first candidate sees it, and write the rubric before the first trial. Where the tasks already differ, grade each against the same criteria item by item (framing, what got checked, what got refused, what was kept), and treat the comparison as weak evidence rather than a ranking. Two tasks means two procedures, which is also the harder position to defend.
References
- 1. 29 U.S. Code § 203 - Definitions ✓ law.cornell.edu Section 203(e)(1) defines employee as any individual employed by an employer, and 203(g) defines employ as including to suffer or permit to work.
- 2. Employment Tests and Selection Procedures ✓ eeoc.gov Work samples and job simulations are listed among employment tests and selection procedures, and where a procedure has disparate impact the employer must show it is job-related and consistent with business necessity.
- 3. 29 CFR 1607.14 - Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) ✓ ecfr.gov Content validity holds to the extent the selection procedure is a representative sample of the content of the job.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues from repositories they knew took 19% longer with AI tools allowed, and afterwards still believed the tools had sped them up by 20%.
- 5. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range ✓ europepmc.org Revised validity estimates cut most high-ranked selection procedures by .10 to .20 points, and structured interviews emerged as the top-ranked selection procedure.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.