Assessment design

Live Coding, Take-Home, or AI-Allowed Work Sample: Which Catches Real Skill?

Between live coding, an unwatched take-home and an AI-allowed work sample, the authored work sample catches real skill in most fields. Its signal comes from a constraint someone who knows the work wrote into the packet, and a rubric that grades the working record rather than the deliverable. Live coding still samples the job where the work is done in front of people: support, sales engineering, a pairing session. A take-home survives only where the deliverable is a model or dataset that exposes a confident wrong answer, not prose.

The takeFormat shopping is a way of avoiding the expensive part. A format can be bought in an afternoon. The constraint that makes one worth running has to be written by somebody who knows the work, and that person is already busy, so the argument stays on formats, where it is cheap. I cannot prove this from the literature, but a 2022 re-analysis cutting every format's validity estimate reads like a letdown for that reason: it compares instruments, and what varies between two companies running the same instrument is who authored it. The hour you would spend adding a third round is the only part of this nobody can sell you.

Where Olive fits

Open a role and see what the work shows

If you build the third format yourself, the packet and the answer key are the expensive parts, and the second case is the one nobody budgets for. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session, and the candidate is granted the same document.

Rank your shortlist

Which format still catches real skill?

The AI-allowed work sample, in most fields, but only when it is authored rather than bought, and graded on the record rather than the deliverable. Live coding still samples the job for roles genuinely performed in front of other people. An unwatched take-home catches almost nothing where the output is a document. The costs run in the opposite direction to the signal.

FormatWhat it costsWho leavesWhat it can no longer see
Live codingAn interviewer hour per candidate, and it never gets cheaper as volume risesPeople who need to think before they speak, and anyone whose performance drops while watched 1Whether they would have checked anything, given their normal tools and an unhurried hour
Unwatched take-homeCheap to send, expensive to grade honestly (two readers or it is not a score)Candidates with a current job or caregiving hours, at whatever length you setWho wrote it, what was rejected, and what was verified 23
AI-allowed work sampleAuthoring: the packet, the answer key, and a second case before the first leaks 5Fewer than a long take-home loses, if it runs under an hour and the brief is honestAnything outside the session it observes: one occupation, one afternoon 4

Three correlation figures circulate for exactly this comparison (roughly 0.6 for live coding, a little lower for take-homes, higher for work samples), and they appear across a dozen vendor blogs with no source attached to any of them. Treat unsourced coefficients as folklore. The published estimates moved the other way: a 2022 re-analysis correcting a systematic overcorrection for range restriction cut mean validity across selection procedures by .10 to .20 points, and structured interviews came out top-ranked in the revised table 4.

That last finding is the uncomfortable one, because it says the cheapest instrument in the building is competitive with the expensive ones. The format argument is worth less than the authoring argument.

Why does live coding punish deliberate thinkers?

Because being watched costs performance, and it costs the careful more than the quick. In a randomized controlled trial with 48 computer science students on the same whiteboard problem, 61.5% failed with an interviewer sitting in the room against 36.3% working privately, and the median score fell from three passing test cases to one 1. The watched group also finished sooner.

That combination is the tell. Public-setting participants took about a minute and a half less on average. That gap is not significant by itself, but paired with half the correctness it describes a specific failure: under observation, people commit early and stop checking. The authors' own conclusion is that interviewers "may be filtering out qualified candidates by confounding assessment of problem-solving ability with unnecessary stress" 1. Their post-hoc note that no woman solved the task in the watched setting while all four did privately rests on single-digit samples and should be read as a reason to look, not a result.

So the question is which fields reward the behavior that observation suppresses. Analytics, underwriting, audit and legal work all pay people to stop and re-read; a watched hour measures the opposite instinct and calls it ability. Support, sales engineering and client-facing consulting are the honest exceptions, because composure in front of a stranger is a task the job actually contains.

If you keep a live round, three changes cost nothing. Drop the whiteboard and let the candidate use the editor, the docs and the assistant they use daily. Running a live session with an AI assistant open changes what you learn rather than what you lose. Give the problem in writing and leave the room for the first ten minutes. And take the think-aloud retrospectively, walking the finished work, rather than demanding narration while the thinking happens.

When does a take-home stop telling you anything?

The moment the deliverable is a document. In a preregistered experiment with 444 college-educated professionals on occupation-specific writing tasks, ChatGPT raised graded quality by 0.4 standard deviations and compressed the spread between workers, because it helped the weakest most 2. A memo, brief, spec or analysis write-up now comes back at roughly the same quality from everyone who submits one.

The paper puts it flatly: the tool "mostly substitutes for worker effort rather than complementing worker skills" 2. Read that as a warning about your rubric: the structure and cleanliness you have been scoring are no longer evidence about the person who sent them. This is the whole of the problem behind take-homes that all come back polished.

The usual patch, a page asking candidates to describe how they used AI, does not survive contact with the evidence either. METR ran a randomized trial with 16 experienced open-source developers on 246 real issues in repositories they already knew; with AI tools allowed the work took 19% longer, and afterwards the developers still believed the tools had sped them up by 20% 3. If people inside the work cannot feel a 39-point swing, their written account of their own process is not a measurement.

Take-homes still carry signal in one case: where the deliverable is not prose. A service with a failure mode nobody named, a model whose numbers have to reconcile, a dataset whose answer turns on a join that is never mentioned in the brief. All three still separate people, because the assistant produces something confidently wrong and the artifact says so. Whether a take-home tells you anything at all with AI in play comes down to that distinction, not to the honor system.

On drop-off, measure your own and distrust everyone else's. No trustworthy cross-company completion curve is published, and the figures that circulate come from vendors with a length to sell. Log three numbers a week (invited, started, submitted) against the number of minutes your brief promises, and you will have a better curve for your funnel within a month than any survey can give you.

What makes an AI-allowed work sample work?

Someone has to write your field's real constraint into the material. An AI-allowed sample is not a take-home with permission attached. Its signal comes from a packet that reads settled until a person opens it, and from a rubric that grades the working record instead of the deliverable. Buy the task off a shelf and you get neither, because the constraint is the part that does not generalize.

What that constraint is depends entirely on the occupation:

  • Software engineering. The input the existing tests do not cover, which the assistant will not think to look for.
  • Data and analytics. A result an assistant explains fluently and a dataset that will not correct it.
  • Financial analysis. Which comparable set, and the reason it is not the three the model reached for first.
  • Marketing. The most quotable statistic in the folder, sourced from a vendor's own small customer survey.
  • Legal operations. A playbook, a system record and a signed precedent that do not agree with each other.
  • Underwriting. An application and a survey that do not describe the same building.

The reason the packet has to contain a real trap is that fluency and accuracy come apart even inside tools built for the job. A preregistered evaluation found Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time 7. Nothing in the output announces which third you are holding, so the only observable act is somebody going outside the conversation to settle it.

Write the rubric as acts, not adjectives. The Uniform Guidelines are direct about why: a selection procedure holds up on content validity "to the extent that it is a representative sample of the content of the job," and one "based upon inferences about mental processes cannot be supported solely or primarily on the basis of content validity" 6. Judgment is named there too, among the constructs a content strategy cannot carry. "Shows good judgment" is a claim about someone's mind. "Re-derived the segment figure against the underlying source, and the recommendation moved" is a sampled piece of job content. Working out what an AI assessment should actually measure starts at that line.

Count the cost before you commit. Federal guidance lists work samples as costly to develop in both time and money, needing periodic updating, and time-consuming and expensive to administer because someone has to observe and rate performance, and it recommends them where the number of applicants being tested is limited 5. Budget a second case before the first has gone out thirty times. The one thing the format wins outright is how it reads to candidates: applicants often perceive work samples as very fair 5.

Pick the format by field, not by fashion

Match the instrument to what the work looks like when nobody is hiring. If the job is performed live in front of other people, a watched session samples the job. If the output is a document produced over days, watching someone type it tells you nothing worth an interviewer hour, and the material has to carry the constraint instead. Most roles are the second kind.

FieldUseWhy
Software engineeringLive pairing on unfamiliar code, assistant openThe job is collaborative and the acceptance behavior shows in minutes
Data and analyticsAuthored work sample, AI allowedThe work is deliberate by construction; observation measures the opposite trait 1
Consulting and marketingAuthored work sample with a buried source errorThe deliverable is prose, so the artifact separates nobody 2
Legal, audit, underwritingAuthored work sample over conflicting recordsThe fastest confident answer is the wrong one
Support and sales engineeringLive sessionComposure in front of a stranger is a task the role contains

Two rules survive whichever you pick. Two readers score independently and compare before anyone talks, because a rubric that two people cannot apply the same way is a rubric problem rather than a candidate problem. Calibrate it on five old submissions first. And keep the panel honest about what it is reading: managers who do not use these tools themselves reliably read polish as skill, which no format fixes.

The limit worth stating out loud is that none of this is a validity claim about your loop. One exercise is one observation of one occupation on one afternoon, and the revised meta-analytic estimates say every procedure on the table predicts less than the older numbers promised 4. Run two formats where the decision is expensive, cut the third, and spend the saved hour on the packet rather than on another round.

See what gets scored

Common questions

Can you run all three formats in one loop?

It fits in a loop, and it is usually the wrong trade. Three formats is three to five candidate hours plus two interviewer hours, and the added drop-off falls hardest on people with a current job. Pick two with different jobs: one that samples the work under realistic conditions, and one short live conversation about the specific decisions in what came back. If a live coding round and a work sample would both test the same act, the work sample is the one carrying more, because the record outlives the room.

Should candidates be allowed to use AI in a live session?

Allow it, and say so in the invitation. A ban measures a version of the job that no longer exists, and on any unwatched format it is unenforceable anyway, so all it produces is a quiet split between candidates who followed the instruction and candidates who did not. Allowing it openly is also what makes the interesting behavior legible: what they asked for first, what they refused, and what they checked outside the conversation. Watching someone accept a confident wrong answer takes about four minutes and tells you more than the finished code.

How long should each format be?

Keep a live session under an hour, keep an unpaid take-home under ninety minutes, and pay for anything longer. Length is the single lever with the clearest effect on who withdraws, and the candidates it removes first are the ones with jobs and caregiving hours rather than the ones without skill. If the task genuinely needs a day, make it a paid trial and state the rate in the invitation. Adding process artifacts to an existing brief means cutting scope somewhere else, not stacking requirements.

Are the validity numbers quoted for these formats reliable?

The specific coefficients passed around comparison blogs are not, because none of them names a study. The published research went in the opposite direction from the marketing: a 2022 re-analysis found that common corrections for range restriction had substantially overestimated validity, cutting mean estimates by .10 to .20 points across selection procedures, with structured interviews ranking highest afterwards. Use that as a warning about precision generally. Any number offered for a format without a citation should be treated as a claim rather than a finding.

Does this hold for roles with no technical artifact?

Yes, and often more strongly, because a document is exactly the deliverable AI compresses. Every occupation has material that reads settled until someone opens it: a compensation benchmark with the wrong geography, a pin cite to a paragraph that does not say what the memo claims, a market figure attributed to the whole category. Build the packet around that. If you cannot name yours, ask the two strongest people on the team what they re-check before signing anything, and write the trap from their answer.

References

  1. 1. Does Stress Impact Technical Interview Performance? Behroozi, Shirolkar, Barik and Parnin, ESEC/FSE 2020 (NC State author copy), 2020. chrisparnin.me Randomized controlled trial with 48 computer science students: 61.5% failed the whiteboard task in the watched setting against 36.3% in private, median correctness fell from 3 passed test cases to 1 (Mann-Whitney U, p=0.038, d=0.57), watched participants finished about 1m36s sooner on average, and the authors conclude interviewers may be filtering out qualified candidates by confounding problem-solving ability with stress.
  2. 2. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence Noy and Zhang, MIT (working paper; later published in Science), 2023. economics.mit.edu Preregistered experiment with 444 college-educated professionals on occupation-specific writing tasks: output quality rises 0.4 SDs, inequality between workers decreases because the tool benefits low-ability workers most, and it mostly substitutes for worker effort rather than complementing worker skills.
  3. 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers on 246 real issues took 19% longer with AI tools allowed, and still believed afterwards that the tools had sped them up by 20%.
  4. 4. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range Sackett, Zhang, Berry and Lievens, Journal of Applied Psychology, 2022. europepmc.org Correcting systematic overcorrection for range restriction reduced mean validity estimates by .10-.20 points across selection procedures, with structured interviews emerging as the top-ranked selection procedure.
  5. 5. Assessment and Selection: Work Samples and Simulations U.S. Office of Personnel Management, 2024. opm.gov Work samples may be costly to develop in both time and money and may require periodic updating; they may be time consuming and expensive to administer and require individuals to observe and sometimes rate performance; applicants often perceive them as very fair; best used where a limited number of applicants is being tested.
  6. 6. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Content validity holds to the extent the procedure is a representative sample of the content of the job; a procedure based upon inferences about mental processes cannot be supported solely or primarily on content validity, and judgment is named among the constructs a content strategy cannot carry.
  7. 7. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools Magesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford RegLab and HAI (arXiv), 2024. arxiv.org Preregistered evaluation finding Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinate between 17% and 33% of the time.

7 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.