Assessment design
Take-Home Assignment or Live Working Session: Which Shows AI Judgment?
Neither a take-home nor a live session wins outright at showing AI judgment; decide by how long the real work takes. If the job's unit of work finishes inside an hour and leaves visible moves, run the live session with the assistant open, ten private minutes before anyone watches, and no stopwatch. If the work spans days, an hour of watching shows you setup: cap a take-home, require a decision trail written during the work, and take twenty minutes on the candidate's file. Score both against one rubric.
The takeThe format argument is a proxy for a loss nobody wants to name: the deliverable stopped being evidence the moment a tool could produce one. Watching does not repair that. It moves the guess earlier. What survives in both formats is refusal, the part the candidate cut, because that is the one act the assistant does not perform on their behalf. Nobody has measured whether rejection predicts performance on the job, so take this as a stated bet: it tells you more per minute than anything else you can fit in a brief, and it is the line most rubrics still forget to write.
Where Olive fits
Open a role and see what the work shows
A walkthrough is still the candidate's account of decisions made off-camera, however well it is run. Olive puts the assignment where the record gets made instead (40 to 60 minutes of occupational work with an AI assistant available, think-aloud spoken or typed), and a human reviewer writes six separately-evidenced findings, each anchored to a moment in the session, with the candidate granted the identical document.
Rank your shortlistWhich Format Fits the Task Horizon?
Match the format to how long the real work takes. When the job's unit of work finishes inside an hour and leaves visible moves (a draft, a query written and run, a triage call), a live session shows AI judgment directly. When the real work runs across days, a live session catches the setup and nothing that follows, and a short take-home with a walkthrough gets closer.
Sort your own roles into two buckets before arguing about formats.
- Fast and observable. A marketer building campaign copy from a research packet. A support lead working a ticket queue. An analyst writing one query and reading the result. A recruiter drafting a screen. The whole loop (frame, generate, check, cut) fits in 45 minutes, and every move in it is something you can watch someone make.
- Long-horizon. A three-day financial model. A migration across a codebase nobody on the team wrote this year. A diligence memo over a 200-page packet. The judgment you care about is spread across hours: which of six threads to abandon on day two, which assistant output to stop trusting after it was wrong twice.
What makes this more than a preference is that people are unreliable narrators of their own long-horizon AI work. METR randomized 246 real issues from experienced open-source developers' own repositories to allow or forbid AI tools. With the tools allowed the developers took 19% longer, and after finishing they still believed the tools had made them 20% faster 1. So a candidate's account of a three-day build is worth something, but not as measurement, and neither is 45 minutes of watching the first hour of it. Measure How Fast They Work With AI, or How Well They Judge It? works through the trap next door.
What a Live Session Shows, and What It Costs
It shows the acts a finished document cannot: the first thing asked of the assistant, whether a source was opened or only cited, what got refused, what got checked. The cost is a measurement error most loops never account for. In a randomized trial with 48 computer science students, median correctness fell by more than half when an interviewer simply watched the candidate work 2.
The same trial measured higher stress and cognitive load in the watched condition, and none of the women in the public setting solved the problem while all of the women in the private setting did 2. Its authors propose two fixes, and both survive the move to AI-assisted work: give the candidate private time on the problem first, then bring them back to narrate what they did.
A live session earns its place when the work is observable, on three conditions.
- Private minutes first. Ten minutes alone with the problem and the assistant, then you join. You lose the opening move in real time; you gain a candidate who is solving rather than performing.
- The assistant stays open. If they will use it on Tuesday, they use it here. Banning it turns the session into a memory test and hides the delegation you called the meeting to see. How to Run a Coding Interview With an AI Assistant Open covers the mechanics for engineering roles.
- Nothing is timed to the minute. Speed under observation is the one thing this format measures reliably, and it is the thing you least want to select on.
What a Take-Home Still Tells You
It tells you the shape of long-horizon work: how the candidate scoped a task with nobody standing over them, what they sequenced first, and what got cut when the time ran out. What it cannot tell you is who decided any of it. A finished artifact reads the same whether the candidate wrote it, edited it, or accepted it whole.
The defect you are hunting is not obvious badness. In Stack Overflow's 2025 developer survey the most common frustration with AI tools, at 66%, was output that is almost right but not quite 3. Almost-right compiles, renders and reads well, then fails on the single figure the recommendation rests on. A read-through does not find it; tracing one number end to end does.
Three things make a take-home worth the candidate's unpaid hours.
- A cap you enforce. Ninety minutes, stated in the brief, with a deliverable small enough to fit inside it. A take-home is a selection procedure like any other test, so it has to be job-related and consistent with business necessity, and that burden sits with the employer 4. Hours you cannot justify against the job are an exposure and a completion problem at once.
- A stated AI policy. Allowed, or not, in one line. A rule you cannot enforce mostly selects for who ignored it, and Should Candidates Use AI on the Take-Home? works through where each answer holds.
- A trail written while the work happens. Assumptions and their basis, one claim checked outside the assistant, one thing cut and why. How to Grade a Take-Home When Every Submission Is Polished has the grading side of it.
Run the Hybrid: Short Take-Home, Then a 20-Minute Walkthrough
Budget about two hours of candidate time and split it. Ninety minutes on a capped take-home with the assistant allowed, then twenty minutes live on the candidate's own submission. The walkthrough is not a re-test and not a defense. It is you watching someone account for decisions they made three days ago, with the file open in front of both of you.
Five prompts fill the twenty minutes, and each asks for an act rather than an explanation.
- Open the file and show me where this number came from. Pick the figure the recommendation rests on. Someone who put it there can trace it in a sentence.
- Which claim did you go outside the assistant to check, and what came back? Including the case where the check returned nothing useful, which is a real answer and a common one.
- Show me something you cut. Rejection is expensive and specific. A submission that accepted everything looks identical to one where nothing was worth refusing, and this is the question that separates them.
- The deadline just halved. What goes? A fresh decision, made in front of you, on material they already know cold. This is the part of the live format worth keeping.
- Where would you not have used the assistant? The answer that counts names a step and a reason rather than a policy.
That set aims at whatever is left of knowledge work once a tool drafts. A survey of 319 knowledge workers across 936 first-hand tasks found the effort moving from producing an answer toward verifying information, integrating a response and stewarding the task, with higher confidence in the tool tracking with less critical thinking and higher confidence in one's own ability tracking with more 5.
One limit, stated plainly: a walkthrough is a retrospective account, not a record. It is harder to fake than the artifact, because it can be checked against the artifact and the artifact can be checked against nothing. It is still not the same evidence as watching the decision get made.
Score Both Formats Against One Rubric
Rate the same four things whichever format produced the evidence: how the problem was framed before anything was generated, what was checked outside the assistant, what was kept by hand rather than handed over, and what was refused. The format decides where the evidence comes from. It does not change what you are grading, and two rubrics are how two candidates become incomparable.
Structure is the part carrying the weight, and it is worth knowing how much weight there is. Sackett and colleagues re-ran the meta-analytic estimates behind the standard validity table and found the usual correction for range restriction had substantially overcorrected: most procedures kept their rank order but lost .10 to .20 in mean validity, and structured interviews came out top-ranked 6. Selection methods still work. They work less well than the table on your wall says, and the structured ones hold up best.
Four things to fix before the first candidate rather than after the third.
- Write the anchors first. Two or three worked descriptions per item, for this specific case. A key written after reading three submissions is a description of those three submissions.
- Hold the case fixed. Same task, same cap, same walkthrough prompts in the same order, for everyone at that stage. Both formats are selection procedures under the same expectation of job-relatedness 4.
- Rate item by item across candidates. Grade every framing, then every outside check, then every rejection. Reading one submission end to end is how a handsome artifact drags a weak decision through an entire rubric.
- Double-read the first five. Two raters, independently, then compare. Disagreement early is information about the key rather than about the candidates, and it is cheap to fix at that point.
Common questions
If I can only run one, which do I pick?
Pick by the work. If the job's real unit of work finishes inside an hour and leaves visible moves, run the live session and give the candidate private minutes before anyone watches. If the work runs across days, run the short take-home, because a live hour of a three-day task shows you setup. Where you can afford both, a capped take-home plus a twenty-minute walkthrough beats either alone and costs the candidate about two hours.
How long should the live session be?
Twenty to thirty minutes if it is a walkthrough of work they already did, and about forty-five if you are setting a fresh problem. Longer mostly buys fatigue. When you set a new problem, give ten private minutes with the assistant before you join: the evidence on being watched is strong enough that the opening minutes are worth protecting.
Should the candidate use AI during the live session?
Yes, if they will use it in the job. A ban turns the session into a memory test and hides the exact behavior you are trying to see: what gets handed over, what gets kept, what gets refused. Say it in the invitation so nobody has to guess, and ask them to work the way they normally would rather than perform a demonstration of tooling.
What if the candidate can't explain a decision from three days ago?
Ask for the artifact rather than the memory. Requiring a short decision note written while the work happens (assumptions, one outside check, one thing cut) means the walkthrough starts from a document instead of recall. If the note is missing and nothing in the submission can be traced back, that is a finding about the work rather than about a bad memory.
Does a walkthrough just reward candidates who present well?
Partly, which is why every prompt asks for an act rather than a story. Open the file. Show the check. Point at the section that got cut. A confident narrator with no trail runs out of material in about four minutes, and a quiet candidate with a traceable file does not. Rate what got shown, against anchors written before the first session.
Do I have to pay for the take-home?
Keep it short enough that the question stays small, and pay once the ask passes a couple of hours or the output is something you would use. Unpaid hours are also a selection effect: candidates with caregiving duties or a current job drop out first, which changes who you get to compare. A ninety-minute cap protects your completion rate and your defensibility together.
References
- 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues from their own repositories, randomized to allow or forbid AI tools, took 19% longer with the tools and still believed afterwards that the tools had sped them up by 20%.
- 2. Does Stress Impact Technical Interview Performance? ✓ chrisparnin.me Randomized controlled trial with 48 computer science students comparing private and public whiteboard settings: median correctness was cut more than half by being watched, stress and cognitive load were significantly higher, no women solved the problem in the public setting while all did in private, and the authors propose private focus sessions and retrospective think-aloud.
- 3. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite.
- 4. Employment Tests and Selection Procedures ✓ eeoc.gov Work samples and other employment tests are selection procedures that must be job-related and consistent with business necessity, with the employer carrying that burden.
- 5. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org Survey of 319 knowledge workers and 936 first-hand examples: effort shifts toward information verification, response integration and task stewardship, with higher confidence in the tool associated with less critical thinking and higher confidence in one's own ability with more.
- 6. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range ✓ europepmc.org Re-analysis of meta-analytic validity estimates: range-restriction corrections had substantially overcorrected, most procedures keep their rank but lose .10 to .20 in mean validity, and structured interviews emerge as the top-ranked selection procedure.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.