Assessment design

Two Hours Is the Ceiling, and It No Longer Bounds Effort

Two hours of candidate time is the working ceiling for a take-home, and ninety minutes is better. Set the number backwards from what you can grade well, not forwards from what feels fair. Shortening it narrows the exercise to one competency, so plan the rest of the loop to cover the others. Nothing else is lost: the stated time stopped rationing effort, because a candidate working with a model can spend twenty minutes or eight hours and hand back work that looks alike. Bound the artifact instead of the hours.

The takeThe two-hour number survives because it sounds fair, not because anyone measured it. What it actually protects is your grading calendar. Four submissions a month at ninety minutes each is an afternoon of careful reading. Twelve at three hours each is a stack nobody finishes, and unread submissions are the quiet way an assessment stage turns into a formality nobody has retired. The exercise nobody finishes reading measured nothing, whatever it asked candidates to spend.

Where Olive fits

Open a role and see what the work shows

Length is the easy decision here; the answer key and the evidence trail are the ones that take the time. Olive runs a role-grounded assignment sized at 40 to 60 minutes and returns six findings, each anchored to a timestamped excerpt from the session.

Rank your shortlist

How long should a take-home actually be?

Two hours, and ninety minutes is better. The number is a convention rather than a finding, so treat it as the outer edge of what candidates will accept and then cut from there. What the evidence supports is that candidates accept the stage: in a meta-analysis of applicant reactions, work samples rated second only to interviews among ten selection methods, at 3.63 against 3.84 on a rescaled five-point scale 1.

Read that finding narrowly. Those ratings come from a handful of descriptive studies in which people rated written descriptions of methods they had not experienced, the corpus closes in 2004, and there is no row in it for an unpaid multi-hour take-home. What it supports is the category: a job-shaped exercise is well received. Your specific brief sits outside its range.

Three rules, in the order they bind:

  • Two hours of candidate time is the ceiling, and the brief says so. Past that you are asking for an evening, and an evening is not equally available to a person holding a job, a second shift or a bedtime.
  • Ninety minutes if the exercise survives a cut, and most do. What gets deleted is setup, formatting, and the second deliverable nobody argues about in the debrief.
  • Paid, or shorter. If the task cannot go under two hours without becoming trivial, the honest choices are to pay for the candidate's time or to run it live. A longer brief with a gentler time suggestion attached is neither.

Why doesn't the stated time bound effort any more?

Because the artifact stopped carrying the hours. A candidate who spends twenty minutes and one who spends eight can hand back documents a reviewer cannot tell apart, so the number in the brief rations nothing. Researchers submitted unedited model output through a real university examinations system under fake student accounts across five modules, and 94% of it drew no concern of any kind from markers working normally 2.

The setting is psychology coursework and not a hiring take-home. No detector was involved, and the authors describe their method as using AI in the most detectable way available, which makes 94% a floor. The transferable part is the mechanism: unsupervised written work no longer separates the person from the assistant on the strength of the finished page.

What is still enforceable is the artifact. A bound on the deliverable survives contact with an assistant, because a reviewer checks it on the page:

  • one decision and the reasoning behind it, in two pages, with a stated word limit
  • one function against a named signature, plus the test you would write for it
  • five slides, no appendix
  • a single recommendation with the evidence that was rejected, and why

State the intended time anyway, and say in the brief that extra polish earns nothing. Then hold to that when every submission comes back polished, because the candidates who believed you are exactly the ones you cannot afford to lose to a scoring habit nobody wrote down.

Fairness is the other argument the cap cannot settle on its own. A stated two hours binds only the candidates who honour it, and the one who can quietly add four more hours is the one with a free evening, so a shorter ceiling narrows that gap without closing it. The bound on the artifact is the part that applies to everybody, which is the reason to write it down.

Set the length from your grading capacity

Start with your own calendar, not the candidate's. Time yourself reading and scoring three past submissions against the rubric you actually use, multiply by the number of finalists you expect in a month, and check whether the result fits the week you have. If it does not, the exercise is too long, whatever candidates spend on it. That reviewer number is the one that binds, and almost nobody writes it down.

A worked version. A team hiring two analysts a quarter sends out roughly four exercises a month. Careful scoring of a two-page memo against six criteria runs about 25 minutes, and a second reader on the finalists adds half of that again. Four submissions come to something under three hours, which fits an afternoon. Push the brief to three hours of candidate work and the memo becomes a deck with an appendix: scoring runs closer to an hour, four submissions eat most of a day, and that is the point at which they start getting skimmed.

Skimming is the failure mode that matters, because it is invisible. The stage still appears on the pipeline diagram, candidates still spend their evenings on it, and the decision quietly reverts to whatever impression the resume left. A shorter exercise read properly by two people beats a longer one read once at speed, and it is the version you can still defend when somebody asks what the exercise measured. That question has a legal shape in the United States: the Title VII burden-shifting provision added by the Civil Rights Act of 1991 asks whether a challenged practice is job related for the position in question and consistent with business necessity 4. Whether your brief clears that is a question for counsel on your own process.

Two numbers belong on the requisition next to the brief: minutes of candidate time, and minutes of reviewer time per submission. If the second is under ten, the rubric is doing no work. If the monthly reviewer load runs past a day, cut the brief before the next posting opens.

What does a shorter exercise cost you?

Coverage, not depth. A ninety-minute exercise cannot test three competencies, so it tests one, and the debrief has to stop pretending otherwise. Drop the claim that a candidate is broadly strong. You keep a defensible read on the single thing you chose to look at, which is more than most three-hour briefs deliver.

Plan the coverage across the loop instead of inside the brief. Decide which competency the exercise owns, then put the others where they can be observed: reasoning under questioning in the live round, prior judgment in the reference conversation, written communication in the debrief note. A live working session instead of a take-home is the right trade whenever the thing you need to see is how somebody responds to a follow-up.

Comparability moves in your favour too. A long open brief mostly measures how much time each person could give it, and a short bounded one lets the difference you chose to look at show through. The stage is still worth running: the current meta-analytic estimate puts work sample validity at .33, correcting the .54 that still circulates, and lands it close to what a structured interview predicts 3.

Take the hedge with the number. Nearly all of those studies tested people already doing the job, no range restriction correction was applied, and none of them studied an unpaid multi-hour assignment. Read it as evidence that a job-shaped exercise earns its slot, and as nothing at all about how long that exercise should run.

On Monday: open your current brief, delete every deliverable except the one you would argue about in a debrief, then time yourself grading three old submissions. The second number is your ceiling.

See what gets scored

Common questions

Is there a legal limit on how long a take-home can be?

No US law sets a maximum length for a hiring exercise. Federal law asks a different question, one about job relatedness and business necessity under Title VII as amended in 1991, and length reaches it only through that. A very long unpaid assignment raises a separate question, about whether the work is compensable, and that turns on federal and state wage law and on what the business does with the output. Length itself is a design decision and a candidate-experience decision. If the exercise produces something the business will actually use, that is the point to stop and call counsel.

Should I tell candidates how long the exercise should take?

Yes, and state it as an intention rather than a limit you can police. A stated time lets candidates plan, gives the ones with less slack a fair basis to accept or decline, and sets the expectation you will hold to in scoring. Pair it with one sentence saying that additional polish earns no credit, then keep that promise when a submission arrives with a designed cover page. The number is only worth printing if your grading behaves as though it were true.

What if a candidate spends eight hours on a two-hour exercise?

You will not know, and chasing it is wasted effort. Bound the artifact instead: a word limit, a fixed number of deliverables, a named signature. A submission that blows past the stated bound is visible on the page and can be handled as a brief-following issue, which is a real signal about how somebody works. Effort beyond the bound that leaves no trace is not something to score, and treating a polished submission as evidence of overwork is guesswork dressed up as judgment.

Does a shorter take-home give a weaker signal?

It gives a narrower one. Narrower and weaker are different things. Signal comes from what the exercise asks for and how carefully it is read, not from duration. A ninety-minute brief asking for one contested decision and the reasoning behind it, read by two people against written criteria, tells you more than a three-hour brief scanned once at midnight. Decide which single competency the exercise owns and move the rest of the loop's coverage to rounds where it can be observed directly.

Should we pay candidates for take-home time?

Pay when the exercise is long, when it uses live material, or when the output has value to the business. Below about an hour, most candidates treat the exercise as an ordinary cost of applying. Above two hours you are asking for an evening, and the people who cannot give one are systematically the people with the least slack. Paying also disciplines the brief, because a team writing a cheque tends to cut the exercise back to what it actually needs.

References

  1. 1. Applicant Reactions to Selection Procedures: An Updated Model and Meta-Analysis Personnel Psychology, 57(3), 639-683 (Hausknecht, Day and Thomas); accepted manuscript PDF in Cornell eCommons, 2004. ecommons.cornell.edu Supports the claim that applicants rate work samples second only to interviews (3.63 against 3.84 on a rescaled five-point scale), and the caveats about the small descriptive sub-sample behind it.
  2. 2. A real-world test of artificial intelligence infiltration of a university examinations system: A “Turing Test” case study PLOS ONE (Scarfe, Watcham, Clarke and Roesch), 19(6): e0305354, 2024. journals.plos.org Supports the claim that unsupervised written work no longer separates a person's effort from a model's: 94% of the AI submissions drew no concern from human markers.
  3. 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the corrected work sample validity estimate of .33 rather than the widely repeated .54, and the concurrent-design caveat attached to it.
  4. 4. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases Office of the Law Revision Counsel, United States Code (prelim), 1991. uscode.house.gov Supports the FAQ statement of what federal law asks of a hiring exercise: job related for the position in question and consistent with business necessity.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.