Assessment design

A Timer Measures Speed, and Speed Is Now the Cheap Part

Keep a generous outer bound on a hiring assessment so a two-hour exercise cannot become a two-day one, and drop the tight timer that scores speed. Production speed collapsed once AI assistants arrived, so a hard clock now mostly separates who had the faster tool and the quieter room, not who can do the job. The thing worth knowing is what a candidate does with the last twenty minutes: check it, revise it, or send it.

The takeTimers survive by default, not because anyone re-derived the case for them. The old argument was that a clock standardised the test: same task, same minutes, comparable results. That held while everyone produced the work at human speed, and it stopped holding once one candidate's assistant drafts in seconds while another candidate types. What has not moved is the cost side: the accommodation and fairness questions a countdown carries are all still there. A feature that got less informative and no cheaper is a feature running on inertia.

Where Olive fits

Open a role and see what the work shows

An Olive session runs 50 to 70 minutes on the candidate's own clock, and none of the six findings turns on how quickly the work was finished: each one rests on a timestamped excerpt a human reviewer chose and wrote up.

Rank your shortlist

What does the clock actually measure?

Production speed, and not much besides. A timer records how quickly somebody turns a prompt into output, which was a decent proxy for capability while producing that output was the expensive part of the work, and it stopped being one around the time every candidate got an assistant. What a tight clock separates now is who had the faster tool, the better connection and the quieter forty-five minutes, and none of those appear in the job description.

The compression is measured, and it happened exactly where assessments live. In a pre-registered experiment, 444 college-educated professionals completed occupation-specific writing tasks; the half given ChatGPT finished 10 minutes faster, 37% quicker than a control group averaging 27 minutes, and graders scored their output higher, with the largest gains going to the weakest writers 1. That is short professional writing done once, online, for pay, with no revision cycle and no colleagues, so it does not describe a working week. It describes an assessment almost perfectly.

Read the last clause again, because it is the part that matters for a timer. The gains were largest for the weakest writers, which means the tool compressed the spread the clock was there to reveal. A timed exercise on a task an assistant handles well now returns a narrower and noisier reading than it did two years ago, from the same stopwatch.

There is one honest use left. Where a role genuinely has a rate, such as a support queue with a service level or a trading desk, throughput is part of the job and testing it is legitimate. Even then the test has to be the role's actual rate against the role's actual tools, not a generic countdown bolted onto a written exercise.

Keep the outer bound and drop the tight timer

Set a cap generous enough that nobody gains by spending a weekend, then stop scoring against it. Two hours of stated work inside a 48-hour window is a scheduling device and a fairness one, because it lets a candidate with a job and a family pick their own two hours. A countdown that starts when the link opens is a measurement, and it is measuring the wrong variable.

The obvious objection is that a stated cap is unenforceable, and it is. You cannot tell whether somebody spent two hours or five. The design answer is to make extra hours stop paying: fix the deliverable, cap its length, and say in the brief that reviewers read the first two pages. A task where more time buys more polish is a task that rewards free evenings, timer or not, and a countdown papers over that rather than fixing it.

Three lines make a stated cap work in practice:

1. Name the cap in hours and say it is a cap. Not a guide, not an estimate. 2. Say what to hand in if the clock runs out. Submit what exists plus a note on what would come next. Without that sentence the honest candidates are the ones who lose. 3. Tell reviewers what to ignore. Anything past the stated length or the stated scope does not earn credit. This is the line that actually enforces the cap, and it belongs in the answer key as much as in the brief.

Live sessions are the exception that proves the rule. A scheduled hour is bounded because it occupies an interviewer's calendar, which is a logistics constraint rather than a claim about the candidate. Grade what happened in the hour and resist the temptation to read the pace of it as evidence.

Test the part the assistant cannot do

If speed genuinely matters for the role, put the clock on verification rather than production. Drafting is fast now. Deciding whether the draft is right is not. An exercise that hands somebody three claims, one of which fails, and asks which one and what they would do about it, puts a timer on judgment, which is the part that stayed expensive.

Self-reported speed is also unreliable in a way that undercuts the whole premise. In a randomised trial, 16 experienced open-source developers completed 246 real tasks on mature repositories they had worked on for years; allowing early-2025 AI tools made them 19% slower, after forecasting a 24% speedup beforehand and still estimating a 20% speedup afterwards 2. Sixteen developers in one setting on codebases they knew intimately is not a general productivity finding, and the magnitude does not transfer. What transfers is the direction of the gap between what those developers believed and what the clock recorded. If experienced people are wrong about their own pace on work they know by heart, nobody in a hiring loop should trust an intuition about what a candidate's finish time meant.

This reframes what the last twenty minutes of an exercise are for. A candidate who finishes early and spends the remainder checking the number they are least sure about has shown you something a faster submission never could. Build the assignment so that behaviour is possible: leave slack in the cap, and make sure at least one claim in the material is checkable against something in the pack.

The underlying choice, whether to measure how fast somebody works with an assistant or how well they judge what it returns, is worth settling once for the whole loop rather than per exercise: see speed versus judgment with AI.

Can the timer survive a request for extended time?

Only if the result means the same thing with the time extended. Under 29 CFR 1630.11, the US regulation implementing the ADA (2023 edition), an employer must select and administer employment tests so that results reflect the skill the test claims to measure rather than a candidate's impaired sensory, manual or speaking skills, except where those skills are what the test measures 3. So the first request for extra time is when a clock reports what it was actually scoring.

That regulation is live and in force, and it is narrower than a general accessibility duty: it addresses test administration where a disability impairs sensory, manual or speaking skills, and the separate duties around reasonable accommodation sit elsewhere in the same part. It does not mandate any particular timer or format. What it does is make the question unavoidable once someone asks.

The cost objection does not hold up. The Job Accommodation Network surveyed 26,028 employers who had contacted it between 2019 and 2024 and drew its report from 5,406 responses; of the 1,425 who gave cost information, 61% said the accommodation cost nothing to implement and 33% reported a one-time expense with a median of $300 4. Everyone in that sample had already sought accommodation help, so they are more willing than employers at large, and the figures cover workplace accommodations broadly rather than assessments. It is still enough to retire the argument that adjusting a test is expensive.

Decide the extended-time version before the round opens rather than improvising it when a request arrives, and write down what an untimed result means next to what a timed one means. If you cannot write that sentence, keep the generous cap and drop the countdown. The mechanics of offering and documenting this on an AI-based exercise have their own answer in ADA accommodations on an AI assessment.

Whether your timer produces different pass rates across groups is a question about your own data, and no argument about countdowns settles it in either direction. If the assessment is a gate, that is worth measuring rather than reasoning about. Confirm any of this with counsel before relying on it.

See what gets scored

Common questions

Is a timer ever the right choice?

Yes, where the role has a genuine rate and the test reproduces it: a support queue with a service level, a trading desk, a live incident drill. The condition is that the pace being tested is the pace of the job, with the tools of the job. A generic countdown attached to a written exercise fails that test, because no role requires producing a memo in exactly forty-five minutes with no chance to check it.

What time limit should a take-home have?

State a work cap of two to three hours and a return window of two to three days. The cap describes how much effort the exercise deserves; the window lets a candidate fit it around a job. Longer caps do not buy better signal and they cost completions among the busiest candidates. If the assignment genuinely needs a day, it is a paid project rather than a screening step.

Do candidates game an untimed assignment by spending far longer?

Some do, which is why the fix is the deliverable rather than the clock. Cap the length of what gets submitted, fix the format, and tell reviewers to read only what falls inside the stated scope. When extra hours cannot buy extra credit, overspending stops being rational. A countdown does not solve this either, since anyone determined to overspend can simply start, stop and restart.

Does dropping the timer make grading harder?

Only if the timer was doing grading work it should not have been. Without a clock, the differences between submissions have to come from the rubric and the answer key, which is where they should have come from anyway. Teams that find grading harder after removing a timer usually discover their key was thin, and that is a useful thing to learn in a round with four candidates rather than forty.

What do we say when a candidate asks for extra time?

Grant it against a policy you wrote before the round, and do not ask for a diagnosis. Decide in advance how much extra time is available, whether it applies to the work cap or the return window, and who approves it, then apply that answer consistently. The failure mode is improvising per request, which produces inconsistent treatment and a record that is very hard to explain later.

References

  1. 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) MIT Department of Economics, 2023. economics.mit.edu Supports the measured time reduction on short occupation-specific writing tasks, and the finding that gains were largest for the weakest writers, behind the claim that a timer discriminates less than it used to.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the gap between forecast, believed and measured speed, behind the claim that an intuition about somebody's pace is not evidence about their capability.
  3. 3. 29 CFR 1630.11 - Administration of tests (Regulations to Implement the Equal Employment Provisions of the Americans with Disabilities Act) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 2023. govinfo.gov Supports the requirement that test results reflect the skill the test claims to measure rather than an impairment, which is the standard a timed assessment has to meet.
  4. 4. Costs and Benefits of Accommodations (Low Cost, High Impact report) Job Accommodation Network (JAN), funded by the U.S. Department of Labor's Office of Disability Employment Policy, 2025. askjan.org Supports the survey counts and the share of accommodations reported as costing nothing, behind the claim that adjusting a timed assessment is cheap.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.