Pipeline
Budget the Whole Loop in Candidate Hours, Not Round by Round
Four to six hours of candidate time across the whole hiring process, scheduling overhead included, is a defensible ceiling for most non-executive roles, and anything past it should be paid or removed. No professional body publishes that total, so it is your budget, not a benchmark you can quote. Set it before the next intake meeting, make each stage buy its hours by producing evidence no earlier stage produced, and keep reviewer minutes in a second column beside it.
The takeThe number that actually binds is not hours; it is the ratio between the two columns. A take-home a candidate returns in twenty minutes and a reviewer needs ninety to interpret has quietly moved its cost to your side of the table, where nobody is counting. Price both columns or you will keep adding stages that look free to the candidate and keep wondering why the loop got slower.
Where Olive fits
Open a role and see what the work shows
Olive is priced per attempt rather than per seat, and one attempt returns six evidenced findings on a single candidate as an input to your decision. Ten attempts a month are free, so a pilot can run beside the loop you already have.
Rank your shortlistHow many candidate hours is a loop allowed to cost?
Somewhere between four and six hours, end to end, for a non-executive role. That covers the application, a screen, one scoped exercise, a panel and the scheduling overhead around them. No professional body publishes a whole-loop figure, so treat that range as a working default you own, not a standard you can cite at a hiring manager who wants a sixth round.
There is no standard because the genre prices one stage at a time. How long an interview should run. How long a take-home should take. Each answer is defended on its own terms and each is roughly reasonable, so the total grows by one reasonable increment at a time until it lands well past anything the team would have agreed to up front, and no single person made that decision.
Employers already measure what a hire costs. SHRM's benchmarking data, from a random sample of its member organizations, puts median cost-per-hire at $1,244 for nonexecutive roles, with the mean in the same table at $4,683, and the definition sums agency fees, advertising, job boards, referral costs, travel, relocation, recruiter pay and the recruiting system, divided by hires 1. Candidate hours are not in that list. Neither is interviewer time. That distribution is heavily right-skewed, which is why the median and the mean sit so far apart, and why no single figure describes what a hire costs.
So write your own number down before the next intake meeting. The argument coming is about a sixth stage, and you will lose it if the only figure in the room is the one attached to that stage alone.
Track reviewer minutes in a second column
Write two numbers beside every stage: what it costs the candidate and what it costs your side to read. The second column is the one that moved. A polished submission now arrives faster than a reviewer can work out what it demonstrates, so a stage priced as cheap on the candidate's clock can be the most expensive hour in the loop on yours.
Run one live req through the list. Stage, candidate minutes, reviewer minutes, with the second number counting notes and scheduling as well as the meeting itself. Filled in for one engineering loop, it looks like this.
- Application. Eight candidate minutes on a short form, three reviewer minutes if a person reads it at all.
- Recruiter screen. Thirty candidate minutes, closer to forty of yours once notes and rescheduling are counted.
- Take-home. Four candidate hours as briefed, twenty minutes as actually spent, and an hour or more of reviewer time to work out which of those two happened.
- Panel. Three candidate hours, six of yours across four interviewers and the debrief that follows.
The third row is the one that should stop the meeting. It is the stage everyone defends as respectful of the candidate's time and it is now the worst trade in the loop: least evidence per candidate hour, most reviewer hours per candidate hour. Once a stage costs a candidate more than an hour and hands you a document you then have to interrogate, the honest options are to shrink it or to pay for it, and paying candidates for take-home time is a narrower decision than it first sounds.
Which stages actually earn their hours?
A stage earns its hours when it produces evidence no earlier stage produced and someone can name that evidence in one sentence. Run that test stage by stage on a live req. If two stages return the same kind of evidence, the second is buying confirmation rather than information, and confirmation is what you end up paying candidate hours for by accident.
Two failures dominate. One is duplication: two conversations in which a prepared candidate narrates past work are one conversation held twice, whatever the question lists look like. The other is newer and easier to miss, because the artifact still arrives on time and still looks like the thing it used to be.
Exercises assembled from public problem banks are the clearest case of the second. Evaluated against 115 Python problem statements taken from HackerRank, OpenAI's Codex solved 96% of them zero-shot and 100% few-shot, with clear signs it was reproducing memorized code rather than synthesising it 3. The memorization half is the load-bearing one: a task drawn from a public bank was testing a lookup before any model touched it. That is a 2022 model on a curated benchmark, so treat the figure as a floor. The candidate hours are real either way, and what the exercise proves has thinned to almost nothing.
Volume is not the defence people assume either. Around 13% of hires in one large applicant-tracking dataset included a take-home component 2, and that dataset skews to venture-backed technology employers, so the practice is a minority one even where it feels standard. If yours has to survive the evidence test, design against the harder constraint: testing AI skills without adding an hour to the loop.
Who pays for the hours nobody counts?
The candidates least able to absorb them, which makes the total a fairness question and not only a courtesy. Unpaid weekday hours land hardest on people already working, people with caregiving duties and people who cannot move a shift. Those hours never appear in a funnel report, because someone who cannot spare them quietly stops replying.
The reflex when a loop gets too long is to cut the exercise, and candidate-reaction evidence does not support that reflex. Pooling favorability ratings across studies and rescaling everything to a five-point scale, applicants rated interviews highest at 3.84 and work samples next at 3.63, ahead of resumes at 3.57 and references at 3.33 4. Those means come from people rating written descriptions of methods, not sitting through them, from studies published up to 2004, and each rests on between five and ten ratings. Read it as a ranking. That ranking puts the job-shaped exercise near the top of what candidates accept, above the resume screen and the reference call.
So spend the budget where the evidence is and take the hours back somewhere else. A fifth conversation, a second culture round, an async video answer nobody watches closely: those are the hours a shorter loop should buy back, which is the whole of which interview round to cut.
On Monday, take one open req, write both columns, and mark every stage that cannot name evidence unique to it. Cut, shorten or pay for each one. Then put the total in the posting, where it does the most work: a candidate who knows the loop is five hours can plan for it, and a hiring manager proposing a sixth conversation has to argue against a number that is already public.
Common questions
Is four to six hours a benchmark I can quote to my hiring manager?
No, and quoting it as one will backfire the first time someone asks for the source. There is no published whole-loop candidate-hour standard: the guidance that exists prices individual stages, which is why the total drifts. Present the number as your own budget, set deliberately, with the reasoning attached. A figure the team decided together holds up better in an intake meeting than a borrowed one, because nobody can dispute your right to set it.
Should candidates be paid for take-home work?
Pay once the exercise crosses roughly an hour, or shrink it below that line. The threshold matters more than the rate: a short scoped task inside an hour is a normal part of applying, while an unpaid half-day is an ask that quietly filters for people with spare weekday time. If paying is impossible in your organisation, that is a real constraint and the honest response is to shrink the exercise until an apology is unnecessary.
Does scheduling overhead really count as candidate time?
Count it, because the candidate pays it. Holding a two-hour window for a one-hour panel, rearranging a shift, travelling, and waiting for a slot that moves twice all cost time that never appears on the agenda. A practical rule is to add half again to any stage that requires a live appointment. It usually turns a loop that looks like four hours on paper into six, which is exactly the gap between what teams think they ask for and what they ask for.
What about executive or senior roles?
Raise the budget and keep the same test. Senior loops legitimately run longer because more people carry a veto and the evidence needed is broader, so eight to twelve hours is not unreasonable at that level. Name the evidence each conversation produces that no earlier one did. Seniority buys more stages, not more stages that repeat each other, and the meeting-the-board round is usually a sell step with no veto behind it.
How do I count a stage where the candidate uses AI and finishes early?
Count both numbers separately and treat the gap as information. Brief four hours, learn a candidate spent twenty minutes, and what you have learned is about the task: it was work a model could carry. That is a signal to change the exercise, not to police the candidate. Ask what the person decided, rejected and checked, because those are the parts an early finish does not explain away.
References
- 1. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) shrm.org Supports the median and mean cost-per-hire figures for nonexecutive roles and the claim that SHRM's cost-per-hire definition counts recruiting spend, not candidate or interviewer hours.
- 2. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the claim that around 13% of hires in this dataset included a take-home component, so the stage is a minority practice.
- 3. Codex Hacks HackerRank: Memorization Issues and a Framework for Code Synthesis Evaluation arxiv.org Supports the claim that an exercise drawn from a public problem bank no longer functions as a work sample, including the memorization finding.
- 4. Applicant Reactions to Selection Procedures: An Updated Model and Meta-Analysis ecommons.cornell.edu Supports the favorability ordering placing interviews and work samples above resumes and references, cited with its rating counts and its age.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.