Pipeline

Move the Decision to the Stage That Still Produces Evidence

Put the hiring decision at the first stage where the candidate produces something under conditions you set. The resume screen still has a job, sorting on facts you can check, but it stopped carrying the hire once a polished document became cheap to produce. That usually means one evidence stage much earlier than most loops place one, and fewer interview rounds behind it, because those rounds were compensating for a screen that no longer separates people.

The takeThe expensive mistake is not choosing wrong; it is choosing nothing and letting the answer be round three by default. A loop that decides at round three has already spent four interviewers' hours on a pool assembled by a filter nobody trusts, and the people who would have proved themselves in an hour of real work were cut in seconds by a reader looking at prose. Pick the stage on purpose and pay for it on purpose.

Where Olive fits

Open a role and see what the work shows

Olive is priced per attempt rather than per seat, and one attempt returns six evidenced findings on a single candidate, each anchored to a timestamped moment in the session. Ten attempts a month are free, so an evidence stage can be piloted beside the loop you already run.

Rank your shortlist

Which stage is actually deciding your hires right now?

Look for the stage after which almost nobody is turned down. In Ashby's benchmark set, recruiter screens pass roughly 35% of the candidates who reach them, and post-onsite and offer stages pass 95% and 81% 1. A stage that passes nearly everyone is ratifying a decision rather than making one. The decision happened earlier, usually at a document review nobody counts as a stage.

Two cautions before you build on those numbers. They are passthrough rates among candidates who reached each stage, so they do not multiply into anything end to end, and the heaviest cut of all happens at application review, which is not in the series at all 1. The sample is one vendor's own customer base, weighted toward venture-backed technology companies, so the levels will not transfer to your funnel.

The shape does. A loop that filters hard at the top on a document and then converts at 95% after the onsite has put its decision in the place where it collects the least evidence, and has staffed several hours of interviewer time to confirm what was settled upstream. The onsite feels decisive because it is where people argue. By then the argument is about a pool that was assembled somewhere else. That shape is measurable one round at a time by comparing the read going in with the decision coming out, which is whether the screen overturns the resume read applied to the first round of the loop.

So the useful question is not how many rounds to run. Which single stage would you defend to a candidate who asked why they were cut, and did that stage give them anything to do? If the answer names a stage where a person produced work you specified, the loop is sound and probably too long. If it names a document, the loop has no evidence stage at all, and dropping the resume screen in favour of a work sample becomes a live option.

What the resume screen can still be trusted to do

Sorting on facts that are true or false: dates, licences, location, work authorization, a named employer somebody could call. That is a real job and it removes real volume. What the screen cannot do any more is tell you who is good, because the parts of a document that used to carry that signal are now the cheapest parts of it to produce.

The evidence there was thin long before generated prose arrived. A meta-analysis of 81 independent samples put prehire work experience, which is most of what a resume screen reads, at a corrected correlation of .06 with later job performance, and experience with relevant tasks, jobs or occupations did no better at about .07 2. Those are corrected figures, so the near-zero result is not an artefact of under-correction. Read the scope carefully: this is experience measured at hire, which is exactly what a screen sees, and it says nothing about licensure or a legally required minimum qualification, which stay in the screen for reasons that have nothing to do with prediction.

What changed recently is the rest of the page. The summary, the achievement bullets and the register were never strong evidence, and they are now free to produce, so a screen that weights them is partly sorting applicants by which drafting tool they opened. Keep the checkable spine and stop grading the prose: the sequence is worked through in reading an application in order of what you can check.

Narrowing the screen's remit also gives you something to test. A screen that only claims to establish scope can be audited against its own claim, which is the practical route into knowing whether your screen is throwing away the wrong people. A screen that claims to identify quality cannot be audited at all, because nobody wrote down what quality meant before the reading started.

Put a price on the evidence stage before you place it

The cost per candidate decides where an evidence stage can sit, so work it out before you draw the funnel. An hour of a hiring manager's time against a twelve-hundred-application pool is not a plan. SHRM's benchmarking of its own member organizations puts median cost per hire at $1,244 for nonexecutive roles against a mean of $4,683 on a badly skewed distribution 3, and that definition excludes interviewer and manager time entirely.

That skew is the whole problem in one line. A process at the 25th percentile spends $354 a hire and one at the 75th spends $4,375 3, so an evidence stage costing forty dollars a candidate is trivial in the second world and impossible in the first at any real volume. Work out your own number first: candidates entering the stage, cost per candidate, and how much of that cost is somebody's hour and how much is a vendor's invoice.

Then place the stage where the arithmetic survives. Three shapes work in practice:

  • After a cheap checkable gate, on the fifty to eighty people who cleared it. The common case, and the one that keeps per-candidate cost bounded.
  • On everyone, when the pool is small or the role is genuinely unreadable from paper. Expensive per head, cheap in total.
  • Before any screen at all, which only works when per-candidate cost is close to zero. Price it before dismissing it: assessing everyone up front instead of screening is arithmetic, not ideology.

Building a bigger screen runs into a ceiling. Modelling a screening stage made only of methods that scale to volume, with the structured interview left out, Berry and colleagues found the validity-maximising composite topped out at .51 against .61 when the structured interview was in the mix, and pushing that screen-only battery past a four-fifths adverse impact ratio dropped it to .39 4. Those are modelled ceilings, optimally weighted in a way no real screening stage is. They are not numbers to expect from a live process, and resume screening is not in the model at all. The comparative point is the usable one: the methods that scale are collectively weaker than the one that does not, so the top of a funnel cannot be made to carry a decision by adding more of itself.

How do you rebuild the loop around one evidence stage?

Take out one interview round for every evidence stage you add, and say out loud which stage now carries the hire. Adding a stage without removing one is how loops reach seven hours and lose the candidates who had options. The rounds you are cutting were mostly there to compensate for a screen that stopped discriminating, so cutting them is not the sacrifice it sounds like.

1. Write down the stage. One sentence: the hire is decided at X. If two people on the team write different sentences, that disagreement is the finding, and it explains most of your debrief arguments. 2. Audit the artifact. For stage X, name what the candidate produced and which conditions you set. No artifact means no evidence stage, whatever the stage is called on the pipeline board. 3. Count the cost. Candidates entering X, multiplied by cost per candidate, with interviewer hours priced at their real rate rather than at zero. 4. Cut a round. Pick the interview round that has never changed an outcome. Most loops have one, and the round worth cutting is often the second conversation about fit. 5. Size the gate above it. Set the cheap top-of-funnel pass so the number reaching X matches what step three says you can afford.

Cannot afford any evidence stage at step four? Keep one round and make it the evidence. Structured interviews came out top ranked at .42 in the corrected 2022 estimates, with job knowledge tests at .40 and work sample tests at .33 5. Those are pooled correlations with supervisor ratings, not accuracy rates, and the credibility bands around them overlap heavily, so the order is no argument that a work sample is the weaker stage. Take .42 as a reason to structure the round you already run, and nothing more.

One failure mode deserves naming before you start. A team that moves the weight onto an assessment it cannot afford to run at volume will quietly start using the resume screen to decide who gets assessed, which puts the decision back where it began with an extra stage bolted on. The order matters: price the evidence stage, then size the gate in front of it, then write the criteria that gate uses.

See a sample report

Common questions

Does this mean we should stop screening resumes?

No. A screen that sorts on checkable facts is cheap, fast and legitimate, and at any real volume something has to reduce the pile before a person spends real time on it. What changes is what the screen is allowed to conclude. It can establish that someone is real, in scope, and meets a stated minimum. It cannot establish that they are good, so it should not be the stage that ends a candidacy on a judgment about quality. Keep the screen, narrow its remit, and put the quality judgment somewhere the candidate actually does something.

Which stage should carry the decision if we cannot afford an assessment?

A structured interview, run properly and treated as the evidence stage rather than as a conversation. It costs interviewer hours where an assessment costs vendor fees, and it is the best-evidenced single method in the corrected estimates. Treating it as the evidence stage means the same questions in the same order for every candidate, written anchors on the rating scale, notes taken during the round, and scores submitted before anyone talks. That is a different thing from the round most teams already call an interview, and it is the cheapest upgrade available to a team with no assessment budget.

How would we know the resume screen is cutting the wrong people?

Sample the rejects. Take fifty applications the screen turned down on one closed requisition, have the hiring manager mark anyone they would have wanted to meet, and count. A handful is noise from a bar you set deliberately. A large share means the criteria are selecting for something other than the job. The slower and better check is whether the people who cleared your screen comfortably outperform those who cleared it narrowly, which most teams cannot run for lack of hires, and that is why the reject sample is the practical version.

Is a take-home the evidence stage?

Only if the candidate produces something you specified, under conditions you set, and somebody reads it against a written standard. A take-home assembled from public exercises is no longer that, because the exercises are solved. Around 13% of hires in one large benchmark set included a take-home component 1, so it is already a minority practice, and the ones still worth running tend to be short, paid, and built on material that is not sitting in a public repository. If yours is none of those, it is a filter dressed as evidence.

Does moving the decision earlier make the process longer?

It should make it shorter. The trade is one evidence stage in exchange for at least one interview round, and most loops carry rounds that exist because nobody trusted the screen. Elapsed time usually falls as well, because a stage the candidate runs on their own clock does not wait for four calendars to align. What does rise is the cost per candidate at that one stage, which is why the arithmetic belongs before the redesign rather than after it.

References

  1. 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report Ashby, 2026. ashbyhq.com Supports the stage passthrough figures (35% recruiter screen, 95% post-onsite, 81% offer), the point that the heaviest cut happens at application review outside the series, and the 13% of hires including a take-home component.
  2. 2. A meta-analysis of the criterion-related validity of prehire work experience Personnel Psychology, 72(4), 571-598 (Van Iddekinge, Arnold, Frieder and Roth); record and abstract at the University of North Florida Digital Commons, 2019. digitalcommons.unf.edu Supports the claim that prehire work experience, the field a resume screen mostly reads, correlates about .06 with later job performance and about .07 when the experience is task or occupation relevant.
  3. 3. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the cost-per-hire figures used to price an evidence stage: median $1,244 for nonexecutive roles, mean $4,683, 25th percentile $354 and 75th percentile $4,375.
  4. 4. Insights from an Updated Personnel Selection Meta-analytic Matrix: Revisiting General Mental Ability Tests' Role in the Validity-Diversity Tradeoff Journal of Applied Psychology, 109(10), 1611-1634 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2024. filiplievens.squarespace.com Supports the modelled ceiling on a screening stage built only of methods that scale: .51 validity maximising, .61 with a structured interview available, and .39 once the adverse impact ratio is pushed past four-fifths.
  5. 5. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, doi 10.1037/apl0000994; full-text copy opened at gwern.net, 2022. gwern.net Supports the corrected 2022 ranking used to argue for keeping one structured round as the evidence stage: structured interviews .42, job knowledge tests .40, work samples .33.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.