Pipeline
If the Screen Never Overturns the Resume Read, It Is a Calendar Step
Count disagreements, not pass-through. For your last fifty screens, record what the pre-call read of the application was and what the post-call decision was, then count two rates separately: strong-on-paper candidates the call stopped, and weak-on-paper candidates the call advanced. If both sit near zero, the round is a scheduling step. Pass-through cannot answer this, because a screen that ratifies the resume perfectly scores a healthy rate.
The takeThe two rates are not symmetric and should not be averaged. A screen that only ever stops strong-looking candidates is a rejection filter, which is a legitimate design and a much cheaper one to run asynchronously. A screen that also rescues weak-looking ones is doing something a document cannot, and that is the expensive capability worth keeping a live call for. Report them separately or the interesting half disappears into the total.
Where Olive fits
Open a role and see what the work shows
Olive is priced per attempt rather than per seat, and an attempt returns six evidenced findings on one candidate: an input to a decision a person still makes. Ten attempts a month are free, so a pilot can run beside the round you are already measuring.
Rank your shortlistHow would you know the screen is producing information?
Compare the decision before the call with the decision after it, fifty times. For each screen, record the pre-call read of the application (advance, unsure, reject) and the post-call outcome, then count the cases where they disagree. Disagreement is the entire output of the round. A screen that agrees with the resume every time has told you nothing you did not already have.
Split the disagreements into two counts and keep them apart. The first is candidates who looked strong on paper and were stopped by the call. The second is candidates who looked marginal and were advanced because of something they said. Both are information about different things, and a single combined percentage hides whichever one is zero.
Fifty is roughly where a rate of five or ten percent stops being one or two people. If you run fewer screens than that in a quarter, use every screen since the requisition opened and treat the result as a direction. Even then the comparison beats any published benchmark, because it is measured on your own funnel.
The cost side of this is real. Recruiter screens sit where the volume is, and a round that produces nothing is not free: it is the scarcest hour in the process spent confirming a document.
Why doesn't pass-through answer this?
Because pass-through counts bodies rather than information. In Ashby's benchmark across more than 54 million applications, recruiter screens pass around 35% of the candidates who reach them while post-onsite conversion runs at 95% and the offer stage at 81% 1. The 95% is the instructive figure: a stage that almost never stops anyone is doing very little independent work, whatever its rate looks like on a dashboard.
Those are stage-to-stage rates among candidates who reached each stage, they do not multiply into anything end-to-end, and stage names are configured per customer inside one vendor's product. What survives the caveats is the logic. A high pass-through means the decision was effectively made earlier. A low one means the stage is stopping people, which is not the same as the stage knowing something.
So the two failure modes are invisible to the metric everyone publishes. A screen that advances exactly the people the resume already favoured earns a perfectly ordinary rate while producing nothing. A screen that stops most of them looks harsh and may be the only honest round in the process. Pass-through cannot tell those apart, and neither can time-to-hire, cost-per-hire, or any other throughput measure.
Other funnel ratios have the same shape. Which funnel metrics still mean anything works through them, and the pattern repeats: a ratio whose denominator changed silently.
Add one field to the screen form
One field, filled in before the call: what you would have decided from the application alone. Advance, unsure, reject. It takes four seconds and it is the only way the comparison exists later, because a disagreement you never wrote down cannot be recovered from an ATS afterwards. The post-call decision is already stored, so this is the missing half.
Three implementation details decide whether the number is worth having.
1. Record it before the call, not after. A pre-call read reconstructed from memory after a decision is the decision wearing a different label. 2. Let the screener see the application but not anyone else's read. If a coordinator has already marked the candidate strong, you are measuring deference. 3. Keep the same three values all quarter. A scale that gains a level halfway through cannot be compared across the period.
Writing things down at all is the harder half of this in smaller teams. In the same benchmark, scorecard completion runs near 49% at organizations under 25 employees and closer to 72% at organizations of 500 or more, and completion is higher for candidates who get hired, so the record of rejected candidates is systematically thinner 1. A one-field pre-call read survives that better than a full scorecard, which is most of the argument for keeping it to one field.
The same benchmark found around 38% of scorecard pairs carrying at least a one-point difference between interviewers, with nearly half of those differences straddling the yes/no line on a four-point scale 1. Nothing in the data says which interviewer was right, so what it shows is instability at the point where the decision gets made. That is a reason to keep the evidence behind a judgment alongside the judgment itself, and a pre-call read is the cheapest form of that.
Read a low reversal rate three ways
A near-zero reversal rate, meaning the call almost never overturns the pre-call read, has three explanations and only one of them means the round is dead. The resume stage ahead of it may be unusually good, in which case the screen has little left to find. The screener may be reading someone else's verdict before the call, which makes agreement meaningless. Or the round genuinely is a scheduling step. Check the first two before acting on the third.
The first explanation is testable and worth testing on its own terms, because a resume stage that is carrying the whole process is a risk concentrated in one unmeasured place. Whether the resume screen still predicts anything is the same question one stage earlier, and running both comparisons together tells you where the decision actually lives.
If the round survives all three readings and still produces nothing, the choices are to redesign it or to remove it. Redesigning means giving it a stated failure condition and one claim to test, which is what a recruiter screen is supposed to decide. Removing it means moving the disqualifiers onto the form and letting candidates reach a substantive round sooner, which shortens the process for everyone still in it.
One caution about where this measurement stops. It tells you whether the call changes decisions, not whether the decisions are right. Only downstream outcomes answer that, and practitioners report little confidence in their own measurement of them: in LinkedIn's Future of Recruiting 2025, drawing on a September 2024 survey of 1,271 recruiting professionals in management roles across 23 countries, 25% said they felt highly confident in their organization's ability to measure quality of hire effectively 2. Nobody audited those organizations. The 25% records how confident people feel about their own employer, and it comes from a company that sells hiring products. It is still the right expectation to set before anyone promises that a funnel change improved hires.
Add the field to the next screen anyone runs. The comparison needs about fifty calls behind it before it reads as anything, so the cost of putting it off is a quarter.
Common questions
What counts as a good reversal rate?
No band has been published, and a borrowed one would defeat the point of measuring your own round. What matters is whether both directions are non-zero and whether the rate is stable enough to notice a change. A round that stops five in fifty strong-looking candidates and advances three in fifty marginal ones is clearly producing something. A round at zero and zero is not. Between those, compare the round against itself over time rather than against another company.
Does this need an ATS change to measure?
No, not for the first quarter of it. A spreadsheet does the job: a candidate identifier, the pre-call read and the post-call outcome, filled in by whoever runs the screens. The point of putting it in the ATS eventually is that it survives the person who set it up leaving. Start with the spreadsheet rather than waiting on a configuration change, because the measurement takes fifty calls to become readable and the queue for ATS work is usually longer than that.
Doesn't a high stop rate just mean the screen is too harsh?
It might, and the way to tell is to look at what happens to the people who pass. A screen stopping most candidates while the rounds after it stop almost nobody is doing the process's filtering, which may be exactly right if the criterion is defensible and applied the same way to everyone. A screen stopping most candidates while later rounds also stop many is filtering on something the later rounds do not share, which is worth reading as a disagreement about the bar rather than as harshness.
Should the same person do the pre-call read and the call?
For this measurement, yes, and that is a deliberate limitation. One person means the comparison is about whether the call changed that person's mind, which is the question being asked. It also means some of the agreement is consistency rather than accuracy. Splitting the two across people measures something different and more expensive: whether two readers of the same application agree at all. Both are worth knowing; the one-person version is the one you can run this quarter.
References
- 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the stage passthrough figures, the scorecard completion rates and the scorecard-pair disagreement rate used in this article.
- 2. The Future of Recruiting 2025 business.linkedin.com Supports the claim that a quarter of surveyed recruiting professionals feel highly confident measuring quality of hire, with its self-report limits stated.
2 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.