Pipeline
Which Funnel Metrics Still Mean Anything Now That Candidates Use AI?
Of the metrics on a hiring funnel dashboard, three still mean something now that candidates draft with AI. Each stage's pass rate, read as a slope against your own pre-2023 history and never as a level against a published benchmark. Evidence density, the share of advances resting on a work act somebody observed rather than described. And the per-stage records adverse-impact review requires, whose meaning never changed. Everything else now tracks how cheaply a candidate can produce what the stage asks for.
The takeNothing about your funnel got worse this year. It got legible. A conversion rate always measured how expensive an application was to produce. Nobody ever proved that cost tracked effort, but for two decades it held closely enough that nobody had to admit effort was the thing on the chart. The proxy is gone and the chart is still on the wall. Most teams will answer by adding stages, which raises the production cost and buys the proxy back for a while. I doubt the next round buys as much. Stop defending the ratio and decide what you're willing to watch a person do.
Where Olive fits
Open a role and see what the work shows
Evidence density is only countable where a stage records an act rather than a description, and most stages cannot do that. Olive is priced per attempt with ten attempts a month free, and one attempt returns six separately-evidenced findings on a candidate, each anchored to a moment in the session, written by a human reviewer and granted to the candidate as the same document you read.
Rank your shortlistWhy did every stage pass rate move at once?
Because clearing a stage got cheaper for everyone at the same time. A pass rate is the ratio between a bar you set and a population's ability to clear it, and the bar didn't move. What moved is the cost of producing whatever the stage asks for: a tailored resume, a clean cover letter, a structured answer, a finished take-home. Adoption of that shortcut is what your conversion chart tracks now.
That effect has been measured directly on the artifact stage. In a field experiment in an online labor market with nearly half a million jobseekers, those given algorithmic writing assistance on their resumes were hired 8% more often, and the researchers found no evidence that employers were less satisfied with the people they hired 1. Read both halves. The stage moved without the underlying population changing, and the people who cleared it were not worse, which is exactly why the ratio stopped being informative rather than merely becoming wrong.
The denominators moved too. LinkedIn's US job-competition measure has seekers submitting roughly twice as many applications as in late 2019 while the number of jobs per seeker sits close to pre-pandemic levels, so a per-opening count climbs faster than a per-seeker count does 2. An application-to-screen ratio that halved is describing an inbox. Why applications per opening tripled is the top-of-funnel version of the same arithmetic, and it is the one stage where nobody disputes the cause.
So the honest summary of your dashboard is this: every ratio whose numerator is a person clearing a bar made of written output moved in the same direction, over the same period, for reasons that have little to do with your roles. The chart isn't broken. It is measuring a different thing now, and none of the axes are labeled with what.
Which funnel metrics still carry signal?
Three. Each stage's pass rate read as a slope against your own history rather than as a level against anyone's benchmark. Evidence density: the share of advances at that stage resting on a work act somebody observed. And the per-stage selection records you are required to keep anyway, which are the one set of numbers whose meaning did not change at all.
What to retire, and what each one is actually reporting now:
- Resume-to-screen conversion as a quality signal. It reports how many applicants used a tool that formats a resume the way your filter likes. Keep it as a volume forecast; what replaces ATS keyword screening when every resume matches is the redesign question underneath it.
- Interview pass rate as a measure of bar height. A rehearsed structured answer is now a thirty-second generation, so the rate moved without anyone deciding to lower anything.
- Take-home completion and take-home pass rate. Both rise when the task is one a model can finish, and both are silent about who finished it.
- Any external benchmark table. A 2026 funnel benchmark is a weighted average of other companies' roles, tooling and adoption. Publishing fresher percentages re-baselines the numbers without touching the logic that made them uninformative.
- Time-in-stage as anything but scheduling. It measures calendar friction and always did.
The records are a separate matter, and the reason is older than any of this. Under the federal Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4, an employer maintains records disclosing the impact its selection procedures have by race, sex and ethnic group, and where the total process shows adverse impact the individual components should be evaluated for it 5. A clean bottom line does not cover a component: in Connecticut v. Teal, the US Supreme Court held in 1982 that a nondiscriminatory bottom line neither prevents a prima facie case nor provides a defense to one, on a record where the written examination passed 54.17% of black candidates against 79.54% of white candidates while the overall promotion rate favored black candidates 4. Keep the per-stage counts even for stages you have stopped trusting as quality signals, and have counsel read them before you redesign around them.
How do you count evidence density at each stage?
One ratio per stage: advances that rest on a work act someone observed, over total advances. A work act is something the candidate did in front of you. A source opened, a figure recomputed, a model's answer refused with a reason, a failing case run before the patch was accepted. A description of any of those is not one of them, and that distinction is the entire measurement.
At the resume screen the answer is zero, structurally. A resume records claims about acts and never acts, so no rewrite of the filter produces one. That is a finding rather than a gap in your instrumentation: the stage is a volume control, and reading its conversion rate as quality was always a stretch that AI merely made obvious.
What counts as an observed act, by the field you are hiring for:
- Financial analysis. The segment figure was reconciled to the filing, and the recommendation changed when it didn't reconcile.
- Software engineering. A failing case was written and run before the generated patch was accepted, not after.
- Marketing. The vendor survey behind the most on-message statistic was opened, and its sample size named in the brief.
- Legal operations. The signed precedent was pulled when the playbook and the system record disagreed, and the position moved to match it.
- Healthcare revenue cycle. One claim in the queue was conceded rather than appealed, with the reason stated.
- Supply chain. Three quotes written on three sets of terms were re-added onto one, and the cheapest headline price lost.
Recording it costs one field on the stage's scorecard: *what was observed*. Count the non-empty ones. Two reviewers should be able to fill that field separately and agree. "Showed good judgment" cannot be counted twice the same way, while "recomputed the multiple against the 10-K" is a yes or a no. Grading a stage this way is the same discipline as grading take-homes when every submission comes back polished, moved up to the dashboard. See how Olive measures this.
Set a new baseline per job family, not company-wide
Both inputs differ by field. How completely a model can produce what a stage asks for is not the same for a positioning brief and an underwriting file, and neither is how easily that stage can ask for something checkable. Average them into one company-wide number and you get a figure that describes no team, moves for reasons you cannot attribute, and hides the family where the change was largest.
Cut the dashboard by job family (engineering, analysis, operations, clinical, field sales) and record four things per family per quarter: volume in, pass rate at each stage, evidence density at each stage, and the outcome the hires actually produced. The fourth column is the one most funnels have never carried, and without it the first three cannot be scored against anything.
Be honest about how long that takes to read. A slope needs about two quarters per family before it means more than noise, and only where volume is high enough that a stage's pass rate isn't swinging on four decisions. Below that threshold, stop reading slopes and read levels: evidence density is a share you can quote on a single cohort, because it counts what a stage recorded rather than comparing two periods against each other.
Careful with the cut itself. Job families are not org-chart teams, and the useful line is the occupation the work belongs to rather than the manager it reports to. A data analyst embedded in marketing sits with the analysts for this purpose, because what changed is the availability of a fluent first draft in that kind of work, not the reporting line.
Which stage actually predicted who worked out?
Run it backwards. Take the hires from four to six quarters ago, sort them by the outcome you actually care about, and ask which stage's record separated the top from the bottom. That is a different question from which stage rejected the most people, and it is the only one whose answer tells you where to spend the next round of design effort.
Expect a modest, wide result rather than a clean one. In the revised meta-analytic estimates that Sackett and colleagues built after correcting how range restriction had been applied, structured interviews top the list at a mean validity of .42 with an 80% credibility interval running from .18 to .66, and the predictors at the top are the ones specific to individual jobs: structured interviews, job knowledge tests, work sample tests, empirically-keyed biodata 3. Widely used predictors came out lower on average than the prior published estimates, work sample tests by .21. A stage that separates almost nothing is a normal finding, not a broken query.
The backwards pass is also the only check that catches a stage which is inflating and predicting at the same time. If the interview pass rate rose and the people it passed still separate on outcomes, the stage is fine and the number just needs a new baseline. If it rose and they don't, the stage is now measuring the artifact. Why a brilliant round misses what the quarter bills for is what that second case looks like from the manager's side, and whether assessment scores predict performance at all is the version of this question asked before the funnel is built.
Name the limit of all of it before you present it. A stage that only records descriptions has an evidence density of zero however good its questions are. A structured question captures a candidate saying how they would check a confident claim, not checking one. And a rubric written in-house has no grounding beyond your own funnel: a backwards pass can tell you a stage separated your hires, never that it measured what the occupation requires.
Common questions
Should you re-baseline against a published 2026 funnel benchmark?
No. A benchmark table is a weighted average of other companies' roles, tooling and AI adoption, and republishing fresher percentages fixes none of what broke. It restates the same inflated ratios with a newer date attached. Your own pre-2023 numbers are the only comparison that holds your roles, your bar and your sourcing constant. If you have no usable history, skip levels entirely and start counting evidence density, which is readable on its first cohort because it counts what a stage recorded rather than comparing two periods.
Is a rising interview pass rate good news or bad news?
Neither, until you know what changed. A pass rate rises when candidates get stronger, when the bar softens, when sourcing narrows, or when the thing the stage grades gets cheaper to produce, and the last one has been moving hardest. Read the rate beside two other numbers: how many of those passes cite something observed rather than described, and how the same cohort looked two quarters into the job. On its own the rate names no cause, which is why it can't be acted on.
How long before a new baseline means anything?
About two quarters per job family, and only where the family's volume is high enough that a stage's pass rate isn't swinging on four decisions. Below that, read levels instead of slopes. Evidence density is a share you can quote on the first cohort, because it counts what a stage recorded rather than comparing periods against each other. A small team gets more from one honest count of observed acts than from a trend line drawn through noise.
What replaces time-to-hire as a quality metric?
Nothing, because it was never one. Time-to-hire measures scheduling: calendar friction, panel availability, how long a decision sits. It moved this year for several reasons, and none of them describe candidates. Keep it, report it to the people who own the calendar, and stop putting it on the same chart as anything claiming to say who you hired. The metric that belongs there is the outcome column: what the hires from four to six quarters ago actually produced.
Do you still need per-stage pass rates if you don't trust them?
Yes. Under the federal Uniform Guidelines, an employer maintains records disclosing the impact of its selection procedures by race, sex and ethnic group, and the US Supreme Court held in Connecticut v. Teal (1982) that a nondiscriminatory overall result is no defense to a component that screened people out. Those records answer a different question than your dashboard does. Distrusting a number as a quality signal is not a reason to stop counting it; check with counsel before changing what you record.
Can you measure evidence density at a resume screen?
No. It is zero by construction. A resume records claims about acts and never acts, so no filter rewrite produces one, and that is the useful finding rather than a gap in your tooling. Two options follow. Treat the stage as a volume control and stop reading its conversion rate as quality, or move the first observation earlier so something checkable happens before five hours of panel time have been spent on a candidate nobody has watched work.
References
- 1. Algorithmic Writing Assistance on Jobseekers' Resumes Increases Hires ✓ nber.org Field experiment in an online labor market with nearly half a million jobseekers: treated jobseekers who received algorithmic writing assistance were hired 8% more often, with no evidence that employers were less satisfied.
- 2. Labor Market Tightness: LinkedIn's Measure of Job Competition ✓ economicgraph.linkedin.com US job seekers submit roughly twice as many applications as in late 2019 while the number of jobs to job seekers sits close to pre-pandemic levels.
- 3. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors ✓ doi.org Structured interviews top the revised list at a mean validity of .42 with an 80% credibility interval of .18 to .66; the highest-validity predictors are job-specific ones; corrected estimates are lower than prior published figures, by .21 for work sample tests.
- 4. Connecticut v. Teal, 457 U.S. 440 (1982) ✓ law.cornell.edu A nondiscriminatory bottom line neither precludes a prima facie case nor provides a defense to one; the written examination passed 54.17% of black candidates against 79.54% of white candidates while the overall promotion rate favored black candidates.
- 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 ✓ law.cornell.edu Users maintain records disclosing the impact their selection procedures have by race, sex or ethnic group, and where the total selection process shows adverse impact the individual components should be evaluated.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.