Pipeline
The Stage That Opens a Gap Is Rarely the Stage That Shows It
A group gap usually opens earlier in the hiring process than the stage where it becomes visible. A stage rate is computed only over the people who reached that stage, so every stage inherits the composition the ones before it produced, and the last stage displays a gap it mostly did not create. Compute entry counts and pass rates at every stage. A gap roughly constant from the first row belongs to sourcing; one that widens between two adjacent rows belongs to the rule at that boundary.
The takeDebriefs get blamed because they are the stage with a human voice and the thinnest paper trail, which makes them easy to accuse and impossible to clear. Say that out loud in the meeting where the gap is presented, because the argument otherwise runs on whoever speaks with the most conviction. Build the table first. The stage that turns out to have opened the gap is frequently one nobody had counted as a decision at all.
Where Olive fits
Open a role and see what the work shows
Olive is priced per attempt rather than per seat, and an attempt returns six evidenced findings on one candidate: an input to the decision a person makes at that boundary. Ten attempts a month are free, and the candidate is granted the identical report, so the evidence a decision rests on is legible to the person it is about.
Rank your shortlistWhy does the last stage show a gap it did not create?
Because a stage rate is conditional on who reached the stage. Every rule earlier in the funnel has already shaped the population a later stage sees, so that stage inherits a composition and then reports it as though it were a result. The visible gap and the causal gap sit in different rows of the same table, and only one of them can be moved by changing the stage you are looking at.
Across one applicant-tracking vendor's dataset covering over 54 million applications and 93,000 jobs, recruiter screens pass roughly 35% of the candidates who reach them, while post-onsite conversion runs at 95% and the offer stage at 81% 1. A 95% passthrough is not evidence that an onsite is easy. It means the decision was effectively made before the onsite ended, so the onsite does very little independent filtering while sitting exactly where a gap becomes visible to everyone.
Read those numbers with their limits attached. They are stage-to-stage rates among candidates who reached the stage, they do not multiply into anything end to end, the heaviest filtering happens at application review which is not one of these numbers, and stage names are configured per customer, so "onsite" is not a standard construct across the sample. The direction is what carries. Late stages ratify and early stages filter, and reading that backwards has a specific cost. Interviewer training aimed at a stage that inherited its gap gets measured against a number the training cannot move.
Where does the composition get set before anyone screens?
In sourcing and in ad delivery, before a single application is reviewed. Who sees the posting decides who applies, and that step usually sits outside the funnel report entirely because it happens on somebody else's platform. A stage table that starts at application review has quietly agreed not to examine the largest single determinant of who is in it.
A black-box audit ran paired job ads for similar roles at companies with different workforce gender mixes and found statistically significant gender skew in Facebook's delivery of those ads that could not be explained by qualification differences, while failing to find such skew on LinkedIn 2. That is a measurement of who gets shown a job, which happens before screening exists. It is also a field measurement taken on live ad platforms in 2020 and 2021. The LinkedIn trials drew few enough responses that the null there is weak evidence, and Meta has since changed how it delivers employment ads. Neither the positive nor the null result carries to today. The transferable finding is that the composition arriving at your funnel is partly produced by a system you do not operate and cannot inspect.
Referral channels do something structurally similar inside your own process. In the same benchmark dataset, 52% of referred candidates passed initial screens against 35% overall 1. That is not evidence that referred candidates are better. Referral status is visible to the screener, so part of the gap is what the referrer's endorsement does to the reviewer, and a channel converting at 52% while the general pool converts at 35% pulls the composition of hires toward the composition of the people already employed. That dataset carries no demographic variable. It gives you the mechanism and a reason to measure the effect on your own numbers, which are the only place its size can be read.
The same thing happens on a compressed timeline when a requisition hits its application cap: a posting that closed in two days selected on speed of arrival before any rule was applied to anybody.
Build the five-row table
Build one row per stage with four columns: how many entered, the composition of who entered, how many passed, and whether the cell is large enough to read. Include the stages that feel mechanical, because a uniform rule meeting a non-uniform population is exactly where a gap opens without anybody deciding anything. Composition is the column that does the work.
| Stage | Entered | Composition of entrants | Passed | Readable |
|---|---|---|---|---|
| Source | — | — | — | — |
| Application review | — | — | — | — |
| Screen | — | — | — | — |
| Interview | — | — | — | — |
| Offer | — | — | — | — |
Then read the composition column first. Pass rates come second, and only to explain what the composition column has already located. A composition roughly constant from the first row means the gap arrived with the pool and belongs to sourcing. A composition that shifts between two adjacent rows names the boundary where a rule did something, and that boundary is the one place worth opening this quarter.
Three cautions keep the table honest. Cells below your stated minimum get marked, not deleted, because a deleted row silently becomes an all-clear. Stages your applicant tracking system does not model, including application review and any vendor step, still get a row even if the numbers have to be rebuilt by hand. And if every stage rate moved at once for reasons unconnected to any group, the prior question is which funnel metrics still mean anything.
What if the gap really does open after the interview?
Then you have found a real one, and roughly half of the discrimination that shows up in job offers appears after the callback 3. That figure comes from the small number of field experiments that follow applicants past the interview invitation instead of stopping at the callback, which is what makes the two stages separable at all.
Across 12 such experiments, majority applicants received 53 percent more callbacks than comparable minority applicants and 145 percent more job offers 3. The hedges are heavy and they change how the number should be used. Twelve studies, skewed toward roles that could be tested in person and toward older designs; 145 percent is a ratio of offer rates, so it magnifies small absolute differences and must never be subtracted from 53; and the additional discrimination after the callback correlates only weakly with the callback gap, which means a firm cannot infer its later-stage behaviour from its screening numbers in either direction. The authors' own summary is that about half of the discrimination in job offers comes from the application-to-callback stage.
Hold on to that weak correlation: it is why the table has to exist at all. A clean screen licenses no conclusion about the debrief, and a dirty screen licenses none either. Once the boundary is named the work changes shape: the question stops being which stage feels unfair and becomes what the rule at that boundary actually reads, which is what makes a procedure biased rather than a person. Name one boundary per quarter and open the rule there. Two boundaries is an agenda item nobody finishes.
Common questions
Can I just compare pass rates by stage and pick the worst one?
Comparing them is fine; picking the worst one is not. A stage rate is computed over the people who reached the stage, so the worst-looking rate is often a stage inheriting a composition that earlier rules produced. Put an entry-composition column beside the pass rates and read that one first: the boundary where composition shifts is the boundary where something happened. The worst pass rate is still worth explaining, but it is a symptom whose cause frequently sits in a row above it.
Which stages do people forget to put in the table?
Sourcing, application review, and any step that runs inside a vendor's system. All three drop out for the same reason: nobody schedules them, so they never appear on the funnel diagram the software draws. Application review is usually the highest-volume decision in the entire process and often the least documented. Rebuild those rows by hand if the system will not produce them, because a table starting at the recruiter screen has agreed not to look where most of the filtering happened.
What if a stage has too few candidates to compute anything?
Mark the cell unreadable and leave it in the table. Deleting a row silently converts "too few to read" into "nothing wrong here", and those are different findings. Decide the minimum cell size before you look at the data, apply it to every row, and if pooling several quarters is needed to reach it, decide that in advance too. A window chosen after seeing the ratio is not a measurement.
Does a constant gap across every stage mean our process is fine?
It means your stages are not where the gap opened, which is not the same claim. A gap that arrives with the pool and travels through unchanged points at sourcing, posting, referral mix and ad delivery, and those are decisions somebody on your side makes. The stages themselves may still contain rules worth reading for other reasons. What a flat profile rules out is the theory that one boundary is responsible for the number.
Is the debrief usually the problem?
Usually not the origin, and it is worth saying so before the meeting rather than during it. A debrief sits at the end of four earlier decisions and inherits their composition, so it displays a gap it mostly did not create. It also leaves the thinnest written record, which is why an argument about it rarely ends in evidence. Both are reasons to compute its entry composition before anyone argues about it.
References
- 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the stage passthrough figures showing late stages ratifying rather than filtering, and the referral passthrough gap used to explain composition drift.
- 2. Auditing for Discrimination in Algorithms Delivering Job Ads arxiv.org Supports the claim that composition can be skewed at ad delivery, before any application is screened, with the platform and date limits stated.
- 3. Evidence from Field Experiments in Hiring Shows Substantial Additional Racial Discrimination after the Callback academic.oup.com Supports the callback-versus-offer comparison used to argue that a clean screen licenses no conclusion about the later stages.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.