Pipeline
Compute Selection Rates Per Gate: The Total Hides the Bad Gate
Compute a selection rate at every hiring gate that removes a candidate. The denominator at each gate is who entered that gate, not everyone who applied, and the ratio compares each group's pass rate to the highest group's pass rate at that same gate. A pooled end-to-end number can read clean while one gate fails, and the US Supreme Court held in 1982, in Connecticut v. Teal, that a clean bottom line is no defense.
The takeThe per-gate table is the deliverable and the headline ratio is a byproduct nobody should print on its own. Most teams find they cannot build the table: counts exist for the stages the software owns and evaporate at the moments a person made a call in a hallway, reopened a rejected profile, or forwarded a referral straight to the hiring manager. That discovery is worth more than the arithmetic it interrupted. A gate you cannot count is a gate you could not have fixed even if a ratio had told you it was broken.
Where Olive fits
Open a role and see what the work shows
Olive is priced per attempt rather than per seat, and one attempt returns six evidenced findings about one candidate: an input the hiring team weighs, with the decision staying exactly where it was. Ten attempts a month are free, so a pilot can run beside the current round.
Rank your shortlistWhich gates do you compute a rate on, and against what denominator?
Every point where a person or a tool removes someone is a gate, and each gate gets its own rate: the number of people from a group who passed it, divided by the number from that group who entered it. Total applicants is the right denominator only at the first gate. Use it everywhere else and every later rate is diluted by people who never reached that step.
A working gate list for most processes, in order:
1. Sourcing and ad delivery. Who saw the posting at all. This one sits outside the applicant tracking system and usually outside the analysis. 2. Application completion. Who started and did not finish. 3. Resume or application review. The heaviest filter in almost every funnel. 4. Recruiter screen. 5. Assessment, work sample or take-home, where one runs. 6. Interview loop, counted per round rather than as one block. 7. Debrief and offer decision. 8. Offer acceptance, which is a candidate decision and belongs in a separate column.
Before any of it computes, settle who counts as an applicant at each gate and write the definition down. Duplicate applications, candidates who withdrew, sourced profiles who never applied and people rejected by an automated knockout question all land differently, and a definition chosen after seeing the results is not a definition.
Can a clean total hide a failing gate?
Yes, and a clean total is not a defense. In Connecticut v. Teal, decided in 1982, the overall promotion rate favored Black candidates, and the US Supreme Court still held that the written test gating the eligibility list could be challenged on its own: a nondiscriminatory bottom line neither precludes a prima facie case nor provides a defense to one 1. Title VII protects the individual stopped at the gate, not the average of everyone who applied.
The federal selection guidelines do carry a bottom-line paragraph, and it is narrower than it is usually quoted as being. It says that where the total selection process shows no adverse impact, the enforcement agencies will in usual circumstances not expect a user to evaluate the individual components, in the exercise of their administrative and prosecutorial discretion, and the same paragraph then lists circumstances where they will expect exactly that 2. That is a statement about where agencies spend attention. Teal is the statement about liability, and the two are often traded for each other by whoever built the dashboard. None of this is legal advice; what a given row means for your own process is a question for counsel.
The arithmetic reason is duller and it is the one that will bite you. Rates that move in opposite directions cancel when they are pooled. In an audit that sent fictitious applications to 108 large US employers, men and women were contacted at the same average rate, while the standard deviation of gender contact gaps between companies was 2.7 percentage points, with a distribution roughly symmetric about zero: some firms consistently favored men, others consistently favored women, and the market average recorded neither 3. What that study found between employers, a pooled funnel does between its own stages.
Here is the shape, with round numbers chosen to make it visible.
| Gate | Group A in / passed | Group B in / passed | Ratio at that gate |
|---|---|---|---|
| Resume screen | 400 / 100 (25%) | 200 / 30 (15%) | 0.60 |
| Phone screen | 100 / 60 (60%) | 30 / 24 (80%) | 0.75 |
| Onsite loop | 60 / 12 (20%) | 24 / 6 (25%) | 0.80 |
| End to end | 400 / 12 (3%) | 200 / 6 (3%) | 1.00 |
The end-to-end ratio is exactly 1.00. The resume screen is at 0.60, and it is the gate that decided almost everything. Notice too that the group with the highest pass rate changes from row to row, so each row's ratio is computed against whichever group leads at that gate. Fix one comparison group for the whole table and every row where it is not the leader comes out wrong.
Map the gates before you compute anything
Build the map before the spreadsheet. List every removal point in order, including the ones outside the applicant tracking system, then write down two counts per gate per requisition: how many entered and how many came out. The regulation asks for less, and by group: the federal selection guidelines (29 CFR 1607, adopted 1978) expect records disclosing the impact a user's selection procedures have by identifiable race, sex or ethnic group, with safeguards against their misuse 2.
Three gates go missing from almost every map. The first is upstream of everything: an audit that ran paired ads for similar roles at companies with different gender mixes in their workforces found gender skew in one platform's delivery of those ads that differences in qualification did not account for, and failed to find that skew on a second platform, where the trials were small enough to leave the null weak 5. That is who was shown the posting, before an application existed, and it is measured on ad platforms in 2020 and 2021 whose delivery has changed since. The second is the referral shortcut, where a candidate skips two gates and the counts never record that they did. The third is the reopen: a rejected profile pulled back into the process by a hiring manager, which is a selection decision with no gate attached to it.
Attention should follow weight rather than spread evenly. In Ashby's benchmark set of more than 54 million applications, recruiter screens passed roughly 35% of the candidates who reached them, while post-onsite stages converted at 95% and offer stages at 81% 4. Those are passthrough rates among people who got that far, they do not multiply into anything, and each customer configures its own stage labels, so the categories are not standardized across the set. The direction is the useful part: the late gates mostly ratify a decision the early ones already made. The early gate is also where the least evidence is written down, which is why the funnel metrics you had stopped meaning what they meant and why a resume screen can throw away the wrong people invisibly.
The deliverable is a table with one row per gate and one column per group, showing counts in, counts out, the rate, and the ratio against the leading group. The ratio itself is defined in what the four-fifths rule actually is, and the arithmetic there does not change here. What changes is where you point it.
What to do about a gate too small to read
Say so in the output. A ratio computed on four people is not a small result, it is not a result, and printing one beside three honest numbers gives every number on the page the same standing. Write the counts, write one sentence naming why no ratio appears for that row, and move the attention to the gates where the counts are large enough to carry a conclusion.
The regulation anticipated this and offers one narrow route. Where a user's evidence indicates adverse impact but rests on numbers too small to be reliable, it says evidence of the procedure's impact over a longer period, or of its impact when used in the same manner in similar circumstances elsewhere, may be considered 2. The condition is doing real work: pooling three years of a hiring loop that was redesigned twice produces a bigger number about nothing. Pool only where the procedure was genuinely the same, and record which version you pooled.
A cheap test before any row is published: move one candidate across the line and recompute. If the verdict changes, that row is noise, and what to do when the whole process is that small is a different exercise from this one. The same applies to rows where most candidates never answered the demographic question, since an impact ratio is only as readable as its response rate.
The gate that resists the exercise hardest is usually the one worth looking at first. A step where a recruiter or a tool produced a number and the reasoning behind it was never written down cannot be re-read, so a bad ratio there leaves you with nothing to inspect and nothing to change. A step where each decision carries the criterion it was measured against and the evidence that satisfied it can be opened and argued with, one candidate at a time, which is what turns a ratio into something you can act on.
Common questions
What counts as a gate?
Any point where somebody leaves the process because of a decision. That includes automated knockout questions, a recruiter reading a resume, an assessment cut score, an interview round, and the debrief. It also includes decisions nobody logs: a hiring manager pulling a rejected profile back in, a referral skipping two rounds, a sourced candidate entering at the loop. If a decision changes who remains, it is a gate, and the fact that no software recorded it is a finding rather than an exemption.
Should the denominator ever be total applicants?
Only at the first gate, where everyone who applied is also everyone who entered. After that, the denominator is whoever reached that step. Using total applicants for a late gate mixes two effects, the pass rate at that gate and the composition of who survived to reach it, and no single number can separate them afterward. Keep both figures if you want them: the per-gate rate for diagnosis, and the end-to-end rate for context, clearly labelled as different quantities.
Do you compute this per requisition or across the year?
Per selection procedure, which usually means per gate per role family rather than per requisition. Two requisitions that ran the same rubric with the same screeners are the same procedure and can be pooled. Two that used different assessments are not, and pooling them measures a mixture. Record which requisitions went into each row, because a year from now nobody will remember which loop was in force, and an unlabelled pooled number cannot be re-derived.
Does a passing end-to-end number protect you?
No. The US Supreme Court held in 1982, in Connecticut v. Teal, that a nondiscriminatory bottom line neither precludes a prima facie case nor provides a defense to one, because the protection runs to the individual who was stopped at a particular step. The federal guidelines' bottom-line paragraph describes how enforcement agencies allocate their own attention; it says nothing about whether a component can be challenged. Treat a clean total as a reason to look at the gates, not as a reason to stop.
Which groups do you compute rates for?
The federal recordkeeping guidelines name sex and a specific list of race and ethnic groups consistent with the EEO-1 categories, and totals. Compute every cell you have counts for rather than only the comparison somebody asked about. Watch the intersections: a process can clear every single-attribute check while a combination fails, and a table built one attribute at a time is structurally unable to show it. Where a cell is too small, print the counts and no ratio.
References
- 1. Connecticut v. Teal, 457 U.S. 440 (1982) law.cornell.edu Supports the claim that a clean end-to-end result is no defense: promotions ran 22.9% for Black candidates against 13.5% for white candidates, and the Court still held the bottom line neither precludes a prima facie case nor provides a defense.
- 2. 29 CFR 1607.4 - Information on impact (Uniform Guidelines on Employee Selection Procedures, 1978) law.cornell.edu Supports three claims: the recordkeeping expectation by race, sex and ethnic group in 4(A) and its safeguards; the bottom-line prosecutorial discretion in 4(C); and the clause in 4(D) allowing evidence over a longer period where numbers are too small to be reliable.
- 3. Systemic Discrimination Among Large U.S. Employers nber.org Supports the cancellation claim: no average gender contact gap across 108 large employers, with a between-company standard deviation of 2.7 percentage points roughly symmetric about zero.
- 4. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the claim that early gates carry the filtering weight: recruiter screens pass around 35% of candidates who reach them, against 95% post-onsite and 81% at offer, across more than 54 million applications.
- 5. Auditing for Discrimination in Algorithms Delivering Job Ads arxiv.org Supports the claim that a funnel can be skewed before any application exists: gender skew in one platform's job-ad delivery that qualification differences did not account for, and a failure to find that skew on the other platform tested, on smaller trials.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.