Policy
An Outcome Gap Names a Stage to Inspect, Not a Verdict
One group passing at a lower rate does not prove the hiring process is biased. A selection rate is computed over the people who reach a stage, so a rule applied identically to everyone still reports a gap when different people arrive at it. What a gap licenses is an inspection of the rule at one boundary. Three separate things produce the same number: a different pool arriving, a stage rule tracking something not job-related, and a sample too small to carry a ratio.
The takeBoth stock answers are unfalsifiable, which is why they survive. One treats any group difference as proof of bias; the other treats an impact ratio above four-fifths as a clean bill of health. Neither will say what a gap that is not the procedure's fault would look like, so neither can be shown wrong by any real funnel. Write your own version down before this quarter's numbers arrive: the smallest cell you will read, and the ratio that would make you open a rule.
Where Olive fits
Open a role and see what the work shows
Olive has not completed a bias audit, and the site says so: at the volume it runs, a selection-rate ratio would rest on cells too small to read, which is the same arithmetic this article applies to a funnel. What it does return is six findings per candidate, each carrying the timestamped excerpt the human reviewer wrote it from.
Rank your shortlistWhat does a lower pass rate actually prove?
On its own, very little about the rule. A selection rate is computed over the people who reached the stage, so two teams applying an identical rule report different rates when the people arriving at it differ. The gap is still real and still worth acting on. What it identifies is a place to look. The thing to look at is the rule that operates at that boundary.
The 1978 federal regulation everyone reaches for says as much in its own text. A selection rate for any race, sex or ethnic group below four-fifths of the rate for the highest group "will generally be regarded by the Federal enforcement agencies as evidence of adverse impact," and the same paragraph immediately adds that smaller differences may nevertheless be adverse impact where they are significant in both statistical and practical terms, and that larger differences may not be adverse impact where they rest on small numbers and are not statistically significant 1. Evidence, generally, with caveats running in both directions. That describes the point at which an agency starts asking questions.
Two arithmetic points get confused constantly and both change the answer. The ratio is one group's selection rate divided by the highest group's rate, not a difference in percentage points, so a twenty-point gap and a ratio under 0.8 are different quantities that happen to be discussed in the same meeting. And the ratio belongs to a single selection procedure, so a number computed from applications all the way to offers is a blend of every rule in the funnel. The four-fifths arithmetic itself is worth reading once before anybody quotes a figure derived from it.
Why doesn't an impact ratio above four-fifths settle it?
Because the agencies that wrote the rule said it does not. Their 1979 interpretive questions and answers call the four-fifths test a rule of thumb "not intended as a legal definition," a practical way of keeping enforcement attention on serious discrepancies, and answer directly that it "speaks only to the question of adverse impact" rather than resolving whether anything unlawful occurred 2.
The number itself has a history worth knowing before anyone treats it as a threshold. A peer-reviewed history traces the earliest appearance of the four-fifths rule to California regulatory guidance in 1972, six years before the federal Guidelines, and reports finding no official written justification for the value anywhere; the only account they located is a recollection that the test "was born out of two compromises," one of them a way to split the middle between a 70% camp and a 90% camp 4. The underlying law is not arbitrary. The line is an administrative convenience, and treating it as a pass mark inverts what its own drafters said it was for.
An aggregate can also clear four-fifths because gaps running in opposite directions cancelled inside it. Kline, Rose and Walters sent more than 83,000 fictitious applications to entry-level vacancies at 108 large US employers and found no average gender gap in employer contact at all, alongside a between-company standard deviation of 2.7 percentage points that was roughly symmetric about zero: some firms consistently favoured men, others consistently favoured women, and the two cancelled 3. An average of zero there is the sum of discrimination running in opposite directions.
The same experiment found the racial contact gap concentrated in a minority of firms, with the top quintile of discriminating firms accounting for nearly half of all contacts lost to Black applicants 3. The measure there is employer contact within 30 days, at unusually large employers, on entry-level roles, using fictitious applications and names to signal group membership, and the authors could not reject the possibility that every firm in the sample weakly favoured white names, so concentration is not a clean bill for anybody else. What the design does establish is that behaviour varies enormously between employers, which is exactly why an industry statistic cannot answer a question about your funnel.
Compare stage rates before choosing a remedy
Compute selection rates one stage at a time, and put the entry counts and the composition of the entrants beside them. An end-to-end rate blends every rule in the funnel into a single number, which is why it moves for reasons nobody touched and why it can hide a bad stage behind three good ones. Stage rates are conditional on who reached the stage, so the entry column is what makes them readable at all.
1. One row per stage, including the ones that feel mechanical. Source and application review belong in the table even though nobody schedules them. 2. Four columns: entered, composition of entrants, passed, pass rate. The composition column is the one that does the analytical work. 3. A fifth column marking cells below your stated minimum. Decide that minimum before you look. 4. Take the largest readable stage gap first. Not the largest gap, and not the stage where the argument started. 5. Open the rule at that boundary. Announcing a remedy is the step after, and it needs a named rule to attach to.
Step four is harder than it looks, because the stage displaying a gap is frequently inheriting it: which stage a gap actually opens in works through the attribution problem in detail. And if every stage rate moved at once for reasons unconnected to any group, the prior question is which funnel metrics still mean anything, because a table built on numbers that stopped measuring what they used to measure will produce a confident wrong answer.
Which cells are too small to read?
Any cell small enough that one decision moves the ratio. A stage holding eleven candidates from one group is that cell, and small cells are where the false alarms live. Decide the number before you see the data and write it down. The four-fifths paragraph itself allows that larger differences may not be adverse impact where they rest on small numbers and are not statistically significant 1.
Two habits keep small cells honest. Mark them unreadable rather than deleting them, because a deleted row silently becomes a claim that the stage was fine. And decide in advance whether you will pool quarters to reach a readable cell, because pooling after seeing the ratio is choosing the window that returns the answer you already had.
The other reason a gap can be unreadable is that the data sits somewhere else. When the stage in question runs inside a vendor's system, the selection rates you need may not be in your applicant tracking system at all, and running an adverse impact audit when the vendor holds the data turns into a contract question long before it turns into an analytics one.
What none of this licenses is the reverse move. A ratio above four-fifths is not evidence that a stage is fine, a stage nobody can compute is not evidence of anything, and neither entitles anyone to close the question. The whole job of a gap is to point at a boundary. What happens after that is reading the rule at that boundary and asking whether each thing it uses has a job-related reason attached, which is what makes a procedure biased in the first place.
Common questions
Does a ratio below four-fifths mean we are breaking the law?
No. The four-fifths rule is a paragraph in a 1978 United States federal regulation describing when enforcement agencies will generally treat a selection rate difference as evidence of adverse impact, and the agencies' own 1979 questions and answers call it a rule of thumb not intended as a legal definition. The paragraph runs to race, sex and ethnic group, and neither the ADA nor the ADEA carries a four-fifths rule of its own. Adverse impact is one element of a disparate-impact claim rather than the claim itself. Treat a failing ratio as a reason to inspect the stage, and take the legal question to counsel.
Can a gap be caused by the applicant pool rather than the process?
A selection rate is computed only over the people who reached the stage, so a rule applied identically to everyone still reports a gap when the arriving pool differs. Rule that out first, since the entry counts that settle it are already in the stage table. Who arrives is the result of sourcing, posting and referral decisions your own team makes, so the gap does not become somebody else's problem. It relocates the fix from the stage that shows the gap to the stage that shaped the pool.
How small is too small to read?
Small enough that one decision moves the ratio, which for most in-house funnels means anything under a couple of dozen in the smallest cell. Rather than argue the exact threshold, pick one, write it down before you look, and apply it to every stage. The threshold matters much less than deciding it in advance, because a number chosen after seeing the data is the analyst's preference wearing a statistical costume.
Should we run a significance test instead of the ratio?
A test answers a different question: whether a difference this large would be surprising if the rule were even-handed. It is a useful companion to the ratio rather than a replacement, since a large sample makes trivial differences significant and a small one hides real ones. Report both alongside the cell sizes, and let either one send you to a stage to open the rule.
What if the gap disappears when we look stage by stage?
Then the end-to-end number was blending stages, and knowing that is worth the exercise. Aggregates can also run the other way, hiding one bad stage behind three good ones, so a clean total is not evidence either. Keep the stage table as the standing artifact and the total as the headline it summarizes. If no stage shows a readable gap, record the cell sizes that made it unreadable rather than recording an all-clear.
References
- 1. 29 CFR 1607.4 - Information on impact (Uniform Guidelines on Employee Selection Procedures, 1978) law.cornell.edu Supports the text of the four-fifths paragraph and its own caveats in both directions, including that larger differences may not be adverse impact where numbers are small.
- 2. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures (Q.11, Q.19) eeoc.gov Supports the claim that the agencies themselves call the four-fifths test a rule of thumb that resolves neither adverse impact nor unlawful discrimination.
- 3. Systemic Discrimination Among Large U.S. Employers nber.org Supports the bidirectional gender result that cancels in aggregate, and the concentration of the racial contact gap among a minority of employers.
- 4. The four-fifths rule is not disparate impact: A woeful tale of epistemic trespassing in algorithmic fairness facctconference.org Supports the 1972 California provenance of the four-fifths value and the recollected compromise between a 70% and a 90% position.
4 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.