Policy

Run Both: The Ratio Sizes the Gap, the Test Says Whether It Is Real

Judge a hiring disparity on both the selection-rate ratio and a significance test, per gate. The ratio is an effect size: it says how big the gap is. The test says whether a sample this size could tell that gap from chance. They disagree routinely, and your applicant volume decides which of the two misleads you first, which is why neither is the advanced version of the other.

The takeCompliance explainers present these as tiers, the ratio for everybody and significance testing for large employers with an analyst on staff. That ordering is backwards at both ends of the range. A high-volume funnel needs the test because its ratio goes quiet on gaps the enforcement agencies say still count, and a low-volume one needs the test because its ratio shouts at noise. The employer who could safely use only one number is the mid-sized case nobody writes the guide for.

Where Olive fits

Open a role and see what the work shows

Neither statistic explains why a gap opened, and a number standing for a person cannot be asked. Olive returns six findings written by a human reviewer, each carrying the timestamped excerpt it rests on, so what a report claims about a candidate can be reopened and argued with rather than only recomputed.

Rank your shortlist

Which question does each number answer?

The ratio answers how big, the test answers how sure, and a decision needs both. An impact ratio of 0.6 says one group passed at three-fifths the rate of the leading group, whatever the counts behind it. A significance test says whether a difference that size, at these counts, can be told apart from ordinary sampling variation. Neither statement contains the other.

The regulation that gave everyone the ratio names all of this in a single paragraph. Smaller differences in selection rate may nevertheless constitute adverse impact where they are significant in both statistical and practical terms, and greater differences may not constitute adverse impact where they are based on small numbers and are not statistically significant 1. Read the phrase everybody skips: statistical and practical terms. There are three questions in that sentence, not two, and the third one has no formula, which is why it goes missing from the spreadsheet.

Practical significance is the question of how many actual people the gap represents at this gate, and whether a difference that size changes anyone's working life. A one-point gap in a group of forty thousand is four hundred people who did not pass a step they would have passed at the leading group's rate. A twenty-point gap in a group of forty is eight people. Both facts are worth knowing and neither is contained in the ratio or the test.

What the ratio is and how it is computed is settled ground, and the four-fifths rule as a trigger for scrutiny covers the arithmetic and its limits. This article is about the choice you make once you are holding both numbers and they point different ways.

How does your volume decide which one misleads you?

Volume changes what each number can see, in opposite directions. With forty applicants at a gate, a ratio of 0.5 and a coin flip are indistinguishable, and the test will say so. With forty thousand, a one-point difference in pass rates clears a conventional significance threshold while the ratio still reads 0.95. Same two instruments, opposite failure modes, and your funnel decides which one you meet first.

The low-volume case, worked. Thirty candidates from one group reach a gate and twelve pass, a rate of 40%. Ten from another group reach it and two pass, a rate of 20%. The ratio is 0.5, which fails the four-fifths screen and looks alarming. A two-proportion test on those counts comes nowhere near a conventional threshold, and moving a single candidate across takes the ratio to 0.75. The ratio was not wrong. It was reporting an effect size computed on ten people, which is what it is designed to do and not what anyone reads it as.

The high-volume case is the mirror. Twenty thousand candidates from each of two groups reach a gate, one passing at 21% and the other at 20%. The ratio is 0.95, comfortably clear of the screen, and a two-proportion test on those counts is significant at the five percent level. The gap is real and it is small, and the practical question is that one percentage point is two hundred people at this gate. That is the point where a team decides whether it has found an emergency, a research question, or a rounding artifact of how the stage is configured.

Neither case is exotic. A recruiter screen at a company with a busy job board sits in the second one, and the final gate of the same funnel sits in the first, in the same week, in the same table. Which is why the choice belongs to each gate on its own, and computing selection rates gate by gate is the structure that makes it possible.

Read the disagreement rather than picking a winner

Four combinations, and each has a different next move. Write the pair down per gate with the counts beside them, and the pair usually tells you what kind of problem you are holding before anyone argues about method. The temptation is to report whichever number supports the conclusion already reached, which is why the choice of method belongs in writing before the analysis runs.

1. Ratio passes, test significant. High volume, small real difference. Not an emergency, and not nothing. Ask what is different about the gate, and size it in people. 2. Ratio fails, test not significant. Low volume. The measurement cannot see anything yet; recomputing with a single candidate moved across will usually confirm it, and what to measure when the whole process is that small matters more than the decimal. 3. Both fire. The arithmetic was built for this case. What to do about the gate is a question for counsel on your own process. 4. Neither fires. Keep the record and keep looking. A pass is not a clean bill of health, and the agencies who wrote the ratio said so themselves: it is a rule of thumb, "not intended as a legal definition," and it "speaks only to the question of adverse impact, and is not intended to resolve the ultimate question of unlawful discrimination" 2.

Where the number came from is worth knowing. A 2024 peer-reviewed paper tracing the rule's provenance found its earliest appearance in California regulatory guidance in 1972, six years before the federal guidelines, and could locate no official written justification for the value four-fifths; the only account the authors found was a recollection that the test "was born out of two compromises," one of them a way to split the middle between a 70% camp and a 90% camp 3. That anecdote is labelled as anecdotal by the authors, whose argument is aimed at computer scientists who turned the rule into a fairness metric. It does not make adverse-impact law arbitrary. It does mean that treating 0.8 as a scientific boundary, and a significance test as its rigorous upgrade, misreads both.

What neither number can tell you

Why the gap opened. Both numbers are summaries of outcomes, and no summary of outcomes contains a mechanism. That matters legally as well as practically: the statutory test is neither number. Under Title VII since 1991, a practice with disparate impact stands or falls on whether it is job related for the position in question and consistent with business necessity, and on whether a less discriminatory alternative was refused 4. Age and disability claims run under different statutes and standards.

The same statute carries a clause worth reading twice. A complaining party must normally identify the particular practice causing the impact, unless the elements of the decision-making process are not capable of separation for analysis, in which case the process may be analyzed as one practice 4. How far that reaches for an opaque scoring tool is unsettled rather than decided, and the direction is uncomfortable: opacity is not a shield.

Where the gap comes from is at least partly a question about how the evaluation was made. A large resume audit released as a preprint in 2026, covering 36,880 applications to 9,220 advertisements for new US college graduates, reports callback gaps varying with the task content of the job, with callbacks in management occupations 28 to 43 percent lower for Black men, Black women, White women and Hispanic men than for otherwise identical White men, and the widest gaps in roles combining high analytical and interpersonal demands with low routine content 5. It has not completed peer review, the discretion mechanism is a model the authors propose and the experiment does not manipulate it, and the outcome it measures is a callback, with offers and hires unobserved. Read at that strength, it still points somewhere useful: gaps concentrate where the judgment is loose.

Which is why the composition of a gate matters more than the statistic printed under it. A step that produced a number with nothing attached leaves you two numbers and no way to ask a third question, because there is nothing to reopen: no criterion, no excerpt, no record of what the decision rested on. A step whose decisions carry the evidence behind them can be read back one candidate at a time, and a failing ratio there becomes a list of specific rejections somebody can examine. Decide with counsel what the analysis will produce before it runs, since the number that lands in a document is the number that gets read back later, and what an assessment vendor has to hand over in an EEOC inquiry is the same question asked from the other side.

Read the evidence

Common questions

Which number goes in a report to a board or a customer?

Both, per gate, with the counts they were computed from and the response rate for the demographic data underneath them. A single decimal with no denominator is the thing that gets quoted back inaccurately a year later. If space forces one line, write the counts and the gate rather than the ratio, because counts can be recomputed into any statistic a reader wants and a lone ratio cannot be recovered into anything.

Is 0.8 a legal threshold you have to clear?

No. The federal agencies that wrote it called it a rule of thumb rather than a legal definition, and said explicitly that it speaks only to whether adverse impact is present, not to whether discrimination occurred. Clearing 0.8 is not a defense and failing it is not a finding. It marks where enforcement attention has historically been aimed, which makes it useful as a first screen and useless as a target to manage toward.

Which significance test should you actually run?

For a two-group comparison at one gate, a two-proportion z test or a chi-square test is the standard choice, and Fisher's exact test is the usual option when the counts are small. The specific choice matters far less than two disciplines around it: run the same test every time so results are comparable across quarters, and record which test was run alongside the result. A statistic whose method nobody wrote down cannot be reproduced or defended.

What does practical significance mean here?

How many real people the gap represents at that gate, and whether a difference that size changes what happens to them. It has no formula, which is why it is usually skipped, and it is the question that keeps a statistically significant one-point difference in a high-volume funnel in proportion. Convert every gap into a headcount before deciding what it means: a percentage describes a rate, and a rate has never once been the thing anyone was actually worried about.

If both numbers pass, is the gate fine?

It is not flagged, which is a weaker statement. Both instruments measure outcomes for the groups you have data on, at the volumes you have, using the categories people chose to disclose. A gate can clear both while excluding a subgroup too small to appear, or while resting on a criterion nobody can justify as job related. Keep the record, note the sample, and treat a clean pair as a reason to look at the criterion rather than as a reason to stop looking.

References

  1. 1. 29 CFR 1607.4 - Information on impact (Uniform Guidelines on Employee Selection Procedures, 1978) Code of Federal Regulations, via Cornell Legal Information Institute, 1978. law.cornell.edu Supports the claim that the regulation itself names statistical and practical significance alongside the ratio, in both directions: smaller differences may still be adverse impact, and larger differences may not be where numbers are small.
  2. 2. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures (Q.11, Q.19) U.S. Equal Employment Opportunity Commission (joint EEOC-DOJ-OPM-DOL-Treasury document, OLC Control Number EEOC-NVTA-1979-1), 1979. eeoc.gov Supports the quoted agency statements that the four-fifths figure is a rule of thumb rather than a legal definition and does not resolve the ultimate question of unlawful discrimination.
  3. 3. The four-fifths rule is not disparate impact: A woeful tale of epistemic trespassing in algorithmic fairness Elizabeth Anne Watkins and Jiahao Chen, in Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24), 2024. facctconference.org Supports the provenance claim: earliest mention traced to California regulatory guidance in 1972, no official written justification for the value, and the anecdotal account of a compromise between a 70% camp and a 90% camp.
  4. 4. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases Office of the Law Revision Counsel, United States Code (prelim), 1991. uscode.house.gov Supports the statutory test of job relatedness for the position in question and business necessity, the less-discriminatory-alternative route, and the clause covering elements not capable of separation for analysis.
  5. 5. Hiring Discrimination and the Task Content of Jobs: Evidence from a Large-Scale Resume Audit arXiv preprint arXiv:2604.01933 (Braun and co-authors), version 2, July 2026, 2026. arxiv.org Supports the claim that measured callback gaps widen where evaluation is subjective: 36,880 applications to 9,220 advertisements, callbacks 28 to 43 percent lower in management occupations for four applicant groups, widest where routine content is low.

5 sources, numbered by first appearance. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.