Pipeline
Too Few Hires for a Ratio: Measure the Process Instead
At fifteen hires a year a selection-rate ratio is not a usable number: one candidate moving across the line swings it further than any bias worth finding. Recompute the ratio with one candidate moved each way before you report it, and when the verdict changes, report the counts and no ratio. What a company that size can measure honestly is the process: whether every candidate met the same written criteria, and whether each decision carries the evidence behind it.
The takeThe pressure here comes from outside: a security review, an insurance renewal, an investor asking whether hiring has been checked. Whoever asked wants a yes, and the useful answer is neither a yes nor a refusal. It is the sample, stated plainly, with the thing you did measure attached. A reviewer who reads that there were fifteen selections, that one person moving flips the verdict, and that counts are reported instead learns more about the company than a decimal would have told them.
Where Olive fits
Open a role and see what the work shows
Olive sits in the same position and says so on its own site rather than implying a pass: too few sessions have run for a four-fifths ratio to carry meaning, so no bias audit has been performed. What one attempt returns is six findings in words, each anchored to a timestamped excerpt from the session, which is a record a fifteen-person company can still re-read a year later.
Rank your shortlistRun the flip test before anything leaves the room
Recompute the ratio twice before you write it anywhere: once with a single candidate moved from rejected to selected, once with a single candidate moved the other way. If either move changes the verdict, the number is a property of this year's applicant pool, and it belongs in no document. This takes about two minutes and it settles the question the rest of the analysis is arguing about.
Work it through. Twenty candidates from one group reached the final gate and six were hired, a rate of 30%. Fifteen from another group reached it and three were hired, a rate of 20%. The ratio is 20 over 30, or 0.67, which fails the four-fifths screen and reads like a finding. Now move one candidate across: four hires out of fifteen is 26.7%, the ratio becomes 0.89, and the finding has disappeared without anything about the company changing.
Read a flip as a statement about the instrument. It cannot answer the question at these counts, which is worth establishing before anyone goes looking for a cause. The federal agencies that wrote the four-fifths rule described the same move in their 1979 interpretive questions and answers: it is generally inappropriate to require validity evidence or to take enforcement action where the number of persons and the difference in selection rates are so small that selecting one different person for one job would shift the result from adverse impact against a group to that group holding the higher selection rate 2. Their condition is narrower than the flip test above, so anything they would set aside fails the flip test too. Two responses follow, and both beat publishing the decimal: report the counts with no ratio attached, or pool a longer run of the same procedure, with the limits in the next section.
Why does a small-sample ratio move so much?
Because the denominator is in the low dozens, the rate moves only in large jumps. One person moving changes a rate computed on fifteen candidates by nearly seven percentage points, and that smallest possible step is larger than most of the effects anyone is hunting for. The 1978 Uniform Guidelines, the US federal regulation carrying the four-fifths rule, set a limit on small-number evidence in the paragraph everybody quotes for the eighty percent figure.
Section 1607.4(D) states that greater differences in selection rate may not constitute adverse impact where the differences are based on small numbers and are not statistically significant. Where the evidence indicates adverse impact but rests on numbers too small to be reliable, the same paragraph allows impact over a longer period of time, or the impact of the same procedure used in the same manner in similar circumstances elsewhere, to be considered in determining adverse impact 1. That is a route to better evidence rather than an exemption, and a longer run can confirm a pattern as easily as it dissolves one. Pooling three years of a hiring loop you redesigned twice produces a larger number about nothing.
The 1979 answers go further: they call the four-fifths figure a rule of thumb, "not intended as a legal definition," but a practical means of keeping enforcement attention on serious discrepancies 2. Asked directly whether the rule means the guidelines tolerate up to 20% discrimination, the agencies answered no: it speaks only to the question of adverse impact and is not intended to resolve the ultimate question of unlawful discrimination 2. A number that was never a threshold is a strange thing to compute on single-digit counts.
When the counts are large enough that the flip test holds, a second question opens up, which is whether the gap is big or merely detectable. That is the subject of the ratio and the significance test answering different questions, and at this size it does not arise yet.
What can fifteen hires actually measure?
The process, because each of those fifteen hires came out of a chain of decisions and the decisions are the observations. A company hiring fifteen people a year runs several hundred of them: screens, rejections, interview ratings, debriefs. Those are countable, they are what the ratio was standing in for, and they are where an employer this size has enough data to see something.
Four questions, each with a countable answer:
1. Was the bar written down before anyone read an application? Count the requisitions where written criteria exist with a timestamp earlier than the first review. 2. Does every rejection carry a reason tied to those criteria? Count the rejections whose recorded reason names one of those criteria. 3. Did every candidate at a gate face the same questions and the same exercise? Count the loops where the question set changed mid-process. 4. Which gate has no record at all? It is usually the earliest one.
The fourth question is the one that pays. In one applicant-tracking vendor's dataset of over 54 million applications, the share of expected scorecards actually submitted runs near 49% at organizations under 25 employees, against closer to 72% at organizations of 500 or more 3. The measure is scorecards submitted inside that vendor's product, by its own customers, and completion is higher for candidates who get hired, so the thinnest part of the record is the rejections, which is what a fairness question is about. The same data shows around 38% of scorecard pairs carrying at least a one-point difference between interviewers, with nearly half of those differences sitting between 2 and 3 and crossing the yes/no line on a four-point scale 3. That is disagreement rather than error: the dataset carries no demographic variable and says nothing about any protected group, and it is a good reason to keep the evidence a rating rests on, so a later reader can see what the call was made on.
The field evidence points back at your own records too. Across 108 large US employers, a correspondence experiment on entry-level vacancies found contact-rate gaps against Black applicants concentrated in a minority of firms, with the top quintile of discriminating firms responsible for nearly half the lost contacts and a Gini coefficient around 0.4 4. The authors could not reject that all 108 weakly favored white names, so concentration is not innocence for the rest, and an industry average is the wrong instrument for any one company. If your own records barely exist, the smallest process a very small company can actually defend is the place to start.
Write the sample size into the sentence
State the counts first, then what they can and cannot support. The sentence a security reviewer or an insurer can use names the population, the number of selections and the consequence, and it beats a ratio because it can be checked. Anyone can work out that a decimal computed on single-digit counts is decoration, and almost nobody will say so on your behalf.
A usable form, filled in with your own figures:
> Across last year's twelve requisitions, 34 candidates reached the final gate and 15 offers were made. Selection counts by group are in the appendix. No impact ratio is reported: at these counts, one candidate moving across the line changes the verdict, so a ratio would describe the sample rather than the process. What is reported instead is criteria coverage, rejection-reason coverage and gate-record completeness, per requisition.
Two temptations to name and refuse. The first is raising the count by making the demographic question mandatory, which converts a voluntary instrument into a compelled one and produces answers people did not want to give, and the response-rate problem it is trying to solve has a bounding test that works better. The second is calling the result a bias audit, which is a specific thing with a specific scope, and the difference between a validated assessment and a bias-audited one is exactly the distinction a reviewer will test if they know the field. Whether you owe an audit at all is a question for counsel, and it turns on where you hire and what the tool does.
The question will be asked again next year, by somebody new, with the same expectation of a yes. What survives that repetition is the record of how each decision was made, which grows by several hundred rows a year even at fifteen hires. A ratio you could not defend the first time will not have improved by being asked for twice.
Common questions
How many hires do you need before a ratio means anything?
The 1978 Uniform Guidelines name no minimum, and the federal agencies that wrote them answered the small-numbers question with a test rather than a number. In their 1979 interpretive answers: where the number of persons and the difference in selection rates are so small that selecting one different person for one job would shift the result from adverse impact against a group to that group holding the higher rate, it is generally inappropriate to require validity evidence or to take enforcement action. A lower rate that persists over time is a different matter. Run the wider version first, moving one candidate each way, and if the verdict changes the ratio is describing your sample.
Can you pool several years together to get a usable sample?
Only where the same selection procedure was used in the same manner across those years. The 1978 Uniform Guidelines allow impact over a longer period to be considered in determining adverse impact where the numbers are too small to be reliable, so a longer window is as likely to show a pattern as to clear one, and a loop that changed its assessment, its rubric or its screeners is not one procedure. Record which requisitions and which version of the process went into each pooled row, because nobody will remember a year later, and an unlabelled pooled number cannot be re-derived.
Should you report a ratio you already know is unstable?
No. Publishing it creates a written record of a conclusion the data cannot support, in either direction, and you will be held to it by whoever reads it next. Report the counts, name the sample size, and state the consequence in one sentence. That is not a weaker answer than a decimal. It is the same information with the uncertainty visible, which is what makes it usable by a reader who knows what they are looking at.
Is a small employer exempt from any of this?
Coverage thresholds vary by statute and by state, and that question belongs with counsel who can see your headcount and your locations. The practice question is separate from coverage and does not wait for it: whether every candidate at a gate met the same written bar, and whether the reason each one was turned down was recorded at the time, are answerable at any size. An employer below a reporting threshold still makes the decisions the threshold was written about.
What do you say when the answer has to be yes or no?
Say what was measured and what was not, in that order, and let the reader draw the line. "Selection counts by group are reported for the year; no impact ratio is computed, because at 15 selections a single decision changes the result" is an answer, not a dodge. Follow it with the process measures you do have. A reviewer whose job is assessing risk can read the difference between a company that looked and a company that produced a number.
References
- 1. 29 CFR 1607.4 - Information on impact (Uniform Guidelines on Employee Selection Procedures, 1978) law.cornell.edu Supports both small-sample claims in 4(D): that greater differences may not constitute adverse impact where they rest on small numbers and are not statistically significant, and that where small-number evidence indicates adverse impact, a longer period or the same procedure used in the same manner elsewhere may be considered in determining adverse impact.
- 2. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures (Q.11, Q.19) eeoc.gov Supports three claims: that the agencies call the four-fifths figure a rule of thumb rather than a legal definition; their answer that it does not tolerate 20% discrimination or resolve the ultimate question of unlawful discrimination; and their small-numbers answer that it is generally inappropriate to require validity evidence or to take enforcement action where selecting one different person for one job would reverse the result, with the caveat that a lower rate persisting over time does constitute adverse impact.
- 3. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the scorecard completion gap by employer size, measured as the share of expected scorecards submitted inside that vendor product (near 49% under 25 employees against closer to 72% at 500 or more), and the finding that around 38% of scorecard pairs differ by at least a point, with nearly half of those crossing the 2 to 3 boundary.
- 4. Systemic Discrimination Among Large U.S. Employers nber.org Supports the concentration claim across 108 large US employers, including the Gini coefficient of about 0.4, and the authors' caveat that they could not reject weak favoring of white names at every firm.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.