Pipeline

Most Lost Callbacks Come From a Minority of Employers

Hiring discrimination is concentrated in a minority of employers, not spread evenly across them. In Kline, Rose and Walters' experiment sending more than 83,000 fictitious applications to entry-level roles at 108 of the largest US employers, the top quintile of discriminating firms accounted for nearly half of all contacts lost to Black applicants, with a Gini coefficient of employer contact gaps around 0.4. Concentration is not a clean bill of health for the others: those authors could not reject that all 108 firms weakly favored White names.

The takeAn industry average is the wrong instrument for a single company, and the Kline, Rose and Walters experiment shows precisely how it fails. Average gender gaps in it were zero, because firms favoring men and firms favoring women cancelled each other out at a between-company standard deviation of 2.7 percentage points. A board deck that opens with a market number and closes with a claim about this company has skipped the only step that matters. Run the measurement inside your own building, or say plainly that you have not run it.

Where Olive fits

Open a role and see what the work shows

Olive is priced per attempt rather than per seat, and an attempt returns six findings on one candidate, each carrying the moment in the session it rests on: an input to your decision and never a gate in front of it. Ten attempts a month are free, so a pilot can sit beside your current process without changing who gets through it.

Rank your shortlist

How concentrated is it?

A fifth of the discriminating firms carried close to half the harm. Kline, Rose and Walters sent more than 83,000 fictitious applications to entry-level vacancies at 108 large US employers: 24 percent drew employer contact within 30 days, distinctively Black names reduced the chance of contact by 2.1 percentage points, and that gap was concentrated among a minority of the firms 1.

The concentration numbers are worth stating precisely, because both the alarming and the reassuring readings of them are wrong.

  • The top quintile of discriminating firms produced nearly half the lost contacts, and the Gini coefficient of employer contact gaps came out around 0.4, which makes discrimination against Black names about as concentrated among these firms as income is among US households 1.
  • 23 of the 108 companies were identified as discriminating while holding the false discovery rate at 5 percent 1. That is a floor produced by statistical power, not a list of the only offenders, and roughly one of the 23 is expected to be a false positive.
  • The other 85 were not cleared. The authors could not reject the hypothesis that all 108 firms weakly favored White names, and they put a lower bound of at least 7 percent of all jobs in the experiment discriminating, rising to at least 20 percent of jobs at the flagged firms 1.

The design limits travel with the finding. These are very large employers, entry-level vacancies, and portals that were easy to audit. The outcome measured is employer contact within 30 days, which stops well short of an offer. Nothing in it observes an interview.

Why doesn't the industry average describe your company?

Because an average is computed across firms whose behavior points in different directions, and it deletes exactly the variation you are trying to find. The same experiment found no average gender gap in employer contact at all, and underneath it a between-company standard deviation of 2.7 percentage points, roughly symmetric about zero: some firms consistently favored men and others consistently favored women 1.

Take that as a warning about your own reporting. Any average taken over units that behave differently will cancel opposing gaps and return a number that describes none of them. Averaging across your req families will do to you exactly what averaging across firms did to the gender result there, and the detection problem compounds it: bidirectional gaps lower the probability that any single firm looks significant, so a clean report can mean low power rather than even-handed behavior.

Where the gap sits also moves with the kind of work. A large resume audit of 36,880 applications to 9,220 advertisements for new US college graduates found callbacks 28 to 43 percent lower in management occupations for Black men, Black women, White women and Hispanic men than for otherwise identical White men, with the widest gaps in roles combining high analytical and interpersonal demands with low routine content 2. That work is a 2026 preprint under review, and the discretion mechanism is the authors' proposal; the experiment does not manipulate it. Taken carefully, it points where you would expect: the less specified the evaluation, the more room there is for something other than the work to move it. That is the practical case for structure, and it is the same argument that runs through which stage the gap opens.

Cut your own numbers by req family before you cut them by company

Cut the numbers at the level where a decision actually gets made, which sits below the company. A company-wide ratio is an average over req families that hire through different processes, use different screeners and attract different pools, and it will cancel opposing gaps the same way the market average cancelled the gender result. Cut the numbers by job family and by stage first, then look at the company total to see what it hid.

1. Fix the unit before you compute anything. One req family, one stage, one time window. A number without a unit invites a comparison to an industry figure that was computed on a different unit entirely. 2. Report the denominator beside the ratio. Forty applications is not a measurement. Publishing the count alongside the rate is what stops a spreadsheet becoming a claim. 3. Keep what each decision rested on. The criterion applied, and the specific thing in the application that met or missed it. A rate tells you a difference exists; only the record tells you what produced it, and only the record can be disagreed with. 4. Write down what you did with the result. California's amended employment regulations make evidence, or the lack of evidence, of anti-bias testing relevant to a claim that an automated-decision system or a selection criterion discriminated, and to any defense of that claim, including the quality, recency and scope of the effort and the response to the results. The same amendments extend the employment-records retention period from two years to four, naming selection criteria and automated-decision system data expressly 4. Testing and then doing nothing is the position that reads worst. That is California regulation, effective October 1, 2025, and how far it reaches your own process is a question for employment counsel.

When the tool doing the screening belongs to a vendor and the data sits on their side, the mechanics change and the questions get sharper: how to run an adverse impact audit when the vendor holds the data is the version of this for bought software.

What if your volume is too small to measure?

Say so, and stop there. A ratio computed on nine hires and four rejections is arithmetic rather than evidence, and publishing it invites a conclusion the data cannot support in either direction. The honest posture is to state the number of decisions behind any figure you release, and to treat a sample that cannot separate signal from noise as a reason to keep better records.

The record and the ratio are different assets. A ratio needs volume before it says anything. A written reason attached to one decision is complete on the day it is written, works at any scale, and is the only thing that lets a specific candidate's outcome be examined at all. Small employers have no access to the first and full access to the second, which reverses the usual assumption that fairness measurement is something only large companies can afford.

The market-wide picture is worth carrying only as context. A meta-analysis of every available US field experiment on hiring discrimination against African Americans or Latinos, 28 studies and 55,842 applications, found Whites receiving on average 36 percent more callbacks than African Americans since 1989, with no detectable change over the following 25 years 3. Carry it as the background your own numbers sit against; it describes no individual employer. The two practical follow-ons are what to do when you have too few hires to measure bias and, for the smallest teams, the smallest hiring process you can defend.

See a sample report

Common questions

Does concentration mean most employers do not discriminate?

No. Concentration means the size of the gap varies enormously between firms, with a fifth of the discriminating employers producing close to half the lost contacts. In the experiment that produced that finding, Kline, Rose and Walters could not reject the hypothesis that all 108 firms weakly favored White names, and their lower bound is that at least 7 percent of jobs in the experiment discriminated, rising to at least 20 percent at the firms they could flag. The 23 identified firms are the ones the statistics could separate, not the complete set.

Can an industry benchmark tell you whether your own hiring is fair?

It cannot. A market-wide figure is an average over employers whose behavior differs by an order of magnitude, and an experiment that measured 108 large employers individually found gaps concentrated in a minority of them. A benchmark is useful for one thing: deciding whether a question is worth asking at all. Once you are asking it, only your own numbers, cut by stage and by req family, can answer it, and only your own records can say what produced them.

The gender numbers come out even. Does that mean the process is even-handed?

An even average is also what two opposite biases look like once they cancel. In the large-employer experiment the average gender contact gap was zero and the between-company standard deviation was 2.7 percentage points, roughly symmetric about zero. The same arithmetic runs inside one company, across req families and across screeners. Cut the number those ways before reading it as parity.

Which stage should be measured first?

The one where the most people leave and the least is written down, which in most funnels is the first contact decision. It is also the stage where the research base is strongest, since correspondence experiments can measure it directly. Measuring it first has a second benefit: it is usually the stage where the fix is cheapest, because the criterion being applied is a small number of written rules rather than a room full of interviewers.

Is measuring your own hiring risky if you find something?

Employment counsel owns that one, and the answer varies by jurisdiction. What is publicly stated in California's amended employment regulations, effective October 1, 2025, is that evidence, or the lack of evidence, of anti-bias testing is relevant to a claim that an automated-decision system or a selection criterion discriminated, and to any available defense, including the response to the results. The sharp edge is not testing. It is testing, finding a disparity, and doing nothing about it, which is the case those regulations single out.

References

  1. 1. Systemic Discrimination Among Large U.S. Employers NBER Working Paper 29053 (revised May 2022); published in the Quarterly Journal of Economics, 137(4), 1963-2036, 2022. nber.org Supports the 83,000 applications to 108 employers, the 24 percent contact rate, the 2.1 percentage point gap, the top-quintile and Gini concentration results, the 23 firms flagged at a 5 percent false discovery rate, the lower bounds of at least 7 percent of jobs and at least 20 percent at flagged firms, and the 2.7 percentage point between-company standard deviation of gender gaps.
  2. 2. Hiring Discrimination and the Task Content of Jobs: Evidence from a Large-Scale Resume Audit arXiv preprint arXiv:2604.01933 (Braun and co-authors), version 2, July 2026, 2026. arxiv.org Supports the 36,880 applications to 9,220 advertisements and the 28 to 43 percent lower callbacks in management occupations, cited as a preprint under review.
  3. 3. Meta-analysis of field experiments shows no change in racial discrimination in hiring over time (PubMed record, PMID 28900012) Proceedings of the National Academy of Sciences, 114(41), 10870-10875; abstract retrieved from the U.S. National Library of Medicine, 2017. eutils.ncbi.nlm.nih.gov Supports the 28 studies and 55,842 applications, the 36 percent average callback advantage for Whites since 1989, and the absence of a detectable change over 25 years.
  4. 4. Final Unmodified Text of Proposed Employment Regulations Regarding Automated-Decision Systems (Attachment B), 2 CCR sections 11009, 11013 California Civil Rights Department, Civil Rights Council, 2025. calcivilrights.ca.gov Supports the claim that anti-bias testing evidence and the response to its results are relevant to a FEHA discrimination claim, and that the retention period moved from two years to four with automated-decision system data included.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.