Screening

Name-Blind Screening Under-Delivers, and the Field Trials Say Why

Name-blind screening reduces bias less than its reputation suggests. In the randomized trial of anonymized applications run by the French public employment service, across about 600 volunteering firms, the interview gap widened rather than closed. Redaction also leaves school, postcode, dates, employment gaps and the register of the writing untouched, and it acts on one stage: a gap that opens at the interview or the offer is unaffected. Measure pass rates by stage before buying anonymization.

The takeBlinding survives because it is the only bias fix that fits in a settings toggle, and that is also its limit. It acts on one field at one stage, while the criteria the reviewer applies afterward stay exactly as they were. The French result is a warning rather than a verdict: removing context can take away a correction somebody was already making in a candidate's favor. A team that ships anonymization and stops has bought a visible intervention and left the stage where its own gap opens unmeasured.

Where Olive fits

Open a role and see what the work shows

No screen can tell you which resume a model wrote, so Olive skips the artifact and assesses the person: an occupational assignment done with an AI assistant, returned as six findings with the timestamped excerpt behind each one. Olive has not completed a bias audit, because attempt volume is too low for a four-fifths ratio to mean anything, and olive.is says so.

Rank your shortlist

Does removing the name reduce bias?

Not reliably, and the randomized trial of the practice points the other way. When the French public employment service randomly assigned about 600 participating firms to receive anonymized or name-bearing applications, the interview gap widened: minority candidates' interview rate fell and majority candidates' rose 1. The firms had volunteered, and volunteers were already interviewing relatively more minority candidates than others.

The authors name two mechanisms and both are about that setting. Firms selected into the program, so the sample was drawn from employers already inclined to interview these candidates. And stripping the name also strips the context a recruiter was using to discount a negative signal, such as an employment gap, for a candidate they thought they could place. Remove the context and the discount goes with it.

The result belongs to its setting. It is a program evaluation in France, on firms that volunteered, and it did not test automated screening or establish a general law. What it establishes is narrower and still useful: removing an identifier is an intervention with its own selection effects, not a neutral safety measure that can only help.

The evidence usually cited on the other side does not say what the pitch says it does. Bertrand and Mullainathan sent 4,870 fictitious resumes to help-wanted ads in Boston and Chicago and found resumes with White-sounding names drew a callback 9.65 percent of the time against 6.45 percent for otherwise equivalent resumes with African-American-sounding names, a gap of 3.20 percentage points, or 50 percent 2. It is a real and careful result. It is also a callback measurement from 2001 and 2002 in two cities, and the authors write plainly that they cannot translate it into gaps in hiring rates or earnings 2.

What survives redaction?

School, postcode, dates, employment gaps, visa phrasing, the names of former employers, and the register of the writing itself. Each one correlates with something the redaction was meant to hide, and each stays on the page. A correspondence study varies the name and holds everything else identical, which is exactly what a real application does not do.

The employment gap is the clearest case, because it is priced at the screen and almost nobody thinks of it as an identity signal. A correspondence study sending roughly 12,000 resumes to real online postings across the 100 largest US metropolitan areas found the callback rate falling from roughly 7 percent at one month of unemployment to roughly 4 percent at eight months, about 45 percent lower, after which additional months made almost no further difference 4. Fieldwork ran in 2011, and the authors read the spell as a signal employers use rather than as skills decaying.

Thinning a document is not neutral either, once a model is doing the reading. In a study of resume screening through language-model retrieval, cutting resumes down to a name and job title produced significant race differences in more of the bias tests than full-length resumes did, and the group selection-rate gaps widened 5. The paper's own percentages in that section do not reconcile to whole numbers of tests, so read the direction and leave the decimals alone.

Those documents kept the name and lost the content, which is the reverse of what an anonymization tool does, so what the study measures is thinning. The mechanism is the part that carries over: strip a document down and whatever identity signal is left takes up more of what the model has to read. A thin profile fed to a ranking layer is not automatically the safer object, and school, postcode and dates are usually still on it. That two-layer pipeline is worth understanding before configuring anything: what AI resume screening does to the pile before you open it.

Which stage does your gap open at?

Only your own pass rates can tell you, and the field-experiment average says the screen is not the whole of it. Across the twelve field experiments that follow applicants past the interview invitation, majority applicants received 53 percent more callbacks than comparable minority applicants and 145 percent more job offers 3. Roughly half the discrimination in offers opens after the resume screen, in rounds no redaction reaches. Blinding acts on one stage only.

Those numbers need their limits attached. Twelve studies is a small and selected pool, skewed toward roles that can be audited in person or by telephone, and 145 percent is a ratio of offer rates rather than a percentage-point difference. The additional discrimination after the interview correlates only weakly with the callback gap, so a firm cannot read its own later-stage behavior off its screening numbers in either direction.

Which is the operational point. Before buying anonymization, compute pass rates stage by stage for the groups you can lawfully count, and see where your own funnel narrows. What you may collect, and how the analysis is held once it exists, turns on the jurisdiction the role sits in, so settle that with counsel before the first query runs. If the screen passes people evenly and the onsite does not, blinding the screen changes nothing you care about, and the money and the political capital have been spent on the wrong stage.

Two adjacent pieces do this arithmetic properly: finding which stage the gap opens at and measuring bias stage by stage. Both are more work than a toggle and both tell you something a toggle cannot.

Replace blinding with criteria you wrote down first

Write the criteria before the pool arrives, apply them one at a time, and record a reason for each rejection in the words of the criterion. That is slower than a toggle, and it acts on the part blinding cannot reach, which is what a reviewer is free to weigh once the name is gone. Criteria written after reading fifty applications describe those fifty people.

Then measure where your own hiring happens rather than assuming a background rate. In an audit sending more than 83,000 applications to entry-level vacancies at 108 of the largest US employers, contact gaps were highly concentrated: the top quintile of discriminating firms accounted for nearly half the contacts lost to Black applicants, with a Gini coefficient of about 0.4 6. Concentration is not a clean bill of health for everyone else, and the authors say they could not reject the hypothesis that all 108 firms weakly favored White names. What it does establish is that this is a property of particular employers and particular jobs, which means it has to be measured where you hire.

Three habits carry most of the value, and none of them needs a vendor. Decide what a rejection reason may say before the first application is read. Keep the reasons in a field somebody can count later. And review the criteria themselves for the ones that are doing work you never intended, which is the usual home of a gap nobody chose.

Where the screen has stopped separating people at all, the problem is upstream of fairness and is its own question: what to screen on when every resume looks perfect, and how to tell whether the screen is throwing away the wrong people.

See a sample report

Common questions

Should we turn name-blind screening off, then?

Not necessarily, but stop counting it as the fix. Redaction is cheap, it removes one real signal, and in a process with no other controls it is better than nothing. What it cannot do is carry a fairness claim on its own, and treating it as the answer tends to end the work. If you keep it, keep it alongside written criteria and stage-level measurement, and be honest in internal reporting that the intervention covers one field at one stage.

Does blinding help with anything at all?

It helps most where a reviewer would otherwise see the identity signal first and the evidence second, and where nothing else in the process constrains the judgment. It also has a signalling value inside a company that can be worth having. What the evidence does not support is the strong version: that removing the name produces a proportional reduction in the gap, or that a redacted application is a neutral object. The French trial found the opposite direction among the employers who volunteered for it.

What about blinding work samples rather than resumes?

A work sample marked without the author's identity is a stronger version of the same idea, because what gets judged is the work itself. It still leaks: writing register, tool choices, formatting habits and the phrasing of a second-language writer all travel. The gain comes less from the anonymity than from marking against a rubric written in advance, which is what makes two markers comparable and what makes a decision explainable afterward.

Our pipeline is too small to compute pass rates. What then?

Say so plainly rather than reporting a ratio that cannot mean anything. With small numbers, a single hire moves a rate by tens of percentage points, and a four-fifths comparison built on a handful of people measures noise. What is still available is process evidence: written criteria, recorded reasons, consistent questions, and a periodic review of which criteria are doing the filtering. Pool several quarters or several similar roles before computing anything, and label the aggregation.

Does an AI screen remove the human bias problem?

It relocates it and makes it harder to see. A model's ordering reflects the text it learned from, and in measured studies a name alone moves the output. It also gives you one artifact to interrogate instead of many reviewers, which is a genuine advantage if you can get at the artifact. The practical question to ask a vendor is not whether the tool is fair but what was tested, on whose data, at which stage, and what the tool does when the answer is uncertain.

References

  1. 1. Unintended Effects of Anonymous Resumes IZA Discussion Paper 8517 (Behaghel, Crepon and Le Barbanchon); published in American Economic Journal: Applied Economics, 7(3), 1-27, 2014. docs.iza.org Supports the finding that anonymization widened the interview-rate gap across about 600 participating French firms, and the two mechanisms the authors identify: firm self-selection and the loss of context used to discount a negative signal.
  2. 2. Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination American Economic Review, 94(4), 991-1013 (Bertrand and Mullainathan), 2004. jenni.uchicago.edu Supports the published callback figures (9.65 percent against 6.45 percent, a 3.20 percentage point or 50 percent gap) and the authors' own statement that the result cannot be translated into gaps in hiring rates or earnings.
  3. 3. Evidence from Field Experiments in Hiring Shows Substantial Additional Racial Discrimination after the Callback Social Forces, 99(2), 732-759 (Quillian, Lee and Oliver), 2020. academic.oup.com Supports the 53 percent callback and 145 percent job-offer figures across the twelve field experiments that follow applicants past the interview invitation, and the weak correlation between a firm's callback gap and its later-stage behavior.
  4. 4. Duration Dependence and Labor Market Conditions: Theory and Evidence from a Field Experiment National Bureau of Economic Research Working Paper 18387 (Kroft, Lange and Notowidigdo); published in the Quarterly Journal of Economics, 2012. nber.org Supports the employment-gap callback penalty: roughly 7 percent at one month of unemployment falling to roughly 4 percent at eight months, about 45 percent lower, and flat thereafter.
  5. 5. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval Kyra Wilson and Aylin Caliskan, Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2024. arxiv.org Supports the claim that cutting resumes down to a name and job title produced significant race differences in more bias tests than full-length resumes did, with wider group selection-rate gaps. The tested documents retained the name, so this is evidence about thinning a document, not about redacting an identifier.
  6. 6. Systemic Discrimination Among Large U.S. Employers National Bureau of Economic Research Working Paper 29053 (Kline, Rose and Walters); published in the Quarterly Journal of Economics, 137(4), 1963-2036, 2022. nber.org Supports the concentration finding across more than 83,000 applications to 108 large employers, the top quintile accounting for nearly half the lost contacts, the Gini coefficient of about 0.4, and the authors' inability to reject weak favoritism at all 108 firms.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.