Screening

Callback Studies Measure the Screen, Not the Hire

Name-swap resume studies prove differential treatment at the document stage, and nothing past the point where an employer replies. Bertrand and Mullainathan sent 4,870 fictitious resumes to help-wanted ads in Boston and Chicago in 2001 and 2002: White-sounding names drew a callback 9.65 percent of the time against 6.45 percent for African-American-sounding names, a gap of 50 percent attributable to the name manipulation. The same paper's weaknesses section says the result cannot be translated into gaps in hiring rates or earnings.

The takeThe literature is quoted for the wrong conclusion. It gets used to argue that the resume screen is where bias lives, and therefore that stripping names is the fix. The same body of work puts roughly half the offer gap after the invitation, and the one randomized test of anonymized applications made the interview gap wider rather than smaller. A program spending its whole budget at the screen is optimizing one half of its own problem and calling it the whole.

Where Olive fits

Open a role and see what the work shows

A screen reads documents, which is why a gap measured there is a gap in how documents were treated. Olive works on a different object one stage later: a 40-to-60-minute occupational assignment done with an AI assistant, written up by a person as six findings, each carrying the moment in the session it rests on, and granted to the candidate in the same form.

Rank your shortlist

What does a name-swap study actually prove?

That a change in one signal on an otherwise identical document changes the reply rate. Bertrand and Mullainathan mailed 4,870 resumes to over 1,300 help-wanted ads in Boston and Chicago in 2001 and 2002, randomly assigning each one a White-sounding or an African-American-sounding name, and recorded a callback for 9.65 percent of the first group against 6.45 percent of the second 1.

Three details go missing every time the study is retold, and each one changes what the number licenses.

  • The manipulation was not the first name alone. Race-typical surnames were assigned as well: Walsh, Murphy and O'Brien against Washington, Jones and Jackson 1. Anyone summarizing it as a first-name swap has the design wrong.
  • The gap is 50 percent one way and about a third the other. White names received 50 percent more callbacks; African-American names received about a third fewer. Quoting both as though they were separate results doubles a single finding.
  • The outcome is a reply, not a job. The authors write that they are not able to translate their results into gaps in hiring rates or gaps in earnings 1. That sentence sits in the paper's own weaknesses section, and it is the sentence that never travels with the citation.

The population is narrow too: sales, administrative support, clerical and customer-service ads, answered through newspaper listings, which excludes every hire that arrived through a referral or a network. So the study is strong evidence that racially distinctive names were treated differently by employers reading documents in two cities in the early 2000s. It is not a national rate of anything, and it is not a claim about the average African-American applicant, since not every employer will have read the name as a racial signal.

One substantive methodological objection exists, and it comes with its resolution attached. Heckman and Siegelman showed that groups differing in the variance of unobserved productivity can generate spurious evidence of discrimination in either direction, even when the applications are identical on everything the researchers control. Neumark showed an unbiased estimate can be recovered when a study deliberately varies applicant quality, and applying that correction to this data strengthened the evidence of discrimination 6.

Where does the rest of the gap open?

After the invitation, in roughly equal measure. Quillian, Lee and Oliver pooled the twelve field experiments that follow applicants past the callback and found majority applicants received 53 percent more callbacks than comparable minority applicants and 145 percent more job offers 2. Their own reading of that split is that about half the discrimination in offers is already present at the application-to-callback stage.

Two cautions belong with the pair of numbers. The 145 percent is a ratio of offer rates, so it magnifies a small absolute difference when offer rates are low: it is never percentage points, and 53 does not subtract from it. And the pool is twelve studies, selected because they were able to run in-person or telephone stages, which skews the sample toward lower-skill and face-to-face roles.

The line that matters for a hiring lead is the correlation rather than the headline. Additional discrimination after the interview correlates with the callback gap at only r = 0.21 2. A firm with clean screening numbers has learned nothing about its own interview stage in either direction, which is why an end-to-end fairness number and a stage-by-stage one answer different questions. If your pass rates have moved for reasons nobody can name, which funnel metrics still mean anything is the prior question, and how you would know if the screen is throwing away the wrong people is the one after it.

Don't take name-blind screening as the fix

Removing the name is a partial control for one signal, and in the one place it was randomized it backfired. When the French public employment service randomly assigned about 600 participating firms to receive anonymized or name-bearing applications, the interview-rate gap widened: the interview rate of minority candidates fell while that of majority candidates rose 4. Two mechanisms explain it, and both are about that program rather than about anonymity as an idea.

The first is selection. Firms volunteered, and the volunteers were already interviewing and hiring relatively more minority candidates, so anonymization took away a discretion they had been using in the candidate's favour. The second is context. Stripping identifiers also strips the information a recruiter uses to discount a negative signal, such as an employment gap, for a candidate they were minded to place.

The treatment in these studies is not uniform either. A survey experiment measuring how people actually read the names used in correspondence audits found the congruent perception rate for Black first names was 75.0 percent with no surname attached, 82.5 percent with a Black surname and 66.5 percent with a White one 3. Names more common among highly educated Black mothers were markedly less likely to be read as Black at all. That cuts both ways: a weakly racialized name dilutes the treatment and understates discrimination, while a strongly racialized one may carry a class signal riding alongside. Neither reading rescues the simple version, which is why whether blind screening reduces bias is still an open question.

Measure your own funnel one stage at a time

Measure each handoff separately, because a gap can sit inside one stage while the end-to-end number still looks acceptable. Count applications, contacts, interviews, offers and hires as distinct populations, using whatever demographic data you actually hold, and compare the ratio at each transition.

1. Cut it at every boundary. Application to contact, contact to interview, interview to offer, offer to acceptance. A process that passes overall can still contain one transition that does not, and averaging hides it by construction. 2. Keep the reason, not only the rate. A rate tells you a difference exists. The note a screener wrote about why this application stopped here is what you can read six months later when the rate looks wrong, and it is the only artifact anyone can argue with. 3. Do not put two different quantities in one sentence. The 50 percent gap measured in 2001 and 2002 and the 2.1 percentage point gap Kline, Rose and Walters measured at 108 large employers are not the same measurement: the later experiment ran at unusually large employers whose overall contact rate was nearly three times higher, which shrinks proportional differences 15. Quoting them side by side as a trend is an error somebody will catch. 4. Say what your volume can carry. With eleven hires a year, no ratio you compute separates a real gap from noise, and publishing one anyway is worse than publishing nothing.

None of that substitutes for the evidence itself. A callback gap counted in a spreadsheet is a fact about outcomes with no account of how any of them were reached, which is exactly where the correspondence literature is stuck: it can prove that treatment differed and it cannot say what any employer was thinking. Your own process does not have to leave you in the same position. Where that gap concentrates at the employer level is the subject of a minority of employers accounting for most lost callbacks.

See a sample report

Common questions

Is the Emily and Greg study still current?

It is a 2004 publication reporting fieldwork from 2001 and 2002, in two cities, through newspaper help-wanted ads, in sales, clerical, administrative and customer-service roles. Treat it as the landmark rather than the latest. The best-identified recent measurement is Kline, Rose and Walters, who sent more than 83,000 applications to entry-level vacancies at 108 large US employers and found distinctively Black names reduced the chance of contact by 2.1 percentage points, about 9 percent of the Black mean contact rate. The proportional gap is much smaller there, partly because those employers contacted applicants far more often overall.

Which figures should be quoted, 10.08 and 6.70 or 9.65 and 6.45?

The published paper reports 9.65 percent and 6.45 percent, a difference of 3.20 percentage points. The pair 10.08 and 6.70 comes from the July 2003 NBER working paper version of the same study, and it still circulates widely. Both versions state the same headline 50 percent gap on different underlying rates. Cite the published figures, and if the working paper's numbers are needed for some reason, label them as the working paper's.

Does a 50 percent gap mean Black applicants are 50 percent less likely to be hired?

No, and that sentence contains three separate errors. The 50 percent is White names receiving 50 percent more callbacks, which comes out at about a third fewer callbacks for African-American names. A callback is an invitation to interview, not a hire. And Bertrand and Mullainathan state in the paper that they cannot translate the result into gaps in hiring rates or earnings. The study supports a callback claim and only a callback claim.

Are correspondence studies methodologically sound?

They have one substantive objection with a known resolution. Heckman and Siegelman showed that groups differing in the variance of unobserved productivity can produce spurious evidence of discrimination in either direction, even when the applications are identical on everything the researchers control. Neumark showed an unbiased estimate can be recovered if a study deliberately varies applicant quality, and applying that correction to Bertrand and Mullainathan's data produced stronger evidence of discrimination, not weaker. Anyone citing the critique as a debunking is using it backwards.

Does an equal-opportunity statement on the posting mean the screen behind it is even-handed?

It predicted nothing good in the Bertrand and Mullainathan data. Employers whose ads stated they were Equal Opportunity Employers, and US federal contractors, which at the time carried affirmative-action obligations, showed no smaller racial callback gap; each characteristic went with a larger one, marginally significantly so for the contractors. That is an observational comparison inside an experiment rather than a randomized test of the statements themselves, and it says nothing about modern programs. It does mean the line in the ad is not evidence about behaviour behind it.

References

  1. 1. Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination American Economic Review, 94(4), 991-1013; full-text copy opened at the University of Chicago (jenni.uchicago.edu), 2004. jenni.uchicago.edu Supports the 9.65 percent against 6.45 percent callback rates and the 50 percent gap, the race-typical surname design, the authors' own statement that the result cannot be translated into hiring or earnings gaps, and the finding that stated equal-opportunity status went with no smaller gap.
  2. 2. Evidence from Field Experiments in Hiring Shows Substantial Additional Racial Discrimination after the Callback Social Forces, 99(2), 732-759 (Oxford University Press), 2020. academic.oup.com Supports the 53 percent callback gap and 145 percent offer gap across twelve field experiments and more than 13,000 applications, and the r = 0.21 correlation between a firm's callback gap and its later-stage behaviour.
  3. 3. How Black Are Lakisha and Jamal? Racial Perceptions from Names Used in Correspondence Audit Studies Sociological Science, 4, 469-489 (S. Michael Gaddis), 2017. sociologicalscience.com Supports the congruent perception rates of 75.0, 82.5 and 66.5 percent, and the finding that names common among highly educated Black mothers were less likely to be read as Black.
  4. 4. Unintended Effects of Anonymous Resumes IZA Discussion Paper 8517 (Behaghel, Crepon and Le Barbanchon); published in American Economic Journal: Applied Economics, 7(3), 1-27, 2014. docs.iza.org Supports the claim that randomized anonymization across about 600 participating French firms widened the interview-rate gap, and the two mechanisms the authors identify.
  5. 5. Systemic Discrimination Among Large U.S. Employers National Bureau of Economic Research Working Paper 29053 (revised May 2022); published in the Quarterly Journal of Economics, 137(4), 1963-2036, 2022. nber.org Supports the 2.1 percentage point contact gap at 108 large employers and the higher baseline contact rate that makes it a different quantity from the 2004 study's 50 percent.
  6. 6. Detecting Discrimination in Audit and Correspondence Studies The Journal of Human Resources, 47(4), 1128-1157 (David Neumark); copy opened at the University of California, Irvine, 2012. sites.socsci.uci.edu Supports the Heckman and Siegelman variance objection to correspondence designs, Neumark's correction requiring deliberate variation in applicant quality, and the fact that applying it to Bertrand and Mullainathan's data strengthened the finding.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.