Screening

Models Screening Resumes Show Measured Name Preferences

Bias in language models screening resumes has been measured, several times, by people who are not selling anything. An audit of three open text-embedding models screening 554 resumes against 571 job descriptions found White-associated names significantly preferred in 85.1 percent of 27 racial bias tests. A newsroom experiment ordering eight matched resumes a thousand times per role found names distinct to Black women placed first for a software engineering job 11 percent of the time, against a one-in-eight baseline.

The takeA vendor's impact ratio and an evidence base are different objects, and the gap between them is where a recruiter gets hurt. Under New York City's 2023 rule an employer that has never used a tool may post an audit computed entirely on other employers' applicants, or on synthetic data. That document can be accurate, current and completely uninformative about how the thing will behave on your pool. Ask what it was computed on before you ask what it says.

Where Olive fits

Open a role and see what the work shows

Olive has published no bias audit, and the site says so: attempt volume is too low for an impact ratio to mean anything yet. What it produces instead is six findings written by a person about one candidate's session, each carrying the timestamped excerpt behind it, which is the material a later disparity could be argued with rather than only counted.

Rank your shortlist

What has actually been measured?

Three independent lines of evidence, none of them produced by a vendor. In a 2024 audit presented at the AAAI/ACM conference on AI, Ethics and Society, three open text-embedding models screened 554 resumes against 571 job descriptions across nine occupations, and resumes carrying White-associated names were significantly preferred in 85.1 percent of 27 racial bias tests, Black-associated names in 8.6 percent 1.

Read the denominator before the headline. The 27 is three models times nine occupations, so 85.1 percent is the share of statistical tests showing a significant preference, not the share of candidates rejected and not a selection-rate gap. Every case where Black names won came from one of the three models, so "the model" is not one thing. Race and gender were signalled only by a first name over a constant surname. Selection meant clearing a similarity cut, and nobody in the study was hired or turned down.

Two results inside that audit matter more than the headline does. Resumes with White male names were preferred over resumes with Black male names in all 27 intersectional tests, while the White male against White female comparison produced a significant difference in fewer than half 1. A fairness check run one protected attribute at a time can pass while the intersection fails. And bias widened when the document carried less content: a name and job title alone produced more significant differences and wider group gaps than full-length resumes did 1.

The second line is a newsroom test. Bloomberg had two GPT versions order eight near-identical resumes a thousand times per job description for four real Fortune 500 roles, assigning demographically distinct names at random: resumes with names distinct to Black women were top-ranked for the software engineering role 11 percent of the time against a one-in-eight baseline, and at least one adversely impacted group appeared in every role except retail manager under GPT-4 2. Four job descriptions, synthetic resumes, and June 2023 API snapshots that no longer exist in that form.

What do the tests find on attributes nobody audits?

The largest measured gaps in this literature sit on attributes almost nobody checks. A 2023 replication testing three deployed chat models on 334 real resumes found no significant true-positive-rate gap between White and African American names or between male and female names, and large gaps on maternity-leave employment gaps, pregnancy status and political affiliation, frequently above 30 percent 4.

That paper is a genuine counterweight to the two above, and the disagreement is methodological: a different task, a different metric, different models and a much smaller corpus. The authors' explanation for the clean race and gender result is a hypothesis rather than a measurement, that the models had been tuned on the most obvious attributes. The structural lesson survives either way. A system can be made to pass the tests everyone runs and still fail the ones nobody runs.

Disability is the clearest case of an unwatched attribute. Asked to order a control CV against the identical CV plus a disability-related leadership award, a scholarship, a panel presentation and an organizational membership, GPT-4 put the objectively stronger version first in only 15 of 70 trials, and in none of the ten autism trials 3. One CV, one job description, ten trials per condition, and a model that called two identical CVs a tie in only 70 percent of baseline runs, so some of that is noise. The signal that is not noise: additional achievements were read as a penalty when the achievement named a disability.

Why doesn't a vendor's bias audit answer this?

Because a published bias audit certifies something much narrower than it sounds. New York City's 2023 rule defines it as a selection rate and an impact ratio for each category, computed separately for sex, for race and ethnicity, and for the intersections, and it lets an independent auditor drop any category representing less than 2 percent of the audit data 5. No other US jurisdiction requires that arithmetic.

Three limits follow, and they all point the same way.

  • It measures group selection rates and nothing else. Not accuracy, not job-relatedness, not whether any individual was treated correctly. A tool can pass and still be useless for choosing anyone.
  • The 2 percent carve-out drops the smallest groups, which are frequently the groups a bias audit exists to protect.
  • The audit need not describe your candidates at all. An employer that has never used the tool may rely on an audit built on other employers' historical data, or on synthetic test data where too little real data exists 5. Once you have your own data you can no longer do that, so the audit is least informative exactly when you are deciding whether to buy.

The corpus of posted audits is also in worse shape than it looks. Researchers at the ACLU and three universities collected every publicly available Local Law 144 audit they could find through early November 2024, arriving at 44 reports covering 116 audits, and 83 percent of those audits reported missing race or sex information for some of the applicants the tool had evaluated 8. Much of that is applicants declining to self-identify, a design flaw in what the law asks for. The impact ratio is still computed on whatever demographic data survives. New York State auditors then re-reviewed the same 22 employers the city's own enforcement sweep had covered, using only public information, and identified at least 17 instances of potential non-compliance where the city had identified one: audits not performed by an independent auditor, missing selection-rate and impact-ratio calculations, and failures to explain the use of historical data 6. Those employers were pre-selected by outside researchers because they raised questions, so that rate does not generalize, and none of it has been adjudicated. Whether the rule reaches your own use of a tool turns on your jurisdiction and on how the tool is used, which is a question for counsel. If you are about to interrogate an audit, what to ask an AI screening vendor that says it is bias tested is the question list, and the difference between a validated assessment and a bias-audited one is the distinction most vendor pages blur.

Ask which model version the evidence was run on

Ask three questions before the tool touches an application, and record the answers. Which model version was the published evidence run on, what happens to that evidence at the next release, and what does the tool leave behind for each individual decision. A published measurement expires with the model version it ran on, and vendors rarely volunteer the expiry date.

Two more, checkable without special access.

1. A fairness instruction is not a control. A custom GPT built on diversity and disability-justice principles put the stronger CV first in 37 of 70 trials against plain GPT-4's 15, a real improvement and still wrong about half the time on a comparison where the correct answer was the same CV plus four accomplishments 3. Bloomberg reports the same shape: prompt sensitivity is its first listed limitation, and adding a line stating that discrimination is illegal did not change the findings 2. 2. The ordering is not stable. In the same tests both models favored whichever resume they saw first, GPT-3.5 naming the first candidate most qualified 56 percent of the time and GPT-4 28 percent, against the one-in-eight share expected if input order did not matter 2. Shuffling the order fixes that and does nothing about the demographic result.

Effect sizes deserve the same care. Across more than 2 million generated hiring emails covering 300 names, 820 prompt templates and 41 occupations on five models, the absolute difference between the highest and lowest group acceptance rate ran between 1.60 percent and 3.78 percent: statistically significant, small in magnitude, and pointing in different directions under different templates 7. A 2 to 4 point difference can fail an impact ratio when base rates are low and be trivial when they are high, so the ratio and the difference are not interchangeable.

What all of this leaves a recruiter with is a records problem. A tool that returns an ordering and nothing else gives you a number per candidate, and a number cannot be interrogated: if a disparity shows up in six months, there is no account of any individual decision to read. A tool that returns the specific thing in the application it acted on gives you something a person can disagree with. That is also why what Amazon's recruiting tool actually taught is a story about inspection.

See a sample report

Common questions

Does 85.1 percent mean 85 percent of Black applicants were rejected?

No. It is the share of 27 statistical tests, three models across nine occupations, in which a significant preference for White-associated names appeared. It is not a rejection rate, not a selection-rate gap and not a four-fifths impact ratio. A single model behaving badly across nine occupations moves that headline by a third on its own, which is why the paper's own breakdown by model matters more than the summary figure.

Do these studies apply to the screening tool your company bought?

Not directly. The embedding audit tested three open research models nobody sells as a hiring product, and the newsroom test used general-purpose API snapshots from 2023. What transfers is the mechanism rather than the number: name signals move retrieval and ordering, thin documents amplify that, prompts move results, and input order moves them too. If a vendor's product is built on a general-purpose model, ask which one and which version, because the published evidence is attached to a specific release.

Does instructing the model not to discriminate fix it?

It shifts the result without making it explainable or stable. A custom GPT built on diversity and disability-justice principles improved the disability comparison from 15 of 70 trials to 37 of 70, a significant change and still a failure on a comparison where one CV was strictly stronger. Bloomberg separately reports that adding a statement about the illegality of discrimination did not change its findings. A prompt is a nudge to a system, not a control on a selection procedure.

Does stripping names from resumes solve it?

No, and the one measurement pointing at this question points the other way. In the embedding audit, resumes reduced to a name and job title produced more significant differences and wider group gaps than full-length resumes did, because removing content makes the remaining signal a larger share of what the model sees. Names are also not the only carrier: schools, dates, employers, gaps and language patterns all encode the same information. Name removal is a partial control for one signal, not a fix.

Which measurement should be trusted when the studies disagree?

Read three things before the number. The task, since ordering candidates against each other, classifying a resume into a job category and writing an acceptance email are different jobs that produce different results. The metric, since a selection-rate gap and a true-positive-rate gap answer different questions. And the model version and date, because every one of these results is attached to a specific release, and the release schedule outruns the publication schedule.

References

  1. 1. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval Kyra Wilson and Aylin Caliskan, Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2024), 2024. arxiv.org Supports the 554 resumes against 571 job descriptions, the 85.1 percent and 8.6 percent test shares, the intersectional result in all 27 tests, and the finding that title-only documents produced more significant differences than full-length ones.
  2. 2. OpenAI's GPT Is a Recruiter's Dream Tool. Tests Show There's Racial Bias Bloomberg (Leon Yin, Davey Alba and Jackie Nicoletti), read via the Internet Archive Wayback Machine, 2024. web.archive.org Supports the 11 percent top-ranking rate for names distinct to Black women in the software engineering role, the adversely impacted group in every role except retail manager under GPT-4, the 56 percent and 28 percent order effects, and prompt sensitivity as the first listed limitation.
  3. 3. Identifying and Improving Disability Bias in GPT-Based Resume Screening Glazko, Mohammed, Kosa, Potluri and Mankoff, University of Washington, ACM FAccT 2024, 2024. arxiv.org Supports the enhanced CV placing first in 15 of 70 trials and none of the ten autism trials, the 70 percent tie rate on identical CVs, and the custom GPT's improvement to 37 of 70.
  4. 4. Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT Veldanda, Grob, Thakur, Pearce, Tan, Karri and Garg, 2023. arxiv.org Supports the insignificant race and gender true-positive-rate gaps across 334 resumes and the large gaps on maternity-leave employment gaps, pregnancy status and political affiliation, frequently above 30 percent.
  5. 5. Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY 5-300 et seq.) NYC Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us Supports what a Local Law 144 bias audit computes, the exclusion of categories under 2 percent of the audit data, and the rule permitting an employer that has never used a tool to rely on other employers' historical data or synthetic data.
  6. 6. Enforcement of Local Law 144 - Automated Employment Decision Tools, Report 2024-N-6 Office of the New York State Comptroller, Division of State Government Accountability, 2025. osc.ny.gov Supports the re-review of the same 22 employers finding at least 17 instances of potential non-compliance where the city identified one, and the three categories of defect found.
  7. 7. Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender? An, Acquaye, Wang, Li and Rudinger, Proceedings of ACL 2024, 2024. arxiv.org Supports the 1.60 to 3.78 percent absolute group differences across more than 2 million generated hiring emails, 300 names, 820 templates and 41 occupations, and the prompt sensitivity the authors report.
  8. 8. Auditing the Audits: Lessons for Algorithmic Accountability from Local Law 144's Bias Audits Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT '25); Gerchick, Encarnacion, Tanigawa-Lau, Armstrong, Gutierrez and Metaxa, 2025. facctconference.org Supports the 44 reports covering 116 publicly available Local Law 144 audits collected through early November 2024, and the 83 percent of those audits reporting missing race or sex information for some of the applicants the tool evaluated.

8 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.