Policy
Amazon's Tool Learned Its Bias From Amazon's Own Hiring
Reuters reported in 2018 that Amazon built experimental resume-rating software from 2014, trained on ten years of its own resumes, and found by 2015 that it penalized resumes containing the word women's and downgraded graduates of two all-women's colleges. Recruiters looked at its output but never relied on those rankings alone, and the team was disbanded by the start of 2017. It was an internal experiment, not a deployed gate. The lesson is the training data: a model fitted to a company's own hiring reproduces it.
The takeThe case is told as an argument against automated screening, and it reads better as an argument for inspectability. The only reason anybody can describe what that system learned is that Amazon built it in house, could watch its behavior on its own resumes, and could see which words it was rewarding. Almost nobody buying a screening tool today has any of those three things, which is the part of the story worth carrying into a purchase decision.
Where Olive fits
Open a role and see what the work shows
Under automated-decision rules, a number attached to a person explains nothing. Olive produces no composite at all: a person writes each of the six findings, every one carries the excerpt from the session it rests on, and a released report exports with its rubric, scorer and bank versions attached.
Rank your shortlistWhat actually happened?
An internal experiment ran for about three years and was shut down. Reuters reported in October 2018, on five anonymous sources, that a team began building resume-rating software in 2014 that scored candidates one to five stars, trained on patterns in resumes the company had received over a decade; by 2015 the tool was not rating technical candidates in a gender-neutral way 1.
The specifics are narrower and more useful than the slogan.
- What it learned. It penalized resumes containing the word women's, as in women's chess club captain, and downgraded graduates of two all-women's colleges 1.
- What it decided. Nothing on the record. Reuters reports no candidate outcome at all and states that recruiters looked at the recommendations but never relied solely on those rankings 1.
- Why it ended. Reuters gives a broader reason than the gender problem. Unqualified candidates were often recommended for all manner of jobs, and with the technology returning results almost at random Amazon shut the project down. Executives had lost hope for it, and the team was disbanded by the start of 2017 1.
The evidentiary base is five people speaking on condition of anonymity to one reporter, and Amazon declined to comment on the recruiting engine or its challenges. That is reporting, not a study, and it should be named as reporting whenever it is cited. It also predates large language models entirely, so it is evidence about a 2014 machine-learning system and about no current product. The numbers that circulate attached to this case, rejection rates and percentages of women screened out, appear nowhere in the reporting.
Why didn't deleting the flagged words fix it?
Because the words were proxies for a pattern with many other carriers. Amazon edited the programs to be neutral to the flagged terms, and the people who described the effort to Reuters said that was no guarantee the machines would not devise other ways of sorting candidates that could prove discriminatory 1. Every attribute on a resume that correlates with the one you deleted is still sitting there, unnamed and unblocked.
One of those other ways is already in the same reporting: the tool favored candidates who described themselves with verbs more common on male engineers' resumes, such as executed and captured 1. The mechanism was documented in the research by then. Applying an association test to publicly released word vectors trained on a web crawl, researchers predicted the share of women in each of the 50 most relevant occupations from the vectors alone, correlating with US Bureau of Labor Statistics occupational data at a Pearson coefficient of 0.90 2. That is not a defect somebody introduced. It is the model being accurate about a segregated labor market, which is why "the model is objective" is the wrong premise. An accurate description of who held which job in the past is a poor instrument for deciding who should hold one next.
And most screening harm in practice is not even a learned pattern. In a survey of 2,275 executives across the US, UK and Germany fielded in early 2020, 48 percent of employers filtered middle-skills candidates out on an employment gap of more than six months, and more than 90 percent of those using a recruitment management system used it to filter or order candidates at initial screening 3. That is a rule a person wrote, applied at scale by software. It names no protected class, which is why it survives review, and its weight falls on caregivers, veterans, people with disabilities and people who were incarcerated. The measured bias in model-driven resume screening is a real and separate problem, and the configured filter is what most of the employers in that survey already had.
Why is this a story about inspectability?
Because the story exists at all only because someone could look. Amazon built the system in house, trained it on its own resumes, and could see which terms it was penalizing and which verbs it was rewarding. A bought tool gives a hiring team none of those three affordances, and the first publicly documented cooperative audit of a commercial screening vendor is narrower than most readers assume.
That audit is worth knowing in detail, because it sets the ceiling on what a claim of having been audited can mean. It was conducted in summer 2020 and scoped to five questions agreed in advance, and the auditors found the vendor did faithfully implement its stated fairness guarantees, with safeguards sufficient to reasonably ensure compliance with the four-fifths rule 4. Agreed in advance to be out of scope: the choice of fairness objective and metric, differential validity, intersectional fairness, groups outside the EEOC categories, the vendor's annual back-testing of deployed models, and data privacy. A pass means the code does what the vendor says it does. It is not a finding that the tool is fair, valid, job-related or predictive of anything.
Responsibility does not follow the same narrow path. In technical assistance the EEOC issued in May 2023 and removed from its site in January 2025, the agency answered the question of whether an employer is responsible for tools designed or administered by someone else with "In many cases, yes", and added that if a vendor is wrong about its own assessment and the tool does discriminate, the employer could still be liable 5. That document never had the force of law and is no longer current federal guidance, so quote it as what the agency said in 2023 and take the question to counsel. As a planning assumption it is the conservative one.
Ask what a tool was trained on, and who can read one decision
Ask two questions of anything you build or buy, and write the answers down. What was this trained on or fitted to, and can anyone outside the vendor read the basis of one individual decision it made. The first question is where Amazon's problem came from. The second is why Amazon could see it and a buyer usually cannot.
1. Name the training set out loud. A tool fitted to your own past hiring decisions will reproduce your own past hiring decisions, including the parts nobody wrote down and the parts nobody would defend. If the answer is the company's own historical data, that is the answer, and it is the risk you are buying. 2. Ask what survives one decision. If the only artifact is a position in a list, then a disparity found later can be counted and nothing more, because there is no record of what happened to any particular person. 3. Ask what a fix would look like. Amazon's was to neutralize the flagged terms, and the people who built the tool told Reuters that was no guarantee against other proxies 1. A vendor who answers instantly, with a term removed from a feature list, is repeating the move the case is famous for. 4. Keep your own record regardless. The criterion applied, and the specific thing in the application that met or missed it, for advances and rejections alike.
The half of the story that gets dropped is the half worth keeping. Amazon did not discover that machines are biased; it discovered what its own hiring had been doing, written down in a form specific enough to argue with. That is an uncomfortable gift and a rare one. Most screening runs on judgments that leave nothing behind, so the pattern never becomes legible at all. If you are about to put that question to a supplier, what you need from a vendor to survive a bias audit or an EEOC inquiry lists what to ask for in writing, and what to ask an AI screening vendor that says it is bias tested covers the claim itself.
Common questions
Did Amazon's tool reject women applicants?
The reporting does not say so, and it is the most common misstatement of the case. Reuters reports no candidate outcome at all, and says explicitly that recruiters looked at the tool's recommendations but never relied solely on those rankings. The opposite overcorrection is also wrong: it was not a system nobody ever saw. The accurate sentence is that an internal experimental tool produced recommendations recruiters could see, was found to be rating candidates in a way that was not gender-neutral, and was abandoned.
Is Amazon's recruiting tool still in use?
The resume-rating tool was abandoned: Reuters reports the team was disbanded by the start of 2017. The same reporting adds two things. Amazon kept a much-watered-down version of the engine for rudimentary chores such as culling duplicate candidate profiles from databases, and a new team in Edinburgh had been formed to try automated employment screening again, this time with a focus on diversity. The technology is 2014 to 2017 machine learning on resume text, which predates large language models, so the case is not evidence about any current product.
Would the same thing happen to a tool trained on your own hiring data?
That is exactly the risk the Amazon case demonstrates. A model fitted to past hiring outcomes learns the criteria that actually drove those outcomes, including the ones nobody wrote in a scorecard and the ones nobody would defend in a meeting. The problem is not that the model is unfair to the data. It is that the model is faithful to it. Any build-versus-buy conversation should start with what the system will be fitted to before it reaches what the system will output.
Does removing gendered terms from the training data solve it?
No, and Amazon's own attempt is the evidence. Reuters reports the programs were edited to be neutral to the particular terms, but that this was no guarantee the machines would not devise other ways of sorting candidates that could prove discriminatory. The same story shows the tool already keying on something else, favoring verbs more common on male engineers' resumes. The reason is structural: schools, activities, employers, dates, phrasing and gaps all carry overlapping information, so deleting one carrier leaves the pattern intact. Removing a term is a partial control for that term, not a fix for what it stood in for.
Who is responsible if a vendor's tool screens people out?
In technical assistance issued in May 2023 and removed from its site in January 2025, the EEOC answered "In many cases, yes" to whether an employer is responsible for a tool an outside vendor designed or administers, and said the employer could still be liable if the vendor was wrong about its own assessment. That guidance never had the force of law and is not current, so it is a planning assumption rather than a rule. The live litigation on vendor liability is a separate question. Take both to counsel.
References
- 1. Amazon scraps secret AI recruiting tool that showed bias against women web.archive.org Supports the 2014 start, the one-to-five-star rating, the ten years of training resumes, the penalty on the word women's and on two all-women's colleges, the favoring of verbs more common on male engineers' resumes, the statement that recruiters never relied solely on the rankings, the caveat that neutralizing the terms was no guarantee against other proxies, the reported reason for the shutdown (unqualified candidates recommended for all manner of jobs, results almost at random, executives losing hope) and the disbanding by the start of 2017, and the much-watered-down version kept for rudimentary chores alongside the new Edinburgh team.
- 2. Semantics derived automatically from language corpora contain human-like biases arxiv.org Supports the claim that word vectors predicted the share of women in the 50 most relevant occupations against Bureau of Labor Statistics data at a Pearson coefficient of 0.90, which is the mechanism behind a model reproducing a segregated labor market.
- 3. Hidden Workers: Untapped Talent hbs.edu Supports the survey of 2,275 executives fielded in early 2020, the 48 percent of employers filtering middle-skills candidates on an employment gap over six months, and the more than 90 percent of recruitment management system users who used it to filter or order candidates at initial screening.
- 4. Building and Auditing Fair Algorithms: A Case Study in Candidate Screening ccs.neu.edu Supports the scope and the finding of the first publicly documented cooperative audit of a commercial candidate-screening vendor, and the list of questions the auditors agreed in advance not to examine.
- 5. Select Issues: Assessing Adverse Impact in Software, Algorithms, and Artificial Intelligence Used in Employment Selection Procedures Under Title VII of the Civil Rights Act of 1964, Question 3 (archived capture, 2025-01-25) web.archive.org Supports the EEOC's 2023 answer that an employer may be responsible for a vendor-designed selection procedure and could still be liable if the vendor's own assessment was wrong, with the caveat that the document was removed in January 2025.
5 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.