Screening
Most of the Disparity Comes From Settings You Chose, Not Their Model
Bias in a screening tool comes from both the vendor's model and your configuration, and the configuration is where to start, because it holds the movable part and the vendor holds the rest. The vendor supplies a model. You supply the knockout questions, the cutoff, the choice between a ranked list and a pass bar, and whichever optional signals the implementation consultant switched on. A model audit can come back clean while the configured funnel produces a large gap, because it never tested your settings.
The takeCoverage of biased hiring AI is almost entirely about models: training data, embeddings, the withdrawn Amazon experiment. That framing leaves a buyer with nothing to do, because nobody outside the vendor can change a model and almost anybody inside the employer can change a threshold. The most useful thing a hiring team can believe about algorithmic bias is that a real share of it was configured by somebody whose name is still on the account, in a meeting nobody wrote down.
Where Olive fits
Open a role and see what the work shows
Olive has not completed a bias audit, and olive.is says so, because attempt volume is too low for a four-fifths ratio to mean anything yet. What it does publish is its own configuration: six findings written by a person about one candidate's session, each carrying the timestamped excerpt behind it, granted to the candidate in the same form the employer reads, and no part of it advances or rejects anybody.
Rank your shortlistList every setting and put a name against it
Export the configuration and write three columns: the setting, its current value, and who chose it. Sort that third column into what the employer decided, what the implementation consultant decided, and what shipped as a default nobody discussed. The third pile is usually the largest, and it is the one nobody defends, because no one in the room remembers agreeing to it.
Then write the job requirement that justifies each value, next to the value. Where no requirement can be written, the setting goes. The most-cited example is a filter that names no protected class and does not need to: in the Harvard Business School and Accenture survey of 2,275 executives, 48% of employers filtered middle-skills candidates out on an employment gap of more than six months, with the report describing almost half of surveyed companies screening out such a resume on that consideration alone 1. That is employer self-report from early 2020, and 48% is the share of employers applying the filter rather than the share of resumes rejected. It is also a rule a person configured, and its weight falls on caregivers, veterans, people with disabilities and people who were previously incarcerated.
The story everybody reaches for points the other way, and is usually told wrong. Amazon's experimental recruiting engine, built from 2014 and disbanded by the start of 2017, learned to penalise resumes containing the word women's and to downgrade graduates of two all-women's colleges 2. It was never the sole basis for a decision, and the account rests on five people speaking anonymously to one reporter. The durable line is the one that gets cut: the team edited the programs to be neutral to the flagged terms, and that was no guarantee the machines would not devise other ways of sorting candidates that could prove discriminatory 2. Deleting the term does not delete the proxy, which is equally true of every filter in your own configuration.
Why does the cutoff deserve its own pass?
Because a shipped threshold is a product decision about pass rates, and the standard it should meet is about your job. The Uniform Guidelines on Employee Selection Procedures, a 1978 US federal rule, say a cutoff score should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force 3. A default was chosen to produce a workable pass rate, not to describe proficiency in a role its author never saw.
The same section separates two uses of one procedure, and the distinction is worth quoting in a configuration review. Evidence sufficient to support a procedure used on a pass/fail basis may be insufficient to support the same procedure used on a ranking basis, and where ranking carries greater adverse impact than an appropriate pass bar would, the user should have enough evidence of validity and utility to support ranking 3. That condition is the one to check first, because it decides whether moving to a bar candidates either clear or do not also lightens what the evidence has to carry.
An ordered list can carry a second problem when a model produces it. In Bloomberg's tests of two 2023 model snapshots, both favoured whichever resume they saw first: the earlier version named the first candidate most qualified 56% of the time and the later one 28%, against the one-in-eight share expected if input order did not matter 4. Bloomberg disclosed this as a limitation and neutralised it by shuffling, and it says nothing about which candidates are favoured. It was measured on two superseded snapshots under one prompt template, so treat it as a property to re-test on whichever model you actually run. What it shows is that a share of the ordering can be an artifact of arrival time, and a pass bar has no arrival times.
A cutoff is also where one person can be screened out for a reason no group ratio will show you, and under the Americans with Disabilities Act that is an individual question. The EEOC's 2022 technical assistance on the ADA and algorithmic tools, removed from eeoc.gov in January 2025 and never binding law in the first place, gives the worked example: a business requiring 90 percent on a gamified memory assessment rejects a blind applicant who cannot play the game, though that applicant might have a very good memory and be perfectly able to perform the essential functions of a job that requires one 5. Screening someone out is not automatically unlawful, and the employer's defence runs on job-relatedness, business necessity and whether reasonable accommodation was possible. Which of those a specific cutoff can carry is a question for counsel on your own process, and the point for a configuration review is that a group-level ratio has said nothing about it.
Measure pass rates stage by stage, not tool-wide
A whole-funnel number tells you a gap exists and hides the setting that made it. Work out pass rates for each configured step separately: each knockout, the cutoff, the ranking cut, and any optional signal switched on at implementation. The split is what shows whether the gap concentrates at one configured step or spreads across every step, and only the first of those has a fix you can apply this week.
Two mechanics make this cheaper than it sounds. Where the tool exports a stage report, that is the whole job, and where it does not, the applicant tracking system holds the stage transitions even when the vendor holds the scores. Compute the ratio at each transition against the highest-passing category, and write the raw counts next to it: a ratio built on eleven selections is a number rather than evidence, and printing the count beside it stops somebody quoting the ratio alone three months later.
Keep a configuration history with dates alongside the ratios, because a year on nobody reconstructs which change moved which number from memory or from a tool's own log. One row per change: date, setting, old value, new value, who approved it, and what happened to that stage's pass rates. That file is also the only thing that lets you cleanly undo a change that turned out to make the decision worse.
Two adjacent symptoms are usually configuration problems described as tool problems. A keyword filter that stopped separating anybody is one: what replaces an ATS keyword filter when every resume matches. Stage pass rates that all shifted at once is the other, and the question there is which funnel metrics still mean anything. Neither has a model anywhere in it.
What the model is genuinely responsible for
Some of it, and pretending otherwise is its own error. A model contributes the associations it learned from the text it was trained on, the way it treats a document it finds thin, and its stability under changes nobody made deliberately. Those are real, they are not yours to fix, and the response to them is a vendor conversation.
The division resolves in one test. If the disparity survives every configuration change available to you, and survives a stage-by-stage recomputation, it sits upstream and the conversation moves to the vendor: what the tool was validated against, what the last audit covered, and what data you can obtain. If it does not survive, it was yours all along, and the cheapest fix available in hiring is a setting nobody could justify in writing.
Run the test in that order, because the reverse order is expensive and common. A team that opens with the vendor conversation spends weeks on a question it cannot answer, while a knockout question nobody remembers adding keeps removing applicants every day of those weeks. The alternative worth looking for first is a change to a stage before a change of supplier, and that search has an order of its own.
On Monday, export the settings and book thirty minutes. Every setting needs an owner and a written reason by the end of the meeting, and the ones that get neither are the shortlist. The room does not need a data scientist. It needs whoever can log in and whoever knows the job.
Common questions
How do I tell whether a bad ratio came from the model or the configuration?
Recompute stage by stage and change one setting at a time. If the disparity concentrates at a configured step, a knockout or a cutoff or a ranking cut, it is yours and it is fixable this week. If it is spread evenly across every step including ones with no settings, or it survives every change you can make, the question moves upstream to the vendor. Running the test in that order is what makes the vendor conversation short and specific when you eventually have it.
The implementation consultant set most of this. Does that change anything?
Not for responsibility, and it changes a great deal for the review. A consultant configuring a tool is making choices about your selection rates, usually optimising for a workable volume of candidates, and those choices arrive without the job requirement written next to them. Ask for the handover document, put a name against every setting, and treat consultant defaults exactly like vendor defaults: they need an owner and a reason or they come out.
Is a ranked list always worse than a pass bar?
Not always, but it asks more of the evidence. Ranking claims the tool can order people finely enough that position matters, which is a stronger claim than saying a candidate clears a bar, and the 1978 federal validity standard asks for more support where ranking carries greater adverse impact than an appropriate pass bar would. Where a model generates the order, part of the order can also be an artifact of the input rather than of the candidates. If ranking is genuinely needed, keep the bar as the decision and use the order only to sequence outreach.
What counts as an optional signal worth switching off?
Anything the tool measures that nobody in the hiring process asked for: response latency, completion speed, tone or sentiment scoring, social profile enrichment, personality inference and anything derived from how somebody sounds or looks. Each one needs a job requirement written beside it to survive the review, and most cannot get one. Switching them off costs nothing, narrows the surface a candidate can be penalised on for reasons unrelated to the work, and shortens the disclosure conversation.
What if there are too few hires for a ratio to mean anything?
Review the design rather than the arithmetic. A knockout question with no job requirement behind it is worth deleting whether or not a ratio reaches significance, and a cutoff set by a vendor default is worth resetting on the same basis. Keep the counts and keep the configuration history anyway, because both accumulate into something readable over a couple of years. Where the numbers cannot carry the argument, the written record of what was examined has to.
References
- 1. Hidden Workers: Untapped Talent hbs.edu Supports the employment-gap filter as the canonical configured rule: 48% of surveyed employers applying it, stated as a share of employers rather than of resumes.
- 2. Amazon scraps secret AI recruiting tool that showed bias against women web.archive.org Supports the accurate version of the Amazon account and the line usually cut from it: editing out the flagged terms was no guarantee the system would not find other proxies.
- 3. 29 CFR 1607.5 - General standards for validity studies law.cornell.edu Supports the cutoff-score standard at 1607.5(H) and the pass/fail versus ranking evidence distinction at 1607.5(G).
- 4. OpenAI's GPT Is a Recruiter's Dream Tool. Tests Show There's Racial Bias (Methodology, Limitations) web.archive.org Supports the order effect in model-produced rankings, cited as the disclosed limitation it was rather than as the study's headline finding.
- 5. The Americans with Disabilities Act and the Use of Software, Algorithms, and Artificial Intelligence to Assess Job Applicants and Employees web.archive.org Supports the gamified-cutoff example showing that a pass mark is the mechanism creating individual disability exposure, quoted as archived 2022 guidance rather than current law.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.