Policy
A Bias Audit Scores the Tool You Bought, Not the Way You Use It
A bias audit of an AI hiring tool certifies one thing: selection rates and impact ratios for sex, for race or ethnicity, and for their intersections, computed on a data set somebody chose, on one date. It makes no claim that the tool predicts performance in the job it screens for, and says nothing about age, disability or pregnancy. Job-relatedness evidence, accessibility, and the categories the audit never measured stay with the employer, because the duty attaches to how the tool is used, not to the vendor's product.
The takeAn impact ratio computed on a handful of selections is fully compliant and carries no information about the tool, and that limit belongs in the summary you send your own leadership rather than in a footnote nobody reaches. A small employer reading a clean ratio as reassurance has bought the one conclusion the arithmetic cannot support. Say it out loud before somebody quotes the number back to you in a meeting where it decides something.
Where Olive fits
Open a role and see what the work shows
Olive has not performed a bias audit, and olive.is says so: the volume is too low for a four-fifths ratio to carry meaning, and a ratio computed on a handful of sessions would be a number rather than evidence. What a reviewer can be handed instead is six findings written by a person, each carrying the excerpt behind it, exported with the rubric and bank versions the report was produced under.
Rank your shortlistWhat does a bias audit actually compute?
Selection rates and impact ratios by category, and nothing else. Under the New York City rules that define it, a bias audit must compute a selection rate and an impact ratio for each category, separately for sex, for race or ethnicity, and for intersectional categories of sex, ethnicity and race 1. It is group arithmetic on a data set, on a date.
Three limits sit inside that definition and all three point the same way. An independent auditor may exclude a category representing less than 2% of the audit data from the impact-ratio calculation, which most often drops the smallest groups in the pool, the ones a bias audit exists to protect. An employer that has never used the tool may rely on an audit built entirely on other employers' historical data, and any employer may rely on one built on test data where there is too little real data for a statistically significant audit 1. And nothing anywhere in the arithmetic touches accuracy or job-relatedness.
The law that produced this shape is local, and that matters when the summary gets read as a general certificate. In force since January 2023, the New York City rule bars an employer from using an automated employment decision tool on a New York City candidate unless a bias audit was performed within the prior year, a summary of the results is posted publicly, and the candidate received notice at least 10 business days before the tool is used 2. It is a disclosure-and-audit regime rather than a ban: a tool that does badly in its own audit is still lawful to use there, provided the audit is posted.
What does the audit not answer?
It does not say whether the tool predicts anything about performance in this job, whether it works for someone using a screen reader or asking for extra time, or what happens to candidates in a category it could not measure. Those are the three questions a buyer actually has. The last one covers every category the rule never asked the auditor to look at.
Disability is the clearest of the three, and the agency said so plainly while its guidance was still posted. A vendor's bias-free claim usually means steps taken against Title VII adverse impact on race, sex, national origin, color or religion, and those steps are typically distinct from what disability requires: because each disability is unique, an individual can still be screened out regardless of how well other people with disabilities fare on the same assessment 3. That document came off the agency's website in January 2025 and never had the force of law, so it belongs in a memo as what the EEOC said in 2022. The statute underneath it was not touched.
The prediction question is separate evidence entirely, and the two labels get run together constantly. The difference between a validated assessment and a bias-audited one is the short version; whether a vendor's study covers your roles at all is a different check with conditions of its own.
In the first publicly documented cooperative audit of a commercial candidate-screening vendor, the auditors found that pymetrics did faithfully implement its stated fairness guarantees, with safeguards sufficient to reasonably ensure compliance with the four-fifths rule 4. They had also agreed in advance not to question the choice of fairness objective or metric, and did not audit the vendor's annual back-testing of deployed models. Differential validity, intersectional fairness and non-EEOC groups were outside the agreed scope. That is a narrow question answered honestly, in 2020, for one product, and it is not a finding that the tool is fair, valid or job-related.
Ask for the full report and the last retrain date
Ask for the auditor's full report, and ask for the date of the last model change. A compliant summary already has to name any category the auditor dropped under the 2% rule, with that category's own counts and the auditor's justification 1, so check first whether the one you were sent does. The full report is where the method sits. A summary published before a retrain describes a system that no longer exists, and nobody volunteers that.
New York State auditors re-reviewed 22 employers that the city's own compliance sweep had largely cleared, using only public information, and identified at least 17 instances of potential non-compliance where the city had identified one: three audits not performed by an independent auditor, five missing proper selection-rate and impact-ratio calculations, and nine failing to properly explain their use of historical data 5. Those 22 employers were pre-selected by outside researchers precisely because they raised questions, so the rate does not generalise, and none of it has been adjudicated. The failure modes are the useful part, and they are the three things to check yourself: independence, the arithmetic, and whose data it ran on.
Send both requests to every vendor already in the stack, not only to the one under evaluation. Renewal is the moment nobody re-reads anything, and an audit that was current at purchase ages out on a one-year clock 2. Assembling the file a regulator would ask for is a longer list: what an assessment vendor has to hand over.
Why a clean ratio can still be uninformative
Because the four-fifths ratio was never a pass mark, and because a ratio computed on small numbers swings on single decisions. The agencies that wrote the rule said in their own 1979 interpretive questions and answers that it is a rule of thumb and not intended as a legal definition, but a practical means of keeping enforcement attention on serious discrepancies 6. Asked whether it means the guidelines tolerate up to 20% discrimination, they answered no.
The reverse inference gets made anyway, in both directions. Clearing 80% is not a clean bill of health, and falling below it is not a finding of discrimination 6; both readings mistake a screening heuristic for a verdict. Sample size compounds it. The same 1979 document says it is generally inappropriate to require validity evidence or take enforcement action where the numbers are small enough that selecting one different person for one job would flip which group has the higher rate, and its worked example is three men and one woman hired from a pool of twenty men and ten women: the ratio trips four-fifths and the count is still too small to support a determination 6. That is how an audit can be compliant and uninformative at the same time, and why the number deserves its raw counts printed next to it wherever it is quoted.
So the honest summary for whoever signs the contract has two halves. Here is what the audit measured, on whose data, on what date. Here is what it did not measure, and here is the evidence being gathered instead. The second half is where the work sits, and it is the employer's work: the audit describes the tool, and the exposure runs on how the tool is configured and used.
On Monday, send the two requests, then write that half-page: categories covered, data source, audit date, cells the auditor could not fill, and the three questions the document does not answer. It fits on one page, it survives the renewal, and it is the thing somebody will look for in a year.
Common questions
Does a vendor's bias audit satisfy my obligation?
Where an obligation exists, it attaches to the employer's own use of the tool, so a vendor's audit is an input to yours rather than a substitute for it. In New York City an employer that has never used the tool may lawfully rely on an audit built on other employers' data. Once it has used the tool, it may rely on a multi-employer audit only if it gave the auditor its own historical data, and once its own data is statistically sufficient it can no longer fall back on test data. Treat the vendor's document as the starting point and plan for your own numbers.
What is an intersectional cell, and why is mine empty?
It is a combination of categories rather than a single one, such as sex and race together, and it is empty because too few people in the data fell into it. The rule that shaped these reports lets an auditor exclude any category making up less than 2% of the audit data from the impact-ratio calculation. That is a defensible statistical choice and it systematically removes the smallest groups, so an empty cell is information: it tells you which candidates the audit could not describe.
Is a bias audit required outside New York City?
The audit-plus-notice duty in that specific form is a New York City rule, and as of August 2026 no other US jurisdiction requires that exact arithmetic. Other jurisdictions reach the same territory differently, and some make testing evidence relevant to a claim without requiring it. The map changes fast and it turns on where the candidate sits, not where the company does, so treat a compliance answer as a question for counsel.
Can a tool pass a bias audit and still be a bad tool?
Yes, and the two facts are unrelated. A bias audit measures group selection rates, so a tool that predicts nothing at all about job performance can produce even rates across categories and pass. Passing says the outputs were distributed acceptably across the categories measured, on that data, on that date. Whether the tool measures anything worth measuring is a validity question resting on separate evidence, and no audit of this shape is designed to answer it.
How often does an audit need to be redone?
Currency is worth tracking on two clocks, not one. The New York City rule works on an annual clock: a tool may not be used if its most recent bias audit is more than a year old. The second clock is the model itself, because a retrain produces a system the audit never examined regardless of how recently it ran. Ask for both dates, and treat a retrain since the audit as an expired audit whatever the calendar says.
References
- 1. Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY 5-300 et seq.) rules.cityofnewyork.us Supports what the audit must compute, the 2% category exclusion and the disclosure the summary owes when a category is excluded, and the conditions under which an employer may rely on another employer's historical data or on test data.
- 2. Automated Employment Decision Tools: Frequently Asked Questions nyc.gov Supports the three duties (audit within the prior year, published summary, notice at least 10 business days before use) and the one-year currency clock cited in the renewal argument.
- 3. The Americans with Disabilities Act and the Use of Software, Algorithms, and Artificial Intelligence to Assess Job Applicants and Employees web.archive.org Supports the claim that a bias-free claim aimed at Title VII categories does not transfer to disability, because screen-out is an individual question rather than a subgroup rate.
- 4. Building and Auditing Fair Algorithms: A Case Study in Candidate Screening ccs.neu.edu Supports what a passed cooperative audit certifies and what was agreed out of scope in advance, used here to define the boundary of the artifact rather than to praise or fault a vendor.
- 5. Enforcement of Local Law 144 - Automated Employment Decision Tools, Report 2024-N-6 osc.ny.gov Supports the three recurring failure modes a buyer can check without special access: auditor independence, the required calculations, and the explanation of historical data.
- 6. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures (Q.11, Q.19, Q.21, Q.22) eeoc.gov Supports the claim that the four-fifths ratio is a rule of thumb rather than a legal definition, that passing it is not a clean bill of health, and the small-numbers limit at Q.21 and Q.22 including the agencies' worked example of a ratio that trips four-fifths on too few selections to act on.
6 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.