Policy

How Do You Run an Adverse Impact Audit When the Vendor Holds the Data?

You don't need the vendor's data to audit an AI screening tool for adverse impact. Pull every applicant the tool touched from your ATS, tag race, sex and ethnicity, and compute each group's selection rate at the step the tool decides, divided by the highest group's rate. Under 0.80 is evidence of adverse impact rather than proof, above 0.80 clears nothing, and on small groups the ratio can mean nothing. What you can't compute yourself is validity evidence for your occupation, and that half decides whether the tool stays.

The takeNotice which half of the Uniform Guidelines the market has priced. A bias audit is one arithmetic pass over a pool, sellable as a certificate. The validity study behind it is a job analysis per occupation, and nobody wants to buy one of those per job family. So the compliance product on offer is the ratio, while the thing that answers a charge under 29 CFR 1607.3 is the study. I would expect the first employer to lose one of these to lose with a clean audit in hand and no job analysis behind it. A certificate is procurement hygiene. It is not a defense.

Where Olive fits

Open a role and see what the work shows

An audit needs a record of what a tool did to a named applicant, and Olive produces no automated decision to audit at all: a human writes all six findings, each anchored to a timestamped excerpt from the session, and every released report exports with its rubric, scorer and bank versions attached. Olive has not performed a bias audit of its own and says so on its site, because there is not yet enough volume for a four-fifths ratio to mean anything.

Rank your shortlist

What Data Do You Need, and Who Actually Has It?

Most of it is already in your ATS. For every applicant the tool touched you need three things: the tool's own outcome, the date, and the applicant's self-identified race, sex and ethnicity. The vendor holds the score distribution, the threshold and its cross-customer audit. You hold the applicant flow, and the Uniform Guidelines put the recordkeeping duty on the user of a selection procedure rather than on the vendor 1.

Three files, one join key. The first is the applicant flow export from your ATS: every person who reached the step the tool runs at, with a requisition ID and a date. The second is the tool's outcome for each of them, either advanced or not, or a raw score plus the threshold that was applied. The third is your voluntary self-identification data, kept in the sex, race and ethnic categories the Guidelines name and the EEO-1 series uses 1. Join on the application ID rather than on email; candidates reapply.

Count the people you cannot classify, and keep that number in front of you. Applicants who declined to self-identify are not a rounding error. They are a separate line in the summary a New York City bias audit has to publish, which states the number of individuals the tool assessed who fall within an unknown category 5. If that count is large, your ratio describes a subset, and the write-up should say which subset.

Compute at the tool's step, not at the offer. Checking the funnel end to end and stopping when the hires look balanced is tempting, and section 4C of the Guidelines does say the federal agencies will not usually expect component-by-component evaluation where the total selection process shows no adverse impact 1. That is prosecutorial discretion, not a legal defense. In Connecticut v. Teal the Supreme Court held that an employer's nondiscriminatory bottom line neither prevents employees from establishing a prima facie case nor supplies the employer with a defense against one 4. A screen that eliminates people is auditable on its own terms.

Run the Four-Fifths Calculation on Your Own Applicants

Take the applicants the tool assessed, group them by race, sex and ethnicity, and compute each group's selection rate at the step the tool decides. Divide each rate by the highest group's rate. A rate below four-fifths, or 80 percent, of the highest group's rate is generally regarded by the federal enforcement agencies as evidence of adverse impact 1.

Here is one requisition: a data analyst role, 1,180 applicants through an AI resume screen, 345 advanced to a recruiter call.

GroupApplicantsAdvancedSelection rateImpact ratio
White (not Hispanic or Latino)42014735.0%1.00
Asian (not Hispanic or Latino)31010232.9%0.94
Hispanic or Latino1804726.1%0.75
Black or African American (not Hispanic or Latino)1503825.3%0.72
Two or more races (not Hispanic or Latino)361130.6%0.87
Declined or unknown84not appliednot appliednot applied

The arithmetic is one division per row. 147 of 420 is 35.0 percent, the highest rate, so it becomes the denominator for every other row. 38 of 150 is 25.3 percent, and 25.3 divided by 35.0 is 0.72. Two groups sit below 0.80, which is evidence of adverse impact for those groups 1. Rebuild the same table for sex categories, and where Local Law 144 applies, for intersectional categories of sex, ethnicity and race as well 5.

Two caveats travel with the number, and they point in opposite directions. Smaller differences can still constitute adverse impact where they are significant in both statistical and practical terms, or where a user's actions have discouraged applicants disproportionately. Greater differences may not constitute adverse impact where they rest on small numbers and are not statistically significant 1. The 36-applicant row above is in that second category: one more advance moves it to 0.95. A ratio of 0.79 across 4,000 applicants and a ratio of 0.79 across 36 are not the same finding, and reporting them the same way is how an audit loses its reader.

If the tool scores instead of passing and failing, the arithmetic changes shape. New York City's rules compute a scoring rate, meaning the share of a category scoring above the sample's median, and take the ratio against the highest-scoring category 5. Run both where a model produces a score that a human then cuts: the scoring rate describes the model, the selection rate describes what you did with it.

Why the Vendor's Bias Audit Doesn't Cover Your Funnel

Because an impact ratio is a property of the pool it was measured on. The vendor's audit pooled applicants across its customers, with different job families, different sourcing and a different threshold from yours. New York City's rule says so directly: an employer may rely on a bias audit using other employers' historical data only if it gave the auditor data from its own use of the tool, or if it has never used the tool at all 5.

The threshold is the second reason. A vendor's audit is run at whatever cutoff the audited population was scored against. If your account advances the top 20 percent rather than everyone above a fixed score, the ratio you were shown describes a different procedure from the one you operate. The Guidelines are explicit that evidence sufficient to support a procedure on a pass/fail basis may be insufficient to support the same procedure used for ranking, and that a user choosing to rank needs validity and utility evidence supporting that method of use 3. Cutoff scores carry their own condition: they should normally be set to be reasonable and consistent with normal expectations of acceptable proficiency within the work force 3.

The third reason is the job mix. An audit pooled across customers is a weighted average of their job families. If most of the audited volume came from high-volume support hiring and you are screening data analysts, the pooled ratio is telling you about the pool rather than about the occupation. That is the same gap as the difference between a bias audit and a validation study: an audit measures who came through at what rate, and says nothing about whether what was measured matters for the work.

None of which makes the vendor's audit worthless. It is the right artifact to ask for during procurement, alongside the documents worth demanding by name. It is simply not your audit, and it will not answer a charge about your funnel.

What Validity Evidence Do You Need Behind the Number?

A ratio under 0.80 does not automatically retire the tool. Under the Uniform Guidelines a selection procedure with adverse impact is considered discriminatory unless it has been validated 2, so the second half of an audit is the job-relatedness file: a job analysis of the occupation you are hiring for, the map from what the screen measures to an important work behavior in that job, the cutoff derivation, and the alternatives investigated 23.

Job-relatedness is occupation-specific, and that is what makes one vendor study thin across a rollout. A content validity study rests on a job analysis of the important work behaviors of the job in question, and the procedure has to be a representative sample of those behaviors or of the work product they produce 6. The critical work behaviors of a data analyst under SOC 15-2051 and a management analyst under SOC 13-1111 are not the same behaviors, so the job analyses differ, so the sample of behavior the screen takes has to differ too. One study does not stretch across six job families by assertion.

Borrowing a study is allowed and expensive in paperwork. To rely on a criterion-related study run elsewhere, incumbents in your job and in the studied job must perform substantially the same major work behaviors, shown by job analyses on both jobs, and the studies must include a study of test fairness for each race, sex and ethnic group that is a significant factor in your relevant labor market 7. Where there are variables in the other study likely to affect validity significantly, you may not rely on it at all 7.

The alternatives question belongs in the same file. Where two procedures serve the same legitimate interest in efficient and trustworthy workmanship and are substantially equally valid, the Guidelines say to use the one demonstrated to have the lesser adverse impact, and an investigation of suitable alternatives is part of the validity study rather than a separate courtesy 2. That is what turns an audit into a decision. A ratio of 0.72 plus a strong occupational validity file plus no less-impactful alternative is a position you can defend; the same ratio with a certificate and no study is not. It is also why whether the score predicts anything about the job stops being an academic question the moment a ratio comes back low.

What to Do When the Vendor Won't Hand Over the Data

Ask in writing, date the request, and file the refusal with the procurement record. Then run the half you control: the applicant-level outcome sits in your ATS whatever the vendor releases, so selection rates at the tool's step are computable without the vendor's cooperation. What a refusal actually costs you is the validity file, and that is a reason to stop using the tool rather than a reason to stop auditing.

Fix it in the contract, at renewal if not before. Four clauses do most of the work: an applicant-level export of scores, thresholds and outcomes in a machine-readable file; a right to hand your own historical data to an independent auditor, which New York City requires if you intend to rely on a multi-employer audit 5; a cooperation clause binding the vendor to support a charge response at a stated rate; and a legal hold you can trigger yourself. Vendors who hold the file sign those clauses. Vendors who negotiate the clause instead of the price have answered the question.

Do not stop auditing because the vendor stopped answering. Where a user has not maintained data on adverse impact as the documentation requirements call for, a federal enforcement agency may draw an inference of adverse impact from that failure, if the group is underutilized in the job category compared with the relevant labor market 1. Absent records are not neutral records.

While the file is missing, take the tool off the decision. Run it beside the screen you already trust, log both outcomes, and advance nobody on the tool's output alone. That accumulates your own historical data, which is the thing an auditor needs and the thing New York City's reliance rule asks you to supply 5, and it keeps a period when you cannot document job-relatedness from also being a period when a procedure with adverse impact is deciding anything 2. Candidates rejected during a parallel run were rejected by your existing process, which is also the honest sentence to write when a candidate asks why they were screened out.

Read the evidence

Common questions

Does a balanced bottom line mean the screen doesn't need auditing?

No. Section 4C of the Uniform Guidelines says the federal agencies will not usually expect a user to evaluate individual components where the total selection process shows no adverse impact, but that is enforcement discretion rather than a defense 1. In Connecticut v. Teal the Supreme Court held that a nondiscriminatory bottom line neither precludes employees from establishing a prima facie case nor provides the employer with a defense to one 4. A screen that eliminates individuals is challengeable on its own, whoever else was hired downstream of it.

What if most applicants don't report race and sex?

Run the audit on the ones who did and publish the count of the ones who didn't. Self-identification is voluntary, and the Guidelines expect records maintained by sex and by the listed race and ethnic groups 1. New York City's rules require the published summary to state how many assessed individuals fall within an unknown category 5. If a large share declines, treat the ratio as provisional and fix the collection point, which is usually a self-identification form shown after submission rather than inside the application.

Is an impact ratio above 0.80 a safe harbor?

No. Four-fifths is what the federal enforcement agencies generally regard as evidence of adverse impact, not a line that clears a tool. Smaller differences may nevertheless constitute adverse impact where they are significant in both statistical and practical terms, or where a user's actions have discouraged applicants disproportionately on grounds of race, sex or ethnic group 1. On a large applicant pool a ratio of 0.85 can be statistically significant. Run a significance test alongside the ratio rather than instead of it.

Can you rely on the vendor's NYC bias audit for your own roles?

In two situations. Under the Local Law 144 rules an employer may rely on a bias audit using other employers' historical data if it supplied historical data from its own use of the tool to the independent auditor, or if it has never used the tool 5. A first-time user may also rely on an audit run on test data where there is not enough historical data for a statistically significant audit, with the summary explaining why historical data was not used 5. Once you have your own volume, that second path closes.

How many applicants do you need before the ratio means anything?

The Guidelines set no number and address the problem directly: greater differences may not constitute adverse impact where they rest on small numbers and are not statistically significant, and where impact evidence is based on numbers too small to be reliable, impact over a longer period or from the same procedure used in similar circumstances elsewhere may be considered 1. New York City's rules let an auditor exclude a category under 2 percent of the audit data, with a written justification and that category's rate still published 5. Aggregate across requisitions in one job family first.

Who is supposed to run the audit, the employer or the vendor?

Both, for different pieces. Each user maintains the impact records for its own selection procedures 1, and a user obtaining a procedure from a publisher is cautioned that it remains responsible for compliance and should confirm the information supporting validity will be made available to it 7. Where Local Law 144 applies, the bias audit itself must be performed by an independent auditor with no employment relationship and no direct or material indirect financial interest in the vendor or the employer 5. The vendor commissions that audit; you own the file that answers for your funnel.

References

  1. 1. 29 CFR 1607.4 - Information on impact Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov The four-fifths rule and both of its caveats, the recordkeeping duty on the user by sex and by the listed race and ethnic groups, the 'bottom line' paragraph in section 4C, and the inference an agency may draw where a user has not maintained impact data.
  2. 2. 29 CFR 1607.3 - Discrimination defined: Relationship between use of selection procedures and discrimination Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov A selection procedure with adverse impact is considered discriminatory unless validated, and the investigation of suitable alternative procedures with less adverse impact is part of the validity study.
  3. 3. 29 CFR 1607.5 - General standards for validity studies Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov Evidence sufficient for pass/fail use may be insufficient for ranking use, cutoff scores should be consistent with normal expectations of acceptable proficiency in the work force, and documentation is required where a procedure with adverse impact is part of the process.
  4. 4. Connecticut v. Teal, 457 U.S. 440 (1982) Supreme Court of the United States, via Cornell Legal Information Institute, 1982. law.cornell.edu A nondiscriminatory 'bottom line' neither precludes employees from establishing a prima facie case of disparate impact on a component barrier nor provides the employer with a defense to such a case.
  5. 5. Notice of Adoption of Final Rule: Automated Employment Decision Tools (6 RCNY 5-300 to 5-304) NYC Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us The impact ratio and scoring rate definitions, the requirement to compute sex, race/ethnicity and intersectional categories separately, the 2% exclusion, the unknown-category count, the independent-auditor definition, and the section 5-302(a) rule limiting reliance on an audit built from other employers' historical data.
  6. 6. 29 CFR 1607.14 - Technical standards for validity studies Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov Content validity requires a job analysis of the important work behaviors of the job in question, and the procedure must be a representative sample of those behaviors or of the work product.
  7. 7. 29 CFR 1607.7 - Use of other validity studies Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov Borrowing a criterion-related study requires substantially the same major work behaviors shown by job analyses on both jobs plus test-fairness evidence, cannot be relied on where other variables are likely to affect validity significantly, and users remain responsible for compliance when buying from a publisher.

7 sources, numbered by first appearance. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.