Screening

What Do You Ask a Vendor Who Says They're Bias Tested?

A screening vendor's 'bias tested' claim is marketing until you have the audit report. No US law defines what bias testing means and no body certifies it; New York City's Local Law 144 is the only mandate with a written method, and it measures selection rates rather than job-relevance. Ask which jobs the audit covered, whose applicants it ran on, who audited it, and every impact ratio, not the headline. An audit on retail screening proves nothing about your engineering roles, and NYC's posting and candidate notice duties stay yours.

The takeThe audit market grew up around the one city with a written rule, so it measures what that rule counts and stops there. That is why a clean impact ratio keeps arriving with nothing attached about whether the tool measures the job. On what is public so far, no one has shown that audited tools hire better than unaudited ones. So read the report as a disclosure document, not a safety certificate. The question it never asks is the one a plaintiff's counsel opens with: what did this thing measure, and why is that the work?

Where Olive fits

Open a role and see what the work shows

A bias audit tells you whether a tool selected groups at similar rates; it does not tell you whether what the tool measured matters for the roles you are hiring. Olive assesses one candidate's actual work with an AI assistant on an occupational assignment and returns six evidenced findings written by a human reviewer, as an input to your decision rather than a selection step, with the candidate given the same report.

Rank your shortlist

What Does 'Bias Tested' Actually Mean?

Almost nothing on its own. No US law defines a general bias-testing standard for hiring tools, and no certification body issues one. The only mandate with a written method is New York City's Local Law 144, which requires an independent audit within the past year, calculating selection rates and impact ratios separately across sex categories, race and ethnicity categories, and intersectional categories 1.

That method is narrow in ways vendors rarely volunteer. It measures whether groups came through the tool at similar rates. It does not ask whether the tool measures anything relevant to the job, does not require the audit to be run on your applicants, and says nothing about which occupations the audited pool contained. A tool can pass on every ratio and still be sorting on something you would never defend in a deposition.

Everything that method skips is why a bias audit and a validation study answer different questions. An audit is a fairness measurement on a population. A validation study is an argument that the scores predict job performance. Vendors frequently market the first as if it were the second, and the word "tested" is what does the work.

Ask These Twelve Questions

Send them as written. The first five settle whether the audit describes the tool you would actually run, and the rest settle who did the work and what it proves. Next to each is the reply that should end the conversation, or at least move it to your counsel before anything gets signed.

1. Which jobs or job families did the audit cover? Worrying: "The audit covers the tool," or "it is model-level." An audit has a population, and the population has occupations in it. 2. What data did the audit run on: our applicants, other customers' applicants, or synthetic test data? Worrying: test data, offered without a reason. Under the NYC rules, test data is permitted only when there is not enough historical data for a statistically significant audit, and the published summary must explain why historical data was not used 1. 3. How was any test data generated and obtained? Worrying: "proprietary." The same rule requires that description in the summary, so a vendor claiming NYC compliance already owes you the answer. 4. What is the audit date, and which model version was audited? Worrying: an audit date that predates the last retrain, or "the model improves continuously." A continuously changing model has an audit with an expiry date, and NYC sets that at one year 1. 5. What selection threshold was the audit run at, and is it the one we will configure? Worrying: the audit measured the raw score distribution while your deployment applies a cutoff. Impact ratios move with the threshold. 6. Who is the auditor, and what is their relationship to the vendor? Worrying: any answer involving the build. The NYC definition excludes anyone involved in using, developing or distributing the tool, anyone in an employment relationship with the vendor or the employer, and anyone with a direct or material indirect financial interest in either 1. 7. What was the auditor engaged to test, and what were they asked not to test? Worrying: a scope note you cannot see. Scope limitations are where a clean report is manufactured. 8. Send the impact ratio for every category, including intersectional ones. Worrying: sex only, or one summary sentence. NYC requires all three groupings separately, and the intersectional figures are usually the lowest 1. 9. How many assessed people were excluded from the calculations, and why? Worrying: a large unknown-category count, or a category dropped with no justification. The rules allow an auditor to exclude a category under 2% of the data, but the summary must carry the justification and that category's own numbers 1. 10. What did the scores predict, and against what measure of job performance? Worrying: "it predicts which candidates our clients advance." That is a model of the old screen, not evidence about the work. 11. What invalidates the audit on our side? Worrying: "nothing." Changing the job description, the weightings, the question set or the cutoff changes the tool being audited. 12. Who publishes the summary of results, and who sends candidate notice? Worrying: "we handle compliance for you." In New York City those duties sit on the employer, not the vendor 1.

Twelve is the working set, not a ceiling. If a vendor answers all of them cleanly, you have moved from a claim to a document you can hand to legal.

Why Job Coverage Separates Evidence From Marketing

Because an impact ratio is a property of a population, not of a model. A tool audited on a pool of high-volume retail applications was measured against that pool's base rates, its resume conventions and its selection threshold. Run the same tool over a few hundred senior engineering applications and none of those hold. The audit does not transfer, and nothing in Local Law 144 asks it to 1.

Federal guidance is stricter than the city rule here, and it is worth quoting at the vendor. Under the Uniform Guidelines on Employee Selection Procedures, borrowing a validity study from another user requires evidence of job similarity: incumbents in your job and in the studied job must perform substantially the same major work behaviors, shown by job analyses on both 4. A vendor asking you to accept a retail-pool audit for an engineering req is asking for exactly that transfer without exactly that evidence.

The small-numbers problem points the same way. The Uniform Guidelines note that large differences in selection rate may not indicate adverse impact when they rest on small numbers and are not statistically significant 3. Your senior finance pipeline is small numbers. So an audit on your own roles may be honestly inconclusive, which is a different and much more useful answer than a borrowed ratio of 0.94. If you are running an adverse impact analysis on your own funnel, you already know how quickly the cells empty out.

What to Do When the Vendor Won't Answer

Treat the refusal as an answer and write it down. Vendors decline for two distinct reasons: the audit does not exist in the form claimed, or it exists and does not say what the deck says. Both are decision-grade. Ask for the refusal in writing, file it with the procurement record, and price the tool as unaudited for your roles.

A usable middle path exists, and good vendors take it. Ask them to run an audit on a pilot cohort of your own applicants for one req, with an auditor you both accept, before the tool touches a live funnel. That is contractible. It also produces the artifact you will want if anyone asks later, which is a dated report naming the job, the data and the auditor, rather than a badge on a website.

Put the twelve questions in the security-and-compliance section of the RFP rather than the demo, because the demo is answered by a salesperson and the RFP is answered by someone who has to be right. The rest of what to check before buying an AI assessment tool belongs in the same document.

Check What You Have to Publish Yourself

The vendor's audit does not discharge your obligation. Where Local Law 144 applies, you publish the summary of results on the employment section of your own website before use: the source and explanation of the data used, the number of assessed individuals who fell into an unknown category, and the number of applicants, the selection or scoring rates, and the impact ratios for all categories 1.

You also owe candidates notice at least ten business days before the tool is used, with instructions for requesting an alternative process or an accommodation, plus published information on what data the tool collects, where it came from and how long it is kept 1. A vendor that offers to "handle it" is offering something it cannot deliver, because the posting duty attaches to the employer's careers page.

The market norm is not a defense. When 155 investigators checked 391 New York City employers against the law, 18 had posted an audit report and 13 had posted a transparency notice, and nearly every audit that was posted reported an impact ratio above 0.8 2. The authors read that pattern as employer discretion over scope rather than clean tooling. Assume any figure you are shown was produced under the same discretion, and keep the underlying report where you can find it, because the evidence a regulator or plaintiff's counsel will ask for is the report, not the ratio.

See a sample report

Common questions

Is a bias audit the same as a validation study?

Not the same thing. A bias audit compares selection or scoring rates across demographic categories and reports impact ratios. A validation study argues that the scores relate to job performance, and under the Uniform Guidelines that argument is job-specific: borrowing one from another employer requires job analyses showing substantially the same major work behaviors 4. A vendor can hold a clean audit and no validity evidence at all for your occupation. Ask for both, separately, and do not let one answer stand in for the other.

Does a passing impact ratio mean the tool is safe to use?

No. The four-fifths rule is a rule of thumb, and the Uniform Guidelines say so in both directions: smaller differences can still be adverse impact where they are statistically and practically significant, and larger differences may not be where the numbers are small and unstable 3. A ratio above 0.8 on someone else's applicant pool is not a finding about your pipeline. It is a finding about theirs.

Can we rely on the vendor's audit for NYC Local Law 144?

Sometimes, with conditions. An employer may rely on an audit using other employers' historical data only if it has never used the tool, or if it supplied its own historical data to the auditor for that audit 1. Either way the audit must be under a year old, and you still publish the summary of results and the distribution date on your own site and give candidates notice. Reliance covers the audit, not the posting.

What if the vendor used test data instead of real applicants?

Ask why, and expect the answer in the published summary. Under the NYC rules, test data is permitted only when there is not enough historical data to run a statistically significant audit, and the summary must explain why historical data was not used and describe how the test data was generated and obtained 1. Test data with no stated reason is the single clearest signal that the audit was produced for marketing rather than for measurement.

Who counts as an independent auditor?

Someone able to exercise objective and impartial judgment across the whole audit scope. The NYC rules disqualify anyone who was involved in using, developing or distributing the tool, anyone holding an employment relationship with the vendor or with the employer during the audit, and anyone with a direct or material indirect financial interest in either 1. Ask for the engagement letter, not the logo. A consultancy that also sells the vendor implementation work is not independent in the sense the rule means.

References

  1. 1. Notice of Adoption of Final Rule: Automated Employment Decision Tools (6 RCNY §§ 5-300 to 5-304) NYC Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us What a Local Law 144 bias audit must calculate, the independent-auditor definition, the historical-versus-test-data rule, the 2% exclusion, and the employer's own publication and notice duties.
  2. 2. Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability Wright et al., ACM FAccT (arXiv:2406.01399), 2024. arxiv.org Of 391 New York City employers checked by 155 investigators, 18 posted an audit report and 13 posted a transparency notice; nearly all posted audits reported an impact factor above 0.8.
  3. 3. 29 CFR 1607.4 - Information on impact Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov The four-fifths rule and its two caveats: smaller differences can still be adverse impact, larger differences may not be where numbers are small and not statistically significant.
  4. 4. 29 CFR 1607.7 - Use of other validity studies Uniform Guidelines on Employee Selection Procedures, eCFR, 1978. ecfr.gov Borrowing a validity study requires job similarity shown by job analyses on both jobs, plus validity and fairness evidence.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.