Screening

Bias-Free, Bias-Tested and Audited Are Three Different Claims

Vendor fairness claims about AI hiring tools sort into three buckets: claims about the model's inputs, claims about measured outcomes, and claims about who did the measuring. Only the middle bucket can be checked, and only against a named metric and a named population, so make the vendor state both. Bias-free, bias-tested and independently audited are three different sizes of claim, and only the last carries a definition written outside the company, where a rule supplies one.

The takeThe claim that decides everything operationally is the one no roundup mentions: whether the vendor will hand over the data that lets you compute the numbers without them. A vendor that releases the auditor's full report and a candidate-level extract is making a materially different claim from one that offers a summary and a badge. Everything else is a sentence somebody wrote. Sort the shortlist by what each vendor will actually send, and the fairness section stops being a debate about adjectives.

Where Olive fits

Open a role and see what the work shows

Olive applies the same test to itself: no bias audit has been performed, because the volume is too low for a four-fifths ratio to carry meaning, and olive.is says so rather than calling the assessment fair. What can be handed over is the artifact, six findings written by a human reviewer with the timestamped excerpt behind each one, and the candidate is granted the identical report free.

Rank your shortlist

Which claims can be checked?

Only the ones stating a measured outcome, and only when the metric and the population are named. A claim about inputs describes what went into the model. A claim about auditors describes who looked at it. Neither is a result. A selection rate by category, on this population, on this date, computed this way, is a result, and it is the only shape anybody can argue with.

Three of the four phrases collapse immediately under that test. Bias-free names neither a metric nor a population, so nothing in it can be checked, and the argument underneath it is the no-protected-attributes claim below. Bias-tested is an unregulated phrase that can describe an internal notebook nobody outside the company has seen. Fairness-aware names a training technique, and its effect on your own selection rates is unmeasured until somebody measures it on your data. None of the three is dishonest. None carries information.

There is a measured reason to distrust a single test even when one was genuinely run. Across more than 2 million generated hiring emails covering 300 names, 820 prompt templates and 41 occupations on five models, the absolute difference between the highest and lowest group acceptance rate ran between 1.60% and 3.78%: statistically significant, small in magnitude, and with the direction changing under different templates 1. The task there is writing an acceptance email rather than screening a resume, and the models are 2023 vintage, so the mechanism is the part that travels. A measurement of this kind moves when the prompt moves, which is why a bare tested-it claim with no stated method and no stated population carries no information about your own funnel.

Why the no-protected-attributes claim proves little

Because disparity in these systems travels through correlates rather than through a race field nobody entered. School, postcode, employment gaps, phrasing and a first name all carry associations the model learned. Removing the protected attribute removes the label and leaves the mechanism running, which is why the claim is nearly always true and nearly always beside the point.

The name result is the cleanest demonstration available. An audit of three text embedding models screening 554 resumes against 571 job descriptions across nine occupations found White-associated names significantly preferred in 85.1% of 27 racial bias tests, and Black-associated names in 8.6% 2. Read the denominator before quoting it: 27 is three models times nine occupations, so 85.1% counts the statistical tests that showed a significant preference. The models are open-source research models nobody sells as a hiring product, the run was a retrieval simulation over public resume data, and race was signalled only by a first name placed over a constant surname. Nobody was selected or turned down anywhere in it. What it establishes is narrow and sufficient: a name alone moves the output, with no protected attribute anywhere in the system.

An instruction to be fair moves the number without making the decision stable. A custom model briefed on diversity and disability-justice principles placed a strictly stronger CV first in 37 of 70 trials, against 15 for the unmodified model, a significant improvement that still left the stronger CV in second place about half the time 3. The document in question was the same CV plus four additional accomplishments, so an unbiased system should have ranked it first every time. That was an instruction written into a prompt, measured on one CV against one job description, so it does not test fairness-aware training directly. The general point holds either way: changing what goes into a model is an input claim until somebody measures what it did to selection rates.

What does 'independently audited' guarantee?

Independence in the narrow sense a rule defines, and nothing at all about rigour. Under New York City's 2023 rules for automated employment decision tools, an auditor is not independent if they were involved in using, developing or distributing the tool, if they have an employment relationship with the employer or the vendor during the audit, or if they hold a direct or material indirect financial interest in either 4. That is a relationship test rather than a quality test.

Scope is the other half, and it is agreed before the work starts. In the first publicly documented cooperative audit of a commercial candidate-screening vendor, run in 2020, the auditors reported that pymetrics did faithfully implement its stated fairness guarantees, with safeguards sufficient to reasonably ensure compliance with the four-fifths rule, having agreed in advance not to question the choice of fairness objective or metric 5. Differential validity, intersectional fairness, non-EEOC groups and the annual back-testing of deployed models sat outside that agreed scope. Careful work, honestly reported, answering the question it was handed.

Where a rule supplies the definition, independently audited means somebody with no financial stake answered a pre-agreed question on a chosen data set, on a date. That is worth considerably more than bias-tested and considerably less than the phrase implies in a sales deck. Outside such a rule the phrase has no fixed meaning at all, so the follow-up is always scope: which question, whose data, which categories, what date, and what the auditor was asked not to examine. The longer version is what a bias audit covers and what stays the employer's job.

Require an artifact instead of an assertion

Rewrite the fairness section of the questionnaire so every question asks for a document. An assertion costs a vendor nothing and cannot be checked. A document can be read by somebody who was not in the sales call. Four requests do most of the sorting, and none of them needs a statistician to evaluate the response.

  • The auditor's full report, not the posted summary. The exclusions and the caveats live in the body of the document.
  • The metric and the population behind every fairness number. With neither, the claim is a sentence. With both, it is a measurement somebody else could re-run.
  • A candidate-level data extract, or the terms under which you could obtain one. This is the request that separates vendors, because it decides whether you can compute your own numbers without their cooperation.
  • The date of the last model change. Any fairness result predating a retrain describes a system that no longer exists.

Sorting by what comes back takes an afternoon. A vendor that sends the report and agrees terms for an extract has made a checkable claim; a vendor that offers a badge has made a marketing one, and the difference is visible without any statistical training in the room. The worked version of this conversation, for a buyer holding a bias-tested claim right now, is what to ask to verify it. Validity is a separate question, resting on an entirely separate body of evidence.

On Monday, rewrite the four questions, send them to every vendor on the shortlist and to the one already in the stack, and record what each returned in a single table. The table is the deliverable. Nobody in the review meeting will argue with a column headed "sent the report".

See a sample report

Common questions

Does removing the race and gender fields fix the problem?

Barely, and it is reassuring in none of the ways it sounds. It rules out the crudest failure, a model handed a race or sex field directly, which almost no serious vendor does anyway. It says nothing about correlates, and correlates are the mechanism in question: school, postcode, employment gaps, phrasing and names all carry the association without the label. Treat the claim as a floor everyone clears, and move the conversation to measured outcomes on your own candidates.

How do I tell a real audit from a certification badge?

Ask three questions: who performed it, what were they asked to examine, and may the full report be read. A real audit names the auditor, states the scope including what was excluded, states the data it ran on, and carries a date. A badge names a programme. If the answer to the third question is a summary page rather than a report, you have the marketing artifact, which may still sit on top of real work you are not being shown.

Does a fairness-aware model reduce bias in my funnel?

Unknown until someone measures it on your data, and that is the honest answer rather than a dodge. Fairness-aware names a family of training techniques, and the label describes what was done to the model, so it carries no information about what happens to your candidates. A published fairness instruction in a resume-screening study moved the result substantially and still left a plainly wrong ranking in place about half the time. Your own selection rates on your own applicants are the only thing that answers the question.

Who should run the vendor fairness review?

Whoever runs the evaluation, usually a recruiter or a people-ops lead, with the questionnaire rewritten to ask for artifacts. The skill required is not statistical: it is noticing whether an answer names a metric, a population and a date, and recording what arrived. Bring in counsel for what gets written down and how the tool is described to candidates, and bring in a statistically trained reader once the documents are on the table.

What if a vendor refuses to share an extract of my own candidate data?

Write the refusal down, because it is the single most useful answer in the review. Without a data extract you cannot compute selection rates yourself, which means every fairness claim about the tool remains the vendor's to make and yours to accept. That is a commercial term rather than a technical constraint, so raise it during negotiation when it is still cheap. Some vendors will agree to terms, and which ones do is worth more than any adjective on the website.

References

  1. 1. Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender? Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li and Rachel Rudinger, Proceedings of ACL 2024; read on arXiv, 2024. arxiv.org Supports the claim that a fairness measurement of this kind is prompt-sensitive and small in magnitude, so a single unstated test proves little.
  2. 2. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval Kyra Wilson and Aylin Caliskan, Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2024); full text read on arXiv, 2024. arxiv.org Supports the claim that a first name alone moves the output of embedding-based resume retrieval, with the denominator stated as tests rather than candidates.
  3. 3. Identifying and Improving Disability Bias in GPT-Based Resume Screening Glazko, Mohammed, Kosa, Potluri and Mankoff, University of Washington, ACM FAccT 2024; full text read on arXiv, 2024. arxiv.org Supports the claim that a fairness instruction shifts the result without making the decision stable, and is therefore an input claim rather than an outcome claim.
  4. 4. Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY 5-300 et seq.) NYC Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us Supports the definition of an independent auditor as a relationship test covering involvement, employment and financial interest.
  5. 5. Building and Auditing Fair Algorithms: A Case Study in Candidate Screening Christo Wilson, Avijit Ghosh, Shan Jiang, Alan Mislove, Lewis Baker, Janelle Szary, Kelly Trindel and Frida Polli, Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21), 2021. ccs.neu.edu Supports what a cooperative audit certifies and what was placed outside its scope by prior agreement, used to size the phrase independently audited.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.