Assessment design
Same Model, Different Configuration, Different Impact
When a hiring tool produces biased outcomes, the bias lives in both the instrument and the configuration, and the useful split is by who can change what. The instrument's properties are the vendor's to document. The cutoff that advances a candidate, the stage the tool sits at, the outcome it was tuned to predict and whether a person may overrule it are yours; each moves impact, and none is what a vendor's audit measured. Ask the deployment questions first; a vendor's fairness evidence covers the vendor's own configuration.
The takeA bias audit certificate is doing something a buyer rarely inspects: answering a question the vendor and the auditor agreed on in advance, about a system running under conditions that are not yours. That is not a criticism of audits, which are worth considerably more than nothing. It is a warning about substitution. The document arrives with the invoice, gets filed under compliance, and quietly replaces the four or five questions about your own deployment that nobody has written down.
Where Olive fits
Open a role and see what the work shows
Any fairness claim attaches to one configuration and one population, which is why Olive publishes no bias-audit claim of its own and says on its site that the volume is too low for such a ratio to mean anything. What it does publish is the configuration: six findings written by a human reviewer against a stated rubric, with every released report exporting alongside its rubric, scorer and bank versions.
Rank your shortlistWhat does a vendor's bias audit actually certify?
A narrow question that was agreed in advance. The first publicly documented cooperative audit of a commercial candidate-screening product, run in summer 2020, was scoped to five questions and found that the vendor faithfully implemented its stated fairness guarantees, with safeguards sufficient to reasonably ensure compliance with the four-fifths rule 1. Read the scope before reading the verdict.
What sat outside that scope, by prior agreement, is most of the list a buyer cares about: the choice of fairness objective and metric, differential validity, intersectional fairness, groups outside the federal categories, and the vendor's own annual back-testing of models already deployed 1. A pass therefore means the code implements the check the vendor says it implements and that neither human error nor a bad actor subverts it easily. It is not a finding that the tool is fair, valid, job-related, or predictive of anything at all.
New York City's rule shows the same shape written into law. Under the Department of Consumer and Worker Protection's 2023 final rules, a bias audit computes selection rates and impact ratios by sex, by race and ethnicity, and by intersectional category; an independent auditor may exclude any category representing less than 2% of the audit data; and an employer that has never used the tool may rely on an audit built entirely on other employers' historical data, or on synthetic data where too little real data exists 2. Three limits pointing the same way. It measures group rates and nothing else, the carve-out most often drops the smallest groups in the pool, and a posted audit need not describe your candidates at all. The difference between a validated assessment and a bias-audited one is the distinction that collapses here most often.
Which part of the impact do you own?
The cutoff, the stage, the tuning target and the override. Each is a customer setting, each moves the selection rate a tool produces, and none of them appears in a vendor's audit, which ran on the vendor's own settings before you had any. The same weights at a different pass mark select a different set of people, so the impact ratio in a certificate is a fact about the cutoff the auditor used.
Agency guidance names the cutoff as the mechanism directly. The EEOC's 2022 technical assistance on the ADA and algorithmic assessment, which carries no force of law and has been read from an archive since the agency removed it in January 2025, says a tool screens somebody out when a disability lowers their performance on a selection criterion and they lose the opportunity as a result, and its worked example is a programmed pass mark: a business requiring 90 percent on a gamified memory assessment rejects a blind applicant who cannot play the game, though that applicant may have an excellent memory 3. That example is the agency's own hypothetical. Screening out is not automatically unlawful, and job-relatedness with business necessity remains a defence. The design point survives all three caveats, because the number that rejects people is a number somebody chose.
The same document undercuts the substitution that makes a certificate feel sufficient. A vendor's bias-free claim usually means steps taken against Title VII adverse impact on race, sex, national origin, color and religion, and those steps are typically distinct from what disability requires, because each disability is unique and an individual can still be screened out regardless of how well other people with disabilities fare on the same assessment 3. That is an agency position with no study cited for it, and it describes a risk without naming any product. It also means a group-rate audit and an individual accommodation route are two separate obligations, and clearing one says nothing about the other.
Whether the tool orders candidates or merely flags them is a configuration choice too, and it changes what the output even is. In Bloomberg's tests both models favoured whichever resume they saw first, naming the first candidate the most qualified 56% of the time for GPT-3.5 and 28% for GPT-4, against a one-in-eight chance if arrival order did not matter 4. The reporters disclosed that as a limitation and neutralised it by shuffling order, so it is not their headline finding, and it is specific to two 2023 model snapshots under one prompt. What it shows is that an ordered shortlist can be partly an artifact of the order you fed in, which makes the ordering a fact about how you called the tool as much as about the candidates in it.
Ask the deployment questions before the vendor questions
Five questions about your own configuration tell you more than a vendor document does, and you can answer all five without anybody's cooperation. They describe what the tool does to a real applicant pool, which is precisely the thing a certificate about an instrument was never about. Write the answers down before the renewal meeting, because they set the agenda for it.
1. What output advances a candidate? The exact threshold, in the units the tool actually emits. 2. What happens to everyone below the line? Rejected, deprioritised, or reviewed by a person anyway. These are three different products built from the same tool. 3. Which stage does it gate? A tool at the top of the funnel touches every applicant; the same tool at the final round touches a population four rules have already shaped. 4. What was it tuned to predict? If the training label was who got hired, the tool learned your previous hiring rule, including whatever was wrong with it. 5. Who may disagree with the output, and how often do they? An override path nobody uses is not a safeguard, it is a line in a contract.
Using a vendor does not move any of this outside the law. California's Civil Rights Council amended the state employment regulations to cover automated-decision systems, effective in October 2025; the regulations name screening resumes for particular terms or patterns and analysing facial expression, word choice or voice in online interviews as examples, and they make an agent acting for an employer, including a vendor running the system, an employer under the Act 5. That expands who can be sued and leaves the buyer's own exposure in place. It imposes no audit or notice duty of its own, and it reaches employers regularly employing five or more people. Whether it reaches your jurisdiction and your setup is a question for counsel.
The vendor questions come after, and they land harder once the configuration is written down. What to ask a screening vendor that says it is bias tested is a better conversation when you can name which cutoff, which stage and which population you are asking about.
Write your configuration down in five lines
One page, five lines, no tooling: the cutoff, the disposition of everyone below it, the stage, the tuning target, and the override path with its real usage rate. That page is the object any fairness question is actually about, and almost nobody has written it, which is how a vendor's document ends up standing in for it at renewal.
With it in hand you have something to read the vendor's evidence against. Treat any fairness claim as attached to one configuration and one population until it has been re-run on yours. Ask which cutoff the audit assumed, whose applicants it was computed on, which categories were excluded for size, and what it measured that is not a group selection rate. If those answers do not come back, that silence is itself information about what you are buying.
Then run it before it gates anybody. A pilot beside your existing round produces the two things no certificate can: your own selection rates, on your own pool, at your own cutoff, and a sense of what the output looks like when a person disagrees with it. How to pilot an assessment vendor before making it a hiring gate is the mechanics of that, and it is the step most often skipped because the audit document feels as though it already happened.
The frame has a limit worth naming. Splitting the question by who can change what is a way of deciding where to look first, not a way of allocating blame. The employer who set the cutoff and the vendor who built the model are both inside the process, and a candidate rejected at that boundary experiences a single decision.
Common questions
Our vendor is bias audited. Are we covered?
The audit describes the vendor's instrument under the vendor's conditions. Your cutoff, the stage you placed it at, the outcome it was tuned to predict and your override path are not in it, and each of those moves the impact the tool produces. Under New York City's rule a first-time buyer may even post an audit computed on other employers' data. Read the audit as evidence about the product, and keep asking about the deployment.
Does changing the cutoff really change the impact ratio?
Whoever sets the pass mark sets the impact ratio. Selection rates are computed at whatever line separates advance from reject, so moving that line moves every group's rate, and the ratio between them follows unless the groups happen to move in exact proportion. The model is untouched throughout. That is why the EEOC's 2022 worked example of an unlawful screen-out is a programmed pass mark rather than an algorithm, and why the cutoff is the most direct lever a customer holds.
Is a bias audit the same as a validity study?
No, they measure different things. A bias audit, as defined by New York City's 2023 rule, the one US rule that requires one, computes group selection rates and impact ratios and stops. A validity study asks whether the instrument predicts performance on the job in question. A tool can pass the first and fail the second, and something that predicts nothing evenly is still something that predicts nothing. Ask for both, and ask which population each was computed on.
Does a debiased tool handle disability?
Usually not. The EEOC's 2022 technical assistance on the ADA, archived since the agency removed it in January 2025, says steps taken against Title VII adverse impact on race, sex, national origin, color and religion are typically distinct from what disability requires, because the disability question is individual rather than a group rate. A tool with even selection rates can still screen out one person whose disability lowers their result on a criterion. That is why an accommodation route has to exist alongside whatever the audit measured.
Who is responsible if the vendor's tool screens people out?
Both parties sit inside the process, and California's employment regulations make an agent acting for an employer, including a vendor operating the system, an employer under the Act. That widens the set of parties who can be sued, and the buyer's own exposure stays where it was. Practically, the buyer chose the cutoff and the stage, so the buyer is the party who can change the outcome next week. How legal responsibility is allocated in a specific case is a question for counsel.
References
- 1. Building and Auditing Fair Algorithms: A Case Study in Candidate Screening ccs.neu.edu Supports what the first cooperative vendor bias audit found and, more importantly here, the list of questions the auditors agreed in advance not to ask.
- 2. Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY 5-300 et seq.) rules.cityofnewyork.us Supports the definition of a required bias audit, the small-category exclusion, and the rule permitting a first-time user to rely on another employer's data.
- 3. The Americans with Disabilities Act and the Use of Software, Algorithms, and Artificial Intelligence to Assess Job Applicants and Employees web.archive.org Supports the programmed cut score as the screen-out mechanism, and the position that a bias-free claim about Title VII categories does not cover disability.
- 4. OpenAI's GPT Is a Recruiter's Dream Tool. Tests Show There's Racial Bias (Methodology, Limitations) web.archive.org Supports the order-sensitivity finding used to argue that asking a tool for an ordering yields something less stable than asking for a pass or a flag.
- 5. Final Unmodified Text of Proposed Employment Regulations Regarding Automated-Decision Systems (Attachment B), 2 CCR sections 11008, 11008.1 calcivilrights.ca.gov Supports the examples of automated-decision systems named in the regulations and the clause making a vendor acting for an employer an employer under the Act.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.