Assessment design
What's the Difference Between a Validated and a Bias-Audited Assessment?
A validated assessment and a bias-audited one answer different questions; a vendor can hold one while failing the other. Validation is evidence the tool samples or predicts one named job's work, built on a job analysis, and it travels to another job only where job analyses on both show substantially the same major work behaviors. A bias audit is arithmetic: selection rates and impact ratios by sex and race, job-agnostic, and allowed to drop categories under 2% of the data. The audit finds adverse impact; only validity answers it.
The takeThe audit won the vendor page because it's the cheap half. Ratios get computed once on pooled data and published as a badge; a validity file has to be rebuilt for every occupation you hire into, and nobody wants to sell that per role. So the industry standardized on the number that generalizes and quietly dropped the evidence that doesn't. The honest guess is that most enterprise buyers can already name their vendor's audit date and not one occupation its criterion study was run on. That is the wrong file to have memorized.
Where Olive fits
Open a role and see what the work shows
If you are asking a vendor these questions, ask Olive them too: the item banks are authored per occupation and carry that occupation's SOC code, and every finding is written by a human reviewer against a timestamped moment in the session rather than by an automated scorer. Olive has not performed a bias audit, because there is not yet enough volume for a four-fifths ratio to mean anything, and olive.is says so.
Rank your shortlistWhat does each one actually prove?
A bias audit proves arithmetic: it reports the rate at which each sex and race/ethnicity group was selected by the tool, and the ratio of each rate to the highest one 3. Validation proves job-relatedness: evidence that the procedure samples or predicts the work of one particular job 1. One is a statement about groups of people. The other is a statement about one job.
| Validated | Bias-audited | |
|---|---|---|
| The question it answers | Does this procedure sample or predict the work of this job? | Do the tool's selection rates differ across sex and race groups? |
| The unit | One job, or jobs shown to share the same work behaviors | One tool, across whoever it has assessed |
| The evidence | Job analysis, plus a criterion or content study, documented per job | Selection rates and impact ratios by category |
| Travels to another role? | Only with transportability evidence 2 | The arithmetic is job-agnostic |
| What a clean result rules out | That the procedure is unrelated to the work | That measured groups were selected at very different rates |
| What it says nothing about | Whether outcomes differ by group | Whether the tool measures anything about the job |
The confusion is profitable, which is why nobody selling an assessment volunteers the distinction. "Independently bias-audited" is a phrase a vendor earns by paying an auditor to compute ratios, often on data pooled from other customers. It is a real obligation in New York City and it is worth doing. It also says nothing whatsoever about whether the tool predicts performance in the role you are filling.
The trap is the direction of travel. Validity evidence does not generalize to a new job on its own; a bias audit's arithmetic does not need to, because it never made a claim about a job in the first place. So a vendor holding a criterion study on customer-service hires and a clean audit has told you two true things, neither of which is about the data-analyst opening on your desk. Before you buy an AI-skills assessment, sort every claim into one of those two columns and see which column is empty.
Why is validation always validation for one job?
Because the evidence is built out of a specific job's content. A criterion-related study starts with a review of job information to find the performance measures "relevant to the job or group of jobs in question" 1. A content-validity claim holds only "to the extent that it is a representative sample of the content of the job" 1. Change the job and the sample you would have to draw changes with it.
Criterion-related validity is a measured correlation between how people did on the procedure and how they did on the job. The Uniform Guidelines treat that relationship as established when it is statistically significant at the 0.05 level 1, which means the study had real applicants or incumbents, in a real job, scored against a criterion someone had to define. Every one of those choices is local to the occupation studied.
Borrowing the study for a second job is allowed, and the conditions are specific: incumbents must perform substantially the same major work behaviors, shown by appropriate job analyses on both jobs 2. That is a document you can ask for by name. If the answer is that the assessment measures general problem-solving and therefore applies everywhere, you have been handed a construct claim with no construct study behind it.
Work the example through. A customer-service criterion study measured a correlation between test performance and something like supervisor ratings or average handle time. A data-analyst role has neither of those criteria. The work behaviors differ, the criterion measures differ, and the labor market the sample was drawn from differs. Nothing about the earlier study is wrong. It is simply about a different job, and the vendor is not lying when they cite it. Whether one assessment can cover every department turns on exactly this paragraph of the Guidelines.
What does a bias audit measure, and what does it skip?
Selection rates and impact ratios, calculated separately for sex categories, race/ethnicity categories, and intersectional categories of the two 3. That is the whole instrument. New York City's rule requires no job analysis, no criterion study, and no showing that the tool relates to the job it screens for. A clean set of ratios is a clean set of ratios, and it is not a statement that the tool works.
The rule most vendors mean by "bias-audited" has edges worth knowing. The audit must be no more than a year old before use. The auditor has to be independent, which the rule defines negatively: not involved in developing, distributing or using the tool, no employment relationship with the vendor or the employer, no direct or material indirect financial interest in either 3. The published summary has to state the audit date, the source and explanation of the data, the number of people assessed who fell into an unknown category, and the applicant counts, rates and impact ratios for every category 3.
Two of those edges change how you read a summary. An auditor may exclude any category representing less than 2% of the audit data from the impact-ratio calculations, provided the exclusion and its justification appear in the summary 3, so the groups with the fewest applicants are the ones most likely to be missing from the table you are shown. And an employer may rely on an audit run on other employers' historical data only if it gave its own historical data to that auditor, or if it has never used the tool 3. "Our vendor is audited" does not survive that sentence once you have been using the tool for a year. Running your own adverse-impact check is a separate exercise from reading a vendor's.
The compliance claim also outruns the compliance. When researchers checked 391 employers for Local Law 144 compliance, they found 18 that had posted a bias audit report and 13 that had posted the required transparency notice 7. Ask for the URL of the posted summary rather than the assurance that one exists.
How do the two fit together under the law?
Sequentially. A bias audit tells you whether a selection procedure has adverse impact; validation is the defense once it does. Under the Uniform Guidelines, a procedure with adverse impact is considered discriminatory unless it has been validated 4. The audit is the trigger and the validity file is the answer, which is why a vendor holding one without the other has brought you half a file.
The four-fifths rule is where the audit's numbers acquire consequences. A selection rate for any race, sex or ethnic group below four-fifths of the highest group's rate is generally regarded by the federal enforcement agencies as evidence of adverse impact, and smaller differences may still constitute adverse impact where they are significant in both statistical and practical terms 5. An impact ratio of 0.84 is not a pass mark. It is a number you need to be able to explain.
Explaining it is the validity file's job. A selection procedure has to be job-related and consistent with business necessity for the job it is used on 6, and the Guidelines add a duty most buyers miss: where two procedures are substantially equally valid, use the one shown to have the lesser adverse impact 4. That makes the audit an ongoing design obligation rather than an annual filing. The documents to have on hand before an inquiry are the two files side by side.
Two honest limits. Local Law 144 compliance is not a Title VII defense. The city asks for arithmetic and publication, not for a validity study. And the local obligation itself keys off whether the tool substantially assists or replaces discretionary decision-making, which the rule defines as relying solely on a simplified output, weighting it above every other criterion, or using it to overrule human conclusions 3. A step that produces no simplified output has a different question to answer, and what a defense actually looks like still starts with the job analysis.
Ask the vendor for these six documents
Six documents, and the answers are short enough to give on a call. The job analysis behind the validity claim. The occupation the criterion study was run on, with its sample size. The criterion measure: what job performance was scored as. The date of the most recent bias audit. The data that audit ran on. And the categories the auditor excluded, with the reason.
1. The job analysis. Which occupation, done when, by whom. Without it there is no validity claim of any kind, because every strategy in the Guidelines starts there 1. 2. The criterion study, with the occupation named. Sample size, the criterion measured, the significance reported 1. "Validated across 200 companies" is not an answer; an occupation is. 3. The transportability memo, if the study was run on a different job: job analyses on both jobs, showing substantially the same major work behaviors 2. 4. The bias audit summary URL. Dated, with the source of the data stated 3. If the data came from other employers, ask what happens when your own year of usage data exists. 5. The excluded categories. Any category under 2% that was dropped from the ratios, and the auditor's justification for dropping it 3. 6. The impact ratios themselves, for sex, race/ethnicity and intersectional categories 3. The table, not a headline sentence about passing.
Then ask the question that sorts vendors fastest: what would this assessment have to show for you to tell a customer not to use it for a role? A vendor with a real validity file can answer that in a sentence, because the file names the jobs it covers and therefore names the jobs it does not. The rest of the vendor question set works the same way: ask for the document, never the adjective. Olive publishes its own version of that file, including the parts not yet measured: See how Olive measures this.
Common questions
Can a vendor be bias-audited but not validated?
Routinely, and it is the most common shape on the market. A bias audit computes selection rates and impact ratios; nothing in it requires a job analysis, a criterion study, or any evidence the tool measures what it claims to. Ask which occupation the validity study was run on. If the answer names a category of jobs, or names your industry rather than your role, there is no validity evidence for your role, only for someone else's.
Does a clean bias audit mean the assessment is fair?
It means the measured groups were selected at similar rates in the data the auditor saw. That data may be pooled from other employers, may exclude any category under 2% of the sample, and covers only sex and race/ethnicity categories (not disability, not age). A ratio above 0.8 is also not a safe harbor: smaller differences can still be adverse impact where they are significant in both statistical and practical terms.
Does a validated assessment still need a bias audit?
Yes, for two separate reasons. If it is an automated employment decision tool used in New York City, the audit is a legal obligation that does not care how well validated the tool is. And under the Uniform Guidelines, validity does not immunize a procedure that has adverse impact: where two procedures are substantially equally valid, the one with the lesser adverse impact is the one you are expected to use. You cannot know which that is without measuring both.
Whose job is the bias audit, the vendor's or ours?
The obligation lands on the employer using the tool, which is why "our vendor is audited" is not an answer. Vendors commonly commission an audit on pooled historical data and hand customers the summary. An employer may rely on that only if it has never used the tool, or if it gave the auditor its own historical data. Settle who posts the summary, on which website, before the first candidate is screened.
What does 'validated' mean if the vendor won't name the job?
Nothing you can use. Every validation strategy in the Uniform Guidelines is anchored to a job: criterion-related validity starts from a review of job information for the job in question, content validity holds only as a representative sample of that job's content, and borrowing a study for a second job requires job analyses on both jobs. A validity claim with no occupation attached is a marketing sentence wearing a technical word.
References
- 1. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 (Technical standards for validity studies) ✓ ecfr.gov A criterion-related study begins with a review of job information to find measures relevant to the job or group of jobs in question; the relationship counts when statistically significant at the 0.05 level; a content-validity strategy supports a procedure only to the extent it is a representative sample of the content of the job.
- 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.7 (Use of other validity studies) ✓ ecfr.gov Validity evidence may be borrowed for another job only where incumbents perform substantially the same major work behaviors, shown by appropriate job analyses on both jobs.
- 3. Notice of Adoption of Final Rule: Automated Employment Decision Tools (6 RCNY §§ 5-300 to 5-304) ✓ rules.cityofnewyork.us A bias audit calculates selection rates and impact ratios separately for sex, race/ethnicity and intersectional categories; the audit must be under a year old; the independent auditor definition excludes anyone involved in developing, distributing or using the tool or holding a financial interest; a category under 2% of the data may be excluded with a stated justification; the published summary must give the date, data source, unknown-category count, rates and ratios; an employer may rely on a multi-employer audit only if it supplied its own historical data or has never used the tool; 'substantially assist or replace discretionary decision making' is defined by sole reliance, dominant weighting, or overruling human conclusions.
- 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.3 (Discrimination defined) ✓ ecfr.gov A selection procedure with adverse impact is considered discriminatory unless validated in accordance with the Guidelines; where two procedures are substantially equally valid, the user should use the one demonstrated to have the lesser adverse impact.
- 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 (Information on impact) ✓ ecfr.gov The four-fifths rule: a selection rate below 80% of the highest group's rate is generally regarded as evidence of adverse impact, and smaller differences may still constitute adverse impact where significant in both statistical and practical terms.
- 6. Employment Tests and Selection Procedures ✓ eeoc.gov A selection procedure must be job-related and consistent with business necessity for the job it is used on.
- 7. Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability ✓ arxiv.org Across 391 employers checked for Local Law 144 compliance, 18 had posted a bias audit report and 13 had posted the required transparency notice.
7 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.