Policy
How Do You Defend an AI-Skills Assessment to Legal and to Candidates?
Defend an AI-skills assessment with documents, not assurances. Legal needs four: a job-relatedness file showing the assessment samples what the occupation performs, a record of who advanced by race, sex and ethnic group, a named human who makes the decision, and the notice candidates got before sitting it. Candidates need one more, their own copy of the result. Rebuild the file for every occupation, because validity evidence does not transport between jobs, and send notice before the tool runs, ten business days ahead where New York City law reaches you.
The takeThe 1978 Guidelines were written for paper tests, and they still ask the only question that matters here: what this particular job requires. Nothing about a model changes that burden. What changes is when anyone writes it down. On what is public so far, these disputes come down to whether the folder existed before the dispute did. Evidence assembled after a complaint reads as an argument to everyone in the room, however solid it is on the page. So the strength of a defense here is mostly fixed months before anyone thinks to test it.
Where Olive fits
Open a role and see what the work shows
Legal asks who decided and on what evidence; the candidate asks the same question in different words. Olive answers both with one artifact: six findings a human reviewer wrote, each anchored to a timestamped moment in the session, exported with its rubric, scorer and bank versions, and granted to the candidate in identical form.
Rank your shortlistWhat does legal actually need to see before signing off?
Four things, and all four are documents. A job-relatedness file showing the assessment measures what this occupation performs. A record of selection rates by race, sex and ethnic group. A named human who makes the decision the assessment informs. And the notice you gave candidates before they sat it. Legal's question is not whether the tool is good. It is whether you can produce those four on request.
Start with the standard, because it names the burden. When a selection procedure produces disparate impact, the employer has to show it is job-related and consistent with business necessity, which the EEOC states as necessary to the safe and efficient performance of the job, and which it ties to "the particular job in question" rather than to general skill 2. Most rollouts fail on that phrase, not on the choice of vendor.
The Uniform Guidelines put the documentation burden on the user. Any selection procedure that is part of a process with adverse impact has to be validated, with the evidence maintained and available, through criterion-related, content or construct validity 3. A vendor's technical report is an input to that file. It is not the file.
The impact record is separate, and it starts on day one. Section 1607.4 asks each user to keep records disclosing the impact its selection procedures have on people by identifiable race, sex or ethnic group, and treats a selection rate below four-fifths of the highest group's rate as evidence of adverse impact 1. You cannot reconstruct that later if nobody recorded who was invited and who advanced. Stand the adverse-impact numbers up before the first cohort, not after the first complaint.
The fourth document is the one people skip: who decided. Keep the assessment an input a named person reads, with their reasons written down. Where a tool is used to make or substantially inform a hiring decision, notice and audit duties can attach. New York City requires a published bias audit summary and at least ten business days' notice to candidates before use 5.
How do you write the evidence file for one role?
Build it as one folder per role, assembled before the first invite goes out. It holds the task analysis for the occupation, the assessment content mapped task by task, the rubric and its version, the names of the people who review, the decision rule, and the running record of who was invited and who advanced by group. Write it while the reasoning is still cheap to recall.
What goes in, concretely:
- The task analysis. The work behaviors this role performs, their importance and their difficulty, sourced from people doing the job rather than from the job posting.
- The content map. Each part of the assessment against the behavior it samples. A row with no behavior beside it is a row to cut.
- The rubric, versioned. What counts as demonstrated, what does not, and which version was in force on the date each candidate sat it.
- The reviewer record. Who read the session, what they wrote, and the calibration they went through before reading anything that counted.
- The decision rule. What the hiring team does with the result, written before the results exist. "Considered alongside the interview" is a decision rule; so is a cut score, and a cut score needs a source.
- The impact log. Invited, started, advanced, by group, from the first cohort.
The rule of thumb: anything that would sound like a judgment call in a deposition should be a dated document instead. That is also the practical difference between a defensible AI-skills assessment and one that merely produces a defensible-sounding report: the file exists before the dispute, or it does not exist.
If you bought the assessment rather than built it, the same folder still belongs to you. Ask the vendor for the job analysis behind each bank, the rubric version history, and the reliability evidence between reviewers. Then get it in writing before the pilot, because a sales deck is not documentation.
What do you tell candidates, and when?
Before they start, in plain language: what the task is, what is captured, which behaviors are assessed, who reads the result, and how long it is kept. After it is reviewed, give them the result itself. New York City requires at least ten business days' notice before an automated employment decision tool is used, plus a published bias audit summary 5. Treat that as the floor everywhere, not as a New York obligation.
The objection you are answering is specific, and it is not an objection to being assessed. Pew Research Center found Americans oppose the use of AI in making final hiring decisions by 71% to 7%, while 47% said AI would do better than humans at treating all applicants similarly 6. People are not rejecting structure. They are rejecting being decided by a machine with nobody in the room.
So the sentence you owe a candidate is short: a person read your session, wrote what they saw, and here is what they wrote. Report parity is the cheapest defensible thing on this list: the identical document, with the identical excerpts, goes to the candidate. It also disciplines the writing, because nothing gets recorded about a candidate that you would not hand to them. Decide what a candidate-facing report exposes before the first release, not after someone forwards one to a lawyer.
Say how to request an accommodation, and say it before the assessment rather than inside the rejection. A timed task, a microphone, a screen-capture requirement and a fixed session length are each a place where the process can turn a disability into a score. Offer the alternate path up front and record that it was offered.
A finding carries the moment it rests on, so it can be discussed, disputed and corrected. A composite number standing for the person carries none of that. It never goes in the report, and it is the single artifact most likely to be read back to you.
Which questions will legal ask that you can't answer yet?
Five, and they arrive in the same order. Who makes the decision. What the cut is and where it came from. What the selection rates look like by group. What the vendor will put in writing. And whether a less discriminatory alternative existed, because a job-related procedure still loses if a challenger shows one did 2. Most rollouts have an answer for the first four and nothing for the fifth.
The fifth question is answerable, but only with evidence you have to go and collect. Run the assessment beside your existing round rather than in place of it, keep both sets of results, and record what each one added. That comparison is what turns "we chose this" into "we tested this," and piloting before the gate is cheaper than discovering the answer during discovery.
On the cut: a threshold invented in a meeting is the weakest link in the file. If the decision rule is a bar, the bar needs a source: a job analysis that says what level of the behavior the work requires, or expectancy evidence from the pilot. Where neither exists yet, keep the result an input a human weighs and say so in the rule.
Two answers legal will accept faster than any argument about the tool: the reviewer's notes, and the version strings. Evidence anchoring means every conclusion points at a moment somebody can open: a timestamp, a transcript turn, a diff, a written answer. Versioning means the report cannot quietly change after release. Together they turn a defense of the assessment into a defense of a specific record about a specific person, which is the only defense that survives being tested one candidate at a time.
Common questions
Does an AI-skills assessment need a bias audit before we use it?
It depends on where you hire and how the result is used. New York City requires a published bias audit summary and at least ten business days' notice to candidates before an automated employment decision tool is used 5. Federally, the Uniform Guidelines do not require an audit by that name, but they do ask every user to keep records of the impact its selection procedures have by race, sex and ethnic group, and treat a selection rate under four-fifths of the highest group's as evidence of adverse impact 1. Keep the log from the first cohort either way.
Can we use the same assessment for every role?
You can use the same instrument; you cannot reuse the same defense. The Uniform Guidelines Q&A says validity found in one situation does not necessarily hold in different circumstances, and lists four conditions before evidence is borrowed across jobs, including showing the jobs comparable on a work-behavior basis 4. In practice: one task analysis, one content map and one impact log per occupation. The behaviors you assess can stay constant across roles. The occupational material and the evidence behind it cannot.
What do you have to tell candidates before an AI-skills assessment?
What the task is, what is captured, what is assessed, who reads it, how long it is kept, and how to request an accommodation. In New York City that notice has a deadline attached: at least ten business days before the tool is used 5. Everywhere else it is still the right default, because disclosure before the session is the only version a candidate can act on. Send it before any account step, so nobody has to sign up to find out what they agreed to.
Does a human reviewing the result remove the legal risk?
It changes what you have to defend, and only if the review is real. A named person who reads the evidence, writes the reasoning and can be asked why is a decision-maker; a person who clicks through a recommendation is a rubber stamp with a name on it. Either way the underlying procedure still has to be job-related and consistent with business necessity where it produces disparate impact 2. Human review is a control on how the result is used, not a substitute for evidence that it measures the job.
What if there are too few candidates for a four-fifths analysis?
Log the data anyway and say plainly that the sample is too small to interpret. The four-fifths ratio is a rule of thumb federal enforcement agencies generally treat as evidence of adverse impact 1, and a ratio computed on eight candidates is noise wearing a number. Small volumes do not excuse the record, they only postpone the analysis. Record invited, started and advanced by group from the first cohort, revisit as the numbers accumulate, and never publish a ratio the sample cannot support.
Should the candidate get a copy of their own assessment report?
Yes, and give them the same document rather than a summary of it. It answers the objection people actually hold: Pew found Americans oppose AI making final hiring decisions by 71% to 7% 6, and the fix for that is showing who read the session and what they wrote. Parity also disciplines the report, because nothing gets recorded that you would not hand over. The version you would be uncomfortable sending is the version that should not have been written.
References
- 1. 29 CFR § 1607.4 - Information on impact law.cornell.edu The four-fifths rule of thumb, and the requirement to keep records of impact by identifiable race, sex or ethnic group.
- 2. Employment Tests and Selection Procedures eeoc.gov Job-related and consistent with business necessity as the employer's burden, tied to the particular job in question, plus the less discriminatory alternative.
- 3. 29 CFR § 1607.5 - General standards for validity studies law.cornell.edu Users maintain validity documentation for a procedure with adverse impact, via criterion-related, content or construct validity.
- 4. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures eeoc.gov Q43 and Q66 on transportability of validity evidence between jobs; Q58 and Q77 on job analysis for content and construct validity.
- 5. Automated Employment Decision Tools (Updated) rules.cityofnewyork.us Local Law 144 rule: published bias audit summary and at least ten business days' notice to candidates before use.
- 6. AI in Hiring and Evaluating Workers: What Americans Think pewresearch.org Americans oppose AI making final hiring decisions 71% to 7%; 47% say AI would treat all applicants more similarly than humans.
6 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.