Policy

How Do You Defend an AI-Skills Assessment to Legal and to Candidates?

Defend an AI-skills assessment with documents, not assurances. Legal needs four: a job-relatedness file showing the assessment samples what the occupation performs, a record of who advanced by race, sex and ethnic group, a named human who makes the decision, and the notice candidates got before sitting it. Candidates need one more, their own copy of the result. Rebuild the file for every occupation, because validity evidence does not transport between jobs, and send notice before the tool runs, ten business days ahead where New York City law reaches you.

The takeThe 1978 Guidelines were written for paper tests, and they still ask the only question that matters here: what this particular job requires. Nothing about a model changes that burden. What changes is when anyone writes it down. On what is public so far, these disputes come down to whether the folder existed before the dispute did. Evidence assembled after a complaint reads as an argument to everyone in the room, however solid it is on the page. So the strength of a defense here is mostly fixed months before anyone thinks to test it.

Where Olive fits

Open a role and see what the work shows

Legal asks who decided and on what evidence; the candidate asks the same question in different words. Olive answers both with one artifact: six findings a human reviewer wrote, each anchored to a timestamped moment in the session, exported with its rubric, scorer and bank versions, and granted to the candidate in identical form.

Rank your shortlist

Why job-relatedness has to be rebuilt for every role

Because validity evidence does not transport. The Uniform Guidelines Q&A says a procedure found valid in one situation does not necessarily have validity in different circumstances, and sets four conditions before evidence can be borrowed from another job 4. So the file you assembled for software engineering does not defend the same assessment used on underwriters. Each occupation gets its own task analysis and its own record.

The Q&A is also specific about job analysis: a full job analysis is required for all content and construct validity studies, and the analysis has to describe the important work behaviors, their relative importance and their difficulty 4. An AI-skills assessment is a work sample, so content validity is the natural strategy, which means the job analysis is not paperwork around the defense. It is the defense.

In practice that means the assessment's material has to be the occupation's material. For an underwriter, an application and a survey that do not describe the same building. For a financial analyst, a source packet whose claims only settle when someone opens them. The behaviors under assessment can stay constant across roles; the task, the answer key and the evidence file cannot. One assessment stretched across every department is the hardest version of this to defend.

Borrowing evidence between two job families is allowed and it is narrow. The Q&A's conditions include showing the jobs comparable on a work-behavior basis and accounting for variables such as performance standards, work methods and how representative the study sample was 4. If you cannot write that comparison down, you have two families and two files.

How do you write the evidence file for one role?

Build it as one folder per role, assembled before the first invite goes out. It holds the task analysis for the occupation, the assessment content mapped task by task, the rubric and its version, the names of the people who review, the decision rule, and the running record of who was invited and who advanced by group. Write it while the reasoning is still cheap to recall.

What goes in, concretely:

  • The task analysis. The work behaviors this role performs, their importance and their difficulty, sourced from people doing the job rather than from the job posting.
  • The content map. Each part of the assessment against the behavior it samples. A row with no behavior beside it is a row to cut.
  • The rubric, versioned. What counts as demonstrated, what does not, and which version was in force on the date each candidate sat it.
  • The reviewer record. Who read the session, what they wrote, and the calibration they went through before reading anything that counted.
  • The decision rule. What the hiring team does with the result, written before the results exist. "Considered alongside the interview" is a decision rule; so is a cut score, and a cut score needs a source.
  • The impact log. Invited, started, advanced, by group, from the first cohort.

The rule of thumb: anything that would sound like a judgment call in a deposition should be a dated document instead. That is also the practical difference between a defensible AI-skills assessment and one that merely produces a defensible-sounding report: the file exists before the dispute, or it does not exist.

If you bought the assessment rather than built it, the same folder still belongs to you. Ask the vendor for the job analysis behind each bank, the rubric version history, and the reliability evidence between reviewers. Then get it in writing before the pilot, because a sales deck is not documentation.

What do you tell candidates, and when?

Before they start, in plain language: what the task is, what is captured, which behaviors are assessed, who reads the result, and how long it is kept. After it is reviewed, give them the result itself. New York City requires at least ten business days' notice before an automated employment decision tool is used, plus a published bias audit summary 5. Treat that as the floor everywhere, not as a New York obligation.

The objection you are answering is specific, and it is not an objection to being assessed. Pew Research Center found Americans oppose the use of AI in making final hiring decisions by 71% to 7%, while 47% said AI would do better than humans at treating all applicants similarly 6. People are not rejecting structure. They are rejecting being decided by a machine with nobody in the room.

So the sentence you owe a candidate is short: a person read your session, wrote what they saw, and here is what they wrote. Report parity is the cheapest defensible thing on this list: the identical document, with the identical excerpts, goes to the candidate. It also disciplines the writing, because nothing gets recorded about a candidate that you would not hand to them. Decide what a candidate-facing report exposes before the first release, not after someone forwards one to a lawyer.

Say how to request an accommodation, and say it before the assessment rather than inside the rejection. A timed task, a microphone, a screen-capture requirement and a fixed session length are each a place where the process can turn a disability into a score. Offer the alternate path up front and record that it was offered.

A finding carries the moment it rests on, so it can be discussed, disputed and corrected. A composite number standing for the person carries none of that. It never goes in the report, and it is the single artifact most likely to be read back to you.

Common questions

Does an AI-skills assessment need a bias audit before we use it?

It depends on where you hire and how the result is used. New York City requires a published bias audit summary and at least ten business days' notice to candidates before an automated employment decision tool is used 5. Federally, the Uniform Guidelines do not require an audit by that name, but they do ask every user to keep records of the impact its selection procedures have by race, sex and ethnic group, and treat a selection rate under four-fifths of the highest group's as evidence of adverse impact 1. Keep the log from the first cohort either way.

Can we use the same assessment for every role?

You can use the same instrument; you cannot reuse the same defense. The Uniform Guidelines Q&A says validity found in one situation does not necessarily hold in different circumstances, and lists four conditions before evidence is borrowed across jobs, including showing the jobs comparable on a work-behavior basis 4. In practice: one task analysis, one content map and one impact log per occupation. The behaviors you assess can stay constant across roles. The occupational material and the evidence behind it cannot.

What do you have to tell candidates before an AI-skills assessment?

What the task is, what is captured, what is assessed, who reads it, how long it is kept, and how to request an accommodation. In New York City that notice has a deadline attached: at least ten business days before the tool is used 5. Everywhere else it is still the right default, because disclosure before the session is the only version a candidate can act on. Send it before any account step, so nobody has to sign up to find out what they agreed to.

Does a human reviewing the result remove the legal risk?

It changes what you have to defend, and only if the review is real. A named person who reads the evidence, writes the reasoning and can be asked why is a decision-maker; a person who clicks through a recommendation is a rubber stamp with a name on it. Either way the underlying procedure still has to be job-related and consistent with business necessity where it produces disparate impact 2. Human review is a control on how the result is used, not a substitute for evidence that it measures the job.

What if there are too few candidates for a four-fifths analysis?

Log the data anyway and say plainly that the sample is too small to interpret. The four-fifths ratio is a rule of thumb federal enforcement agencies generally treat as evidence of adverse impact 1, and a ratio computed on eight candidates is noise wearing a number. Small volumes do not excuse the record, they only postpone the analysis. Record invited, started and advanced by group from the first cohort, revisit as the numbers accumulate, and never publish a ratio the sample cannot support.

Should the candidate get a copy of their own assessment report?

Yes, and give them the same document rather than a summary of it. It answers the objection people actually hold: Pew found Americans oppose AI making final hiring decisions by 71% to 7% 6, and the fix for that is showing who read the session and what they wrote. Parity also disciplines the report, because nothing gets recorded that you would not hand over. The version you would be uncomfortable sending is the version that should not have been written.

References

  1. 1. 29 CFR § 1607.4 - Information on impact Uniform Guidelines on Employee Selection Procedures, Legal Information Institute, Cornell Law School, 1978. law.cornell.edu The four-fifths rule of thumb, and the requirement to keep records of impact by identifiable race, sex or ethnic group.
  2. 2. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Job-related and consistent with business necessity as the employer's burden, tied to the particular job in question, plus the less discriminatory alternative.
  3. 3. 29 CFR § 1607.5 - General standards for validity studies Uniform Guidelines on Employee Selection Procedures, Legal Information Institute, Cornell Law School, 1978. law.cornell.edu Users maintain validity documentation for a procedure with adverse impact, via criterion-related, content or construct validity.
  4. 4. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures U.S. Equal Employment Opportunity Commission, 1979. eeoc.gov Q43 and Q66 on transportability of validity evidence between jobs; Q58 and Q77 on job analysis for content and construct validity.
  5. 5. Automated Employment Decision Tools (Updated) NYC Rules, Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us Local Law 144 rule: published bias audit summary and at least ten business days' notice to candidates before use.
  6. 6. AI in Hiring and Evaluating Workers: What Americans Think Pew Research Center, 2023. pewresearch.org Americans oppose AI making final hiring decisions 71% to 7%; 47% say AI would treat all applicants more similarly than humans.

6 sources, numbered by first appearance. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.