Policy
Will an AI-Skills Assessment Hold Up if a Candidate Challenges It?
An AI-skills assessment holds up when a rejected candidate challenges it if the file predates the first invite: a job analysis, a map from each task to a work behavior, findings on the challenger, comparator records, impact figures at your threshold, and the notice and alternative-process record. Title VII requires showing it is job related, consistent with business necessity [1]. Two limits: the candidate can still win by naming an equally valid, less-impactful alternative you refused, and content validity carries no trait score and nothing learned on the job [3].
The takeDefensibility gets sold as a feature of the tool, and it has never been one. A vendor can hand you an audit. The job analysis is yours to write, and no contract moves that. The employers I've watched come through a challenge intact did not build the file for the challenge. They built it because they wanted to know what they were buying, and the defense fell out of that. Where a challenge bites, I'd expect it to bite the number that decided on its own. A written finding at least has a story a person can tell about it.
Where Olive fits
Open a role and see what the work shows
Defensibility rests on the link between the task and the occupation, and on a record you can still produce a year later: Olive's cases are authored per occupation and carry that occupation's SOC code, a human reviewer writes all six findings with the timestamped excerpt each one rests on, and a released report exports as a self-contained file alongside archival JSON carrying its rubric, scorer and bank versions. Olive has not performed a bias audit and says so on its own evidence page: there is not yet enough volume for a four-fifths ratio to mean anything.
Rank your shortlistWhat Does 'Hold Up' Actually Mean?
It means surviving a three-step sequence, not passing an inspection. A complaining party first demonstrates that a particular employment practice caused a disparate impact. The burden then moves to you to demonstrate the practice is job related for the position in question and consistent with business necessity. If you carry that, the candidate can still win by naming an equally valid alternative with less impact that you refused to adopt 1.
Most challenges never reach a courtroom. What arrives is an agency charge, a demand letter, a city or state inquiry, or an internal complaint from someone who already read their report. All four are answered out of the same file, and that file either exists or it doesn't. A defense written after the challenge lands is a memo, not a record.
One clause decides how much of your process goes on trial. The complaining party normally has to point at the specific practice that caused the impact, but where the elements of your decision-making process are not capable of separation for analysis, the whole process may be analyzed as one practice 1. If nobody wrote down how the assessment fed the decision, you have handed over that separation.
The Uniform Guidelines put it more bluntly than the statute does: a selection procedure with adverse impact is considered discriminatory unless it has been validated 2. So the arithmetic on your own funnel matters long before anyone challenges anything, and so does knowing that a bias audit and a validation study answer different questions: an audit measures who came through at what rate, and validity evidence is what explains why the thing being measured belongs in the decision at all.
Which Evidence Actually Defends the Decision?
Six artifacts, and all of them have to exist before the first invite goes out. A job analysis naming the important work behaviors. The map from each assessment task to one of them. The reviewer's written findings for the candidate who challenged. The comparator records for everyone else in that requisition. Your impact figures at the threshold you actually ran. The notice and alternative-process record.
1. The job analysis. It names the critical or important work behaviors and the tasks associated with them, and where a behavior produces something, the work product too 3. Written by someone who talked to people doing the job. A competency framework bought off a shelf is not this. 2. The task-to-behavior map. One line per assessment task, naming the work behavior it samples and where that behavior appears in the analysis. This is the artifact counsel asks for first and the one most employers cannot produce. 3. The findings for this candidate. What the assessment showed, in writing, with the moment it rests on. A dimension name plus a number is not a finding; a finding is a described act that a second reader can locate in the record. 4. The comparator records. The same evidence for the other candidates in that requisition, at the same standard. A challenge is comparative by nature, and a file that is thorough for one rejected candidate and thin for the hired one reads exactly as badly as it sounds. 5. Impact figures at your threshold. Selection rates by group at the cutoff you ran, on your applicants, not the vendor's pool. Whether a rate difference is substantial is judged on the numbers you produced. 6. Notice, accommodation and the alternative process. Where New York City's rule applies, notice runs at least ten business days before use and must include instructions for requesting an alternative selection process or an accommodation 7. Even outside it, what you offer a candidate who cannot take the assessment as designed is part of the record.
The vendor holds a lot of this and owes you none of it by default. The EEOC's long-standing position on tests is that a vendor's validity documentation may help, but the employer is still responsible for ensuring its tests are valid 5; its 2023 technical assistance on algorithmic tools went further, saying an employer may be responsible even where the tool was developed by an outside vendor, and that if the vendor is wrong about its own assessment the employer can still be liable 6. That document no longer resolves on eeoc.gov, which is its own lesson: cite the regulation in your file, not an agency web page. The documents to demand from a vendor in writing belong in the contract, with dates.
Why a Generic AI Score Is the Weakest Artifact
Because it carries none of the evidence the middle step needs, and it attracts the rules that bite hardest. The Guidelines rule out by name any assumption of validity based on a procedure's name or descriptive labels, promotional literature, data on how widely it is used, and the testimonials and credentials of sellers or consultants 4. A number called an AI-readiness index is a descriptive label until a study says otherwise.
It also decides which regime you are under. New York City defines substantially assisting or replacing discretionary decision making as relying solely on a simplified output (a score, tag, classification or ranking) with no other factors considered; using that output as one of a set of criteria where it is weighted more than any other; or using it to overrule conclusions derived from other factors, including human decision-making 7. A number that decides is the trigger. A written finding a person weighed among other things is a different shape, and only your own records can show which one you ran.
Where the rule applies the duties are concrete, and they land on you rather than on the vendor: a bias audit within one year of use, a summary of results published on the employment section of your own site before use (including the selection or scoring rates and the impact ratios for every category), and candidate notice at least ten business days ahead 7.
A challenge goes after the evidence first. A finding that a candidate carried an unsourced claim into a client-facing memo at minute 34 is evidence about a work behavior, and it points at a moment two people can go and read. A 68 is a summary of evidence somebody else is holding. When counsel asks what the 68 measured, the answer has to come out of a validity study, and if no study names your occupation, there is nothing to say.
What to Do the Week a Challenge Arrives
Preserve first, reconstruct second. Once a charge is filed, every personnel record relevant to it is kept until final disposition, expressly including the test papers completed by the unsuccessful applicant and by every other candidate for the same position 8. Send the legal hold to the assessment vendor the same day, because a deletion default that fires mid-charge destroys your evidence, not theirs.
1. Freeze everything, vendor included. Exports, not dashboard views. A record you can only see while the subscription is live is not a record you can produce. 2. Pull the whole requisition. Every candidate, every stage, every note, including the informal ones. A challenge is answered comparatively or not at all. 3. Write the decision memo today, dated today. What the assessment showed, what else was weighed, who decided. Never backdate anything; a document with the wrong date on it converts a defensible process into a credibility problem. 4. Recompute the rates at the threshold you ran. On your applicants, for this requisition and for the last twelve months of the role. Do it before you are asked, and do it with someone who will report the number honestly. 5. Answer through counsel. Everything above is the file counsel works from. None of it is legal advice, and the moment a charge is real, the reply is not yours to draft alone.
If the file has holes, the useful move is forward rather than backward. Fix the open requisitions (write the job analysis, map the tasks, record the human decision) and treat the closed one honestly. Most people who challenge a rejection want an explanation they can act on, and a straight explanation of how the assessment fed the decision closes more of these than any citation does. The file is for the ones it doesn't close.
Common questions
Does the vendor's bias audit protect us if a candidate challenges the rejection?
No. An audit reports selection rates and impact ratios for a population; it says nothing about whether what the tool measures matters for your job, which is the step you have to carry. The EEOC's guidance on tests is explicit that a vendor's documentation may help, but the employer remains responsible for ensuring its tests are valid 5, and its 2023 technical assistance added that an employer can be liable even where the vendor developed the tool and was wrong about its own assessment 6. Ask which job families the audit covered and at which threshold.
Do we have to validate the assessment if it shows no adverse impact?
The validation obligation in the Uniform Guidelines attaches where a selection procedure has adverse impact 2, so a clean funnel lowers the immediate exposure. Two cautions. You cannot know your impact without keeping the records that measure it, and a ratio computed on forty applicants can move sharply at four hundred. Write the job analysis anyway. It costs days at the start and it is unbuildable in the week a charge arrives, because the artifact is a study of the job rather than a document about the tool.
Is an AI-skills assessment an automated employment decision tool under Local Law 144?
It depends on how the output is used, not on whether AI is involved. The city rule reaches a simplified output (a score, tag, classification or ranking) relied on solely, weighted more than any other criterion, or used to overrule conclusions drawn from other factors including human decision-making 7. A written finding that a person weighed alongside interviews and references is a different shape from a number that sets the cutoff. Your records decide which one you ran, so write down how the assessment fed the decision. Confirm the analysis with counsel.
What if the assessment tests something we would train a new hire on anyway?
Then it is the wrong thing to gate on. Content validity is not an appropriate strategy where the knowledge, skill or ability is one an employee will be expected to learn on the job 3. A test of your prompt library, your internal tools or a vendor's syntax fails that line, and it also screens out capable people for a gap that closes in week one. Move tool familiarity into onboarding and assess the judgment that does not: framing a problem, demanding a source, refusing a confident answer, checking it against something outside the conversation.
Does telling candidates we use AI in the assessment make it defensible?
No. Notice and job-relatedness are separate duties, and satisfying one does nothing for the other. Where New York City's rule applies, notice runs at least ten business days before use and carries instructions for requesting an alternative selection process or an accommodation 7. That is a disclosure obligation, not a defense. Disclosure is still worth doing everywhere: it sets expectations, it surfaces accommodation requests early, and a candidate who was told what was being assessed challenges less often than one who learned it from a rejection email.
References
- 1. 42 U.S.C. 2000e-2 - Unlawful employment practices, subsection (k) Burden of proof in disparate impact cases law.cornell.edu The burden-shifting sequence: the complaining party demonstrates disparate impact, the employer must demonstrate the practice is job related for the position in question and consistent with business necessity, and the refused less-impactful alternative; plus the clause treating a decision-making process as one practice where its elements are not capable of separation for analysis.
- 2. 29 CFR 1607.3 - Discrimination defined: Relationship between use of selection procedures and discrimination ecfr.gov A selection procedure with adverse impact is considered discriminatory unless it has been validated in accordance with the guidelines, and suitable alternative procedures with less adverse impact are part of the validity study.
- 3. 29 CFR 1607.14 - Technical standards for validity studies ecfr.gov Section 14C: content validity requires a representative sample of the job's work behaviors or work product; the manner, setting and complexity should closely approximate the work situation; a content strategy is not appropriate for traits or constructs such as intelligence, aptitude, personality, commonsense, judgment, leadership and spatial ability, nor where the knowledge, skill or ability is one an employee will be expected to learn on the job. Also the job-analysis requirement in 14C(2).
- 4. 29 CFR 1607.9 - No assumption of validity ecfr.gov Specifically ruled out as substitutes for evidence of validity: assumptions based on a procedure's name or descriptive labels, all forms of promotional literature, data on frequency of usage, and testimonial statements and credentials of sellers, users or consultants.
- 5. Employment Tests and Selection Procedures eeoc.gov A test vendor's validity documentation may be helpful, but the employer is still responsible for ensuring its tests are valid under the Uniform Guidelines; also the job-relatedness and business-necessity standard and the three validation strategies.
- 6. Select Issues: Assessing Adverse Impact in Software, Algorithms, and Artificial Intelligence Used in Employment Selection Procedures Under Title VII of the Civil Rights Act of 1964 web.archive.org Question 3: an employer may be responsible under Title VII even where the tool was developed by an outside vendor, and if the vendor is incorrect about its own assessment the employer could still be liable. Archived because the document no longer resolves on eeoc.gov as of 2026-08-24.
- 7. Notice of Adoption of Final Rule: Automated Employment Decision Tools (6 RCNY 5-300 to 5-304) rules.cityofnewyork.us The three-part definition of substantially assisting or replacing discretionary decision making via a simplified output; the bias audit within one year of use; the published summary of results with selection or scoring rates and impact ratios; and notice at least ten business days before use including instructions for requesting an alternative selection process or accommodation.
- 8. 29 CFR 1602.14 - Preservation of records made or kept ecfr.gov Where a charge has been filed, all personnel records relevant to it are preserved until final disposition, expressly including test papers completed by the unsuccessful applicant and by all other candidates for the same position.
8 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.