Assessment design
How Do You Evaluate an AI-Skills Assessment Before Buying?
Four questions settle whether an AI-skills assessment is worth buying, and the first decides most of it: what job analysis is it benchmarked against, for the occupation you're hiring for? What does the candidate see, and when? Does it output a single number, and what supports sorting by it? What sits behind a finding when two candidates diverge? Put all four in the RFP, and pilot on one requisition before the tool gates anyone. A vendor who cannot answer the first is selling a construct, not a work sample.
The takeThe evidence gap in this category is less a vendor problem than a demand problem. Composites exist because buyers want a shortlist sorted by lunchtime, and nobody builds a document nobody asks for. Ask four vendors for a dated job analysis and one of them starts writing one that week. On what is public so far, the category holds far more bias audits than validity studies, and that is because audits are the artifact procurement teams learned to request. Buyers set the evidence standard here. Most have not noticed they are holding the pen.
Where Olive fits
Open a role and see what the work shows
The first question turns on who wrote the answer key: Olive ships twelve authored cases per occupational bank, each bank grounded in one SOC code, and returns six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification) written by a human reviewer and anchored to a timestamped excerpt rather than blended into a composite. The candidate is granted the identical report.
Rank your shortlistWhat Is the Assessment Benchmarked Against?
Ask for the job analysis. Under the Uniform Guidelines, a work-sample claim rests on an analysis of the important work behaviors of the job in question, and the procedure has to be a representative sample of them 1. A vendor benchmarking against a general AI aptitude construct has no such analysis, because the construct is not an occupation. Get that answer in writing before the demo.
The Guidelines are unusually blunt about the difference. A content-validity strategy is not appropriate for a procedure that purports to measure a trait or construct, and the examples given name aptitude and judgment directly 1. What content validity does support is a sample: the manner, setting and level of complexity of the procedure should closely approximate the work situation 1. "AI aptitude" is the label that cannot be defended by pointing at the task; "this is the memo the job produces" is the one that can.
That gives you a test to run inside the demo. Ask which occupation the case was written for, who wrote it, and what in the role it was drawn from. A diligence memo over a source packet and a change to an unfamiliar codebase are different work situations, and no single case approximates both. An answer along the lines of "it works across every role" is the vendor telling you the benchmark is the construct.
Ask too where the validity evidence came from, because most of it came from someone else's applicants. Borrowing a study is permitted, but only where incumbents in your job and in the studied job perform substantially the same major work behaviors, shown by job analyses on both jobs 2. The liability does not travel with the study either: the Guidelines caution that users are responsible for compliance, whatever the publisher's manual says 2. Settle what the assessment is supposed to measure before you compare two of them on price.
What Does the Candidate See, and When?
Ask three things: what the candidate is told before the session starts, how much notice they get, and whether they ever see the result. Where New York City's rules apply, an employer must give notice 10 business days before an automated employment decision tool is used, and the bias audit summary has to be publicly available 5. A vendor whose invite flow cannot support that timing constrains your process, not theirs.
The notice duty attaches to the employer. So an assessment sent the same afternoon it is scheduled cannot satisfy it, however compliant the tool is. Ask for the invite mechanics in the RFP: how far ahead a link can be issued, how long it stays valid, whether an alternative process exists for a candidate who asks, and what disclosure the candidate reads before anything is captured.
Then ask what the candidate is shown at the end. A candidate-facing report is rare and is not required in most jurisdictions, which is exactly why it is worth pricing: a finding the candidate will read gets written differently from one only your hiring committee sees. Ask for a redacted sample of both documents, and read them side by side. If a sentence about a candidate would embarrass you in front of that candidate, it will eventually be read by one.
Accessibility belongs in this conversation rather than in a security review six weeks later. A timed browser task with audio, screen capture and a fixed window has several distinct accommodation surfaces, and "we extend the clock" covers one of them. Ask which the vendor has actually built, and how a request is made without disclosing a diagnosis to the hiring manager. The detail is worked through in accommodations on an AI-open assessment.
Does It Produce a Single Number?
Ask what comes out, then ask what supports using it that way. The Uniform Guidelines separate the two uses: evidence sufficient to support a procedure on a pass/fail basis may be insufficient to support the same procedure on a ranking basis, and a user choosing to rank should hold evidence supporting that use 3. A composite score is a ranking instrument whether or not the vendor describes it as one.
So the useful follow-up is not whether the number is accurate. It is what the number means and how it was assembled. A blended score has weights, and the weights are a claim about the job: that framing counts twice as much as verification, say, for this occupation. Ask what the weights are, who set them, and against what evidence. "The model learned them" means they were learned from someone's past decisions, which is a model of the old screen rather than of the work.
Cutoff scores carry their own standard. Where they are used, the Guidelines say they should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the workforce 3. A vendor-supplied default cutoff was not set against your workforce, and adopting it makes it yours. Ask what happens to a candidate one point below it, and whether anyone reads the underlying work.
The alternative is a report that stays disaggregated: separate findings on separate behaviors, each readable on its own, none of them added together. That output is slower to sort and harder to put in a spreadsheet, which is the trade you are making. It also survives the question a single number cannot answer, which is what specifically this person did. Whether the number predicts anything at all is a separate claim with its own evidence.
How Is a Divergent Result Evidenced?
Ask what sits behind a finding when two candidates come out differently. The documentation standard is specific: the user should identify the work behavior each part of the procedure is intended to sample, and reliability estimates should be reported for the procedure where available 4. If a vendor cannot say what an item measures, nobody in your company can defend the gap between two candidates, not to the candidates and not to counsel.
Ask for the technical report by name rather than the one-page summary. For a content-validity claim the Guidelines mark particular items essential: the method used to analyze the job, a complete description of the work behaviors and how their importance was determined, an explicit description of what the procedure samples, and a comparison of the manner, setting and complexity of the procedure against the work situation 4. Those are the pages anyone reviewing the decision later will actually want.
Then ask about agreement. Two reviewers reading the same session, or the same scorer run twice, should agree at a rate the vendor can state as a number with a method behind it. "Our rubric is standardized" is a description of intent, not a measurement. If nobody has measured it, that is a real answer. Write it down, and treat the assessment as an input rather than a gate until it exists. What rubric reliability takes to establish is the same work whether you buy it or build it.
A "bias tested" badge is a different document answering a different question, and vendors routinely offer the first when asked for the second. An audit compares selection or scoring rates across groups; a validity study argues the procedure relates to the job. Ask for both, ask which jobs each one covered, and check the audit date against the last model change 5. The distinction is worth settling before the procurement call.
Run the Checklist Against One Real Req
Take the four questions to a live requisition rather than to a category review. Pick the role you are hiring for now, ask the vendor to name the case its bank would open on, and read that case against the job you posted. A vendor selling a work sample can do this inside a week. One selling a construct will offer a general demo and a customization roadmap instead.
Put the questions in the RFP, not the demo. A demo is answered by a salesperson in the room; an RFP is answered by someone whose answer gets filed. Four documents, each with a date on it: the job analysis, the technical report, the bias audit with its scope note, and a redacted sample of both the employer report and the candidate report.
Then run it beside your existing round before it replaces anything. A pilot on one requisition, where both instruments see the same candidates, tells you what the assessment adds and what it duplicates, and it produces the record you will want if anyone later asks how the tool was chosen. Scope the pilot before you sign, because a vendor who will not pilot on your req has told you something about the evidence. The mechanics are in running a pilot before the assessment becomes a gate.
Buying is not the only path, and the comparison is closer than the pricing suggests. The expensive parts of an assessment are the authored case, the answer key and the reviewer time, not the software that serves it, which is why the build-or-buy question turns on authoring capacity rather than on engineering.
Common questions
What should a vendor's validity evidence actually contain?
For a content-validity claim, the Uniform Guidelines mark specific items essential: dates and locations of the job analysis, the method used to analyze the job, a complete description of the work behaviors and their measures of importance, an explicit description of what the procedure samples, and a comparison of the procedure's manner, setting and complexity against the work situation 4. Reliability estimates should be reported where available 4. Ask for that document by name. A marketing one-pager citing a correlation with no job analysis behind it is not validity evidence, whatever number it carries.
Is a bias audit the same as validity evidence?
Not the same document at all. A bias audit compares selection or scoring rates across demographic groups; a validity study argues the procedure relates to the job. New York City requires the audit within one year of the tool's use and a publicly available summary 5, and none of that speaks to whether the tool measures anything relevant to your occupation. A vendor can hold a clean audit and no validity evidence for your role at all. Ask for both documents separately, and check which jobs each one covered.
Does the candidate have to see their result?
In most jurisdictions, no. New York City requires notice before use and a published bias audit summary, not disclosure of an individual result 5. Treat a candidate-facing report as a procurement question rather than a compliance one: it constrains how findings are written, gives candidates something for the hour you asked of them, and removes the gap between what you say about someone internally and what you would say to them. Ask for a redacted sample of both documents before you decide it is optional.
What if the vendor won't share the technical report?
Treat the refusal as the answer and put it in the procurement record. Vendors decline for two reasons: the evidence does not exist in the form the deck claims, or it exists and says less than the deck. Both are decision-grade. Responsibility for compliance sits with the employer using the procedure, not with the publisher that sold it 2, so "proprietary" transfers no risk away from you. Price the tool as unevidenced for your roles, or ask for a validity study on a pilot cohort of your own applicants.
Do these questions still matter if the score is only one input?
Yes, and the Guidelines key off exactly that. The standard scales with how a procedure is used: evidence adequate to support pass/fail use may be inadequate to support ranking 3. "One input among several" describes an intention, not a mechanism. If the number orders your shortlist, determines who gets a second conversation, or breaks a tie between two finalists, it is being used to sort people, and the evidence has to support that use.
References
- 1. 29 CFR 1607.14 - Technical standards for validity studies ecfr.gov Content validity requires a job analysis of important work behaviors; a content strategy is not appropriate for procedures purporting to measure traits or constructs such as aptitude and judgment; the manner, setting and complexity of the procedure should closely approximate the work situation.
- 2. 29 CFR 1607.7 - Use of other validity studies ecfr.gov Borrowing a validity study requires job similarity shown by job analyses on both jobs; users remain responsible for compliance regardless of what a publisher's manual claims.
- 3. 29 CFR 1607.5 - General standards for validity studies ecfr.gov Section 5G: evidence sufficient for pass/fail use may be insufficient for use on a ranking basis. Section 5H: cutoff scores should normally be reasonable and consistent with normal expectations of acceptable proficiency within the workforce.
- 4. 29 CFR 1607.15 - Documentation of impact and validity evidence ecfr.gov Section 15C marks as essential the job-analysis method, the described work behaviors and their importance measures, an explicit description of what the procedure samples, identification of the behavior each item samples, and a comparison of manner, setting and complexity against the work situation; reliability estimates reported where available.
- 5. Automated Employment Decision Tools (Local Law 144 of 2021) nyc.gov A bias audit within one year of use, a publicly available summary of its results, and notice to candidates 10 business days before the tool is used.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.