Assessment design

What Should an AI Skills Assessment Measure Before You Pay?

An AI skills assessment worth paying for measures behavior on your own occupational work: how a candidate frames the problem, which claims they demand a source for, what they keep by hand, and what they check outside the model. Tool familiarity and prompt trivia take a week to learn and score every job the same. Demand the job analysis behind the task candidates are given, a finding on each behavior with the excerpt it rests on, occupational data dated this year, and the same report sent to the candidate.

The takeOf everything on this list, the question I would ask first is the free one: send me the document the candidate receives. A vendor whose findings survive being read by the person they describe has already done the expensive part, because nobody writes a defensible sentence about someone's afternoon without the excerpt sitting under it. Refusing that document is rarely a privacy position. Nobody has measured the correlation, and it still looks like the cheapest tell in the room: the products that will not show a candidate the report are the ones with nothing underneath the findings.

Where Olive fits

Open a role and see what the work shows

Run this checklist against Olive too: twelve authored cases per occupation, and six findings a human reviewer writes by hand (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each carrying the timestamped excerpt it rests on, with the candidate granted the identical report.

Rank your shortlist

What Should an AI Skills Assessment Actually Measure?

Behavior on your own occupational work, broken into dimensions that can be judged separately, each one carrying the moment it rests on. The useful signal is what a candidate does when an assistant is willing to do everything: whether the problem gets framed before anything is generated, which claims get a source demanded, what work is kept by hand, and what gets tested against something outside the chat.

Name the dimensions before you look at a vendor, because the names decide what the report can tell you. A single blended figure hides which part of the work was strong, so it gives a hiring manager nothing to act on and gives a candidate nothing to answer. Separate judgments do both.

The measurement itself is worth four questions:

  • Is the task occupational? A generic "summarize this article" prompt tests nobody's job. The case should look like a Tuesday in the role you are hiring for, with the ambiguity left in.
  • Is the assistant real, and willing to overreach? If the model in the exercise is capped, sanitized, or scripted, the exercise measures compliance with a script.
  • Is each dimension reported on its own? Six findings that disagree with each other are more informative than one number they were averaged into.
  • Is anything about style being scored? Prose quality, prompt phrasing, typing speed and how much AI got used are not signals. Volume of usage is the easiest thing to measure and the least worth knowing.

That last point is where most demos quietly fail. If you are rebuilding the exercise yourself rather than buying, the same four questions govern what a work sample should test now that AI can produce the work sample.

Why Tool Familiarity Is the Wrong Thing to Buy

Because familiarity generalizes and judgment does not. A test of which model to pick, what a token is, or how to phrase a prompt measures something a new hire acquires in a week, and it returns the same result for a paralegal and a data scientist. The Uniform Guidelines are blunt about the consequence: a selection procedure rests on content validity only to the extent it is a representative sample of the content of the job 2.

The same section rules out two things vendors sell. Content validity cannot support a procedure that measures traits or constructs (the Guidelines name intelligence, aptitude, personality, commonsense, judgment, leadership and spatial ability), and it cannot support one built on knowledge or skills an employee will be expected to learn on the job 2. So a product sold as an "AI aptitude" measure needs criterion-related evidence instead: scores that relate to some measure of performance on the job. Ask which of the two arguments the vendor is making. Most decks make neither.

Notice that judgment is on that list of constructs. That is the reason to buy a sample of the work rather than a judgment test: what can be defended is what a person did on an occupational task, not an abstract faculty inferred from twenty questions. It is also why the bar has to be written per role before the assessment arrives, which is a separate exercise: setting a defensible proficiency bar for a specific role is work the vendor cannot do for you.

What Evidence Should Come Back With Each Finding?

Something you can open. Every dimension should arrive with the moment it rests on: a timestamp in the recorded session, a turn in the AI transcript, a diff in the artifact, an answer in the written debrief. A finding with no excerpt behind it is an opinion with a label on it, and it is worth nothing the first time a hiring manager or a candidate asks why.

The vendor's own documentation is the other half. The Guidelines set out what a content validity report contains: the job analysis and the work behaviors it found, a description of the procedure and its content, the relationship between the two, the alternative procedures investigated, and who did the work 3. Ask for that document. It is a different artifact from the accuracy figure on the slide, and a vendor who has one will send it.

Do not accept the claim in place of the evidence, and do not assume the claim transfers to you. The EEOC's own questions and answers on the Guidelines say validity reports printed in test manuals may help, but it is the user's responsibility to determine that the evidence is adequate, and that an unsupported assertion by anyone that a procedure has been validated does not satisfy the Guidelines 1. Responsibility stays with the employer even where an agency or a consultant administers the procedure 1. If the vendor's answer is a compliance badge, note that a bias audit and a validation study answer different questions.

How Recent Does the Benchmark Have to Be?

Recent enough that the tasks in it are the tasks the job had this quarter. Occupational content moves, and the reference data moves with it: O*NET updates its database quarterly, with a primary update each third quarter, across a taxonomy of just over a thousand occupations 5. A case authored against a two-year-old picture of an occupation is testing a job that has changed shape since.

So ask for version strings, not adjectives. Which release of the occupational data the role content was built from. When the case was last re-authored, and what changed. Which rubric version produced the report you are holding. "Continuously updated" is not a date, and a vendor who cannot name one is describing a benchmark rather than showing you it.

Some products benchmark a candidate against the vendor's other candidates rather than against the job, which is the quieter substitution of the two. That is a statement about who else bought the product, and it drifts every time the customer mix does. Job-anchored content survives a change in the customer base; a cohort percentile does not. The same distinction decides whether one assessment can cover every department or the case has to change per role.

Does the Candidate See the Report?

Ask, and treat the answer as a purchase criterion. A candidate who gets the same document you get can correct a factual error in it, and a report you would not want them to read is a report that should not reach a hiring manager either. Vendors who withhold it usually cannot show you the excerpts behind the findings, which is the same defect wearing a policy.

Administration is the other half of that question. Under the ADA regulations, a test given to an applicant whose disability impairs sensory, manual or speaking skills has to be given in a way that makes the results reflect what the test is meant to measure rather than the impairment 4. In practice that is one concrete question to the vendor: is there a typed path everywhere there is a spoken one, built into the product rather than arranged by email afterwards? The difference matters enough to check in the demo, because accommodations in an AI assessment are a design property, not a support ticket.

Then the disclosure duties, jurisdiction by jurisdiction. In New York City, a tool that substantially assists or replaces discretionary hiring decisions triggers a published summary of an independent bias audit, at least ten business days' notice to the candidate, and published information on the data the tool collects and how long it is kept 6. Those duties attach to the employer, not to the vendor who offers to handle them. Ask which rules the vendor believes apply to their tool, get the answer in writing, then take it to your own counsel.

See what gets scored

Common questions

Is a multiple-choice AI literacy test worth buying?

Only as a cheap filter for something you would otherwise ask in a form field. It measures recall of tool facts, which a new hire picks up in days, and it returns similar scores for jobs with nothing in common. Where a decision turns on judgment, the Uniform Guidelines push toward evidence tied to job content: a procedure is supportable on content validity only to the extent it samples the content of the job 2. A quiz samples the content of a quiz.

Can one assessment cover every role you hire for?

The dimensions can be shared; the task cannot. Framing, evidence, delegation and verification describe good AI work in any occupation, so one rubric across departments is reasonable. The case is where the job lives, and content validity is an argument about the specific job's work behaviors 2. A finance case run on an engineering req produces a report about how someone handles a finance case.

Does a passing bias audit mean the assessment is validated?

A passing audit is not validation. They answer different questions. An audit compares selection or scoring rates across demographic groups. Validation argues the scores relate to the job, and the Guidelines require validity evidence where a selection process has adverse impact, with documentation kept by job 3. A vendor can hold a clean audit and no job-related validity evidence for your occupation at all. Ask for both, separately, and do not let one stand in for the other.

Who is responsible if the vendor's validity evidence turns out to be thin?

You are. The EEOC's questions and answers on the Uniform Guidelines put it on the user to determine that validity evidence is adequate, and state that an unsupported assertion by anyone that a procedure has been validated does not satisfy the Guidelines 1. That responsibility stays with the employer even where an agency or consultant administers the procedure 1. Practically: keep the vendor's documentation where your counsel can find it, and file any refusal to send it.

What should a pilot look like before you sign?

One req you are already hiring for, run beside your existing round rather than in place of it, with nobody gated on the result. Read the reports against what the interviews said and what the hire actually did. Ask to see the exact document a candidate would receive. Then price the tool on whether the findings told you something the round had missed, not on whether the demo was impressive.

References

  1. 1. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures U.S. Equal Employment Opportunity Commission, 1979. eeoc.gov Q35: validity reports in test manuals may help, but it is the user's responsibility to determine the evidence is adequate. Q38: an unsupported assertion by anyone that a procedure has been validated does not satisfy the Guidelines. Q63: responsibility stays with the employer where a third party administers the procedure.
  2. 2. 29 CFR § 1607.14 - Technical standards for validity studies Uniform Guidelines on Employee Selection Procedures, Legal Information Institute, Cornell Law School, 1978. law.cornell.edu Section C: content validity requires a job analysis of important work behaviors; a procedure is supportable to the extent it is a representative sample of the content of the job; content validity is not appropriate for traits or constructs (intelligence, aptitude, personality, commonsense, judgment, leadership, spatial ability) or for knowledge and skills learned on the job.
  3. 3. 29 CFR § 1607.15 - Documentation of impact and validity evidence Uniform Guidelines on Employee Selection Procedures, Legal Information Institute, Cornell Law School, 1978. law.cornell.edu Users maintain adverse impact information per job and, where impact is found, evidence of validity; the content validity report must document the job analysis, work behaviors, the procedure and its content, the relationship between them, alternatives investigated, and the researcher.
  4. 4. 29 CFR § 1630.11 - Administration of tests ADA Title I regulations, Legal Information Institute, Cornell Law School, 1991. law.cornell.edu A test given to an applicant with a disability impairing sensory, manual or speaking skills must be selected and administered so results reflect the factor the test purports to measure rather than the impairment.
  5. 5. O*NET Database National Center for O*NET Development, 2026. onetcenter.org The O*NET Data Collection Program updates the database quarterly, with a primary update in the third quarter of each year, across a taxonomy of just over a thousand occupations.
  6. 6. Notice of Adoption of Final Rule: Automated Employment Decision Tools (6 RCNY §§ 5-300 to 5-304) NYC Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us Local Law 144 duties for a tool substantially assisting or replacing discretionary employment decisions: published summary of an independent bias audit, at least ten business days' notice to candidates, and published information on data collected and retention.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.