Assessment design
Six Categories of Hiring Assessment and What Each Can Claim
Six categories cover almost everything sold as pre-employment assessment: cognitive ability, personality, situational judgment, job knowledge, work samples, and structured interviews. They differ less in price than in what they license you to conclude. Pick backwards. Write the sentence you would have to defend to the candidate you rejected, then choose the category that supports that sentence. The first three on that list do not support "this person can do the job" on their own.
The takeCategory shopping is the wrong first move. A vendor comparison grid sorts on features and price because that is what a grid can sort on, and the axis it cannot show is the one that matters: what a defensible rejection would be allowed to say. Half the market sells inference about a person while the buyer hears evidence about the work. Those are different products, and the price gap between them is smaller than the evidence gap.
Where Olive fits
Open a role and see what the work shows
Olive returns six separately evidenced findings from one occupational assignment, each anchored to a timestamped moment in the session rather than to a number. Twelve item banks are live, each grounded in a single occupation, and the candidate is granted the identical report.
Rank your shortlistStart from the sentence you would have to defend
Before comparing products, write the one sentence you would say to the strongest candidate you are about to reject. "You scored below the cutoff on a reasoning test" is a different sentence from "your draft missed the pricing error in the brief." The second names job content. The first names a trait. Only one of them survives the question "how is that related to this job?"
That question is older than every product in the market. Griggs v. Duke Power (1971) holds that a hiring test must measure the person for the job and not the person in the abstract, and that the touchstone is business necessity 1. Nothing in it forbids testing. What it forbids is giving a test controlling force without showing it is a reasonable measure of job performance, and that showing is far easier for some categories than others.
The practical version for a buyer: for each category on the shortlist, finish the sentence a declined candidate would hear. Declined because ___. If the blank fills with something the person did on work resembling the job, the category is producing evidence. If it fills with a percentile, a trait label or a fit number, the category is producing inference, and the inference is what gets argued about afterwards.
Do this before the demo rather than after it. A demo is built to make an output look like a conclusion, and every category in this market produces an output. The sentence test is the cheapest way to find out which outputs a person can stand behind, and it settles most of what an assessment should be measuring before a single vendor call.
What can each of the six categories actually claim?
Each category answers a narrower question than its marketing implies. Cognitive ability tests estimate general reasoning. Personality inventories estimate stable dispositions. Situational judgment tests estimate knowledge of what a good response looks like. Job knowledge tests estimate what someone has already learned. Work samples show a person producing something. Structured interviews record judgment on job scenarios against a fixed scale.
The 2022 re-analysis of the selection literature gives most of them a corrected number and, in the same table, a mean Black-White subgroup difference. Structured interviews estimate at .42 with a d of .23. Job knowledge tests sit at .40 with a d of .54. Work samples come in at .33 with a d of .67, and cognitive ability tests at .31 with a d of .79 2.
Read the pairing rather than either column alone. On these four the two columns run the same way: the strongest predictor carries the smallest subgroup difference, and the weakest carries the largest. That alignment is local to this shortlist and breaks further down the table, so read the second column for whatever you are actually considering rather than assuming a tradeoff has been forced on you.
Both columns carry limits worth stating to whoever asks for the slide. The validities are corrected correlations with supervisor ratings of job performance, pooled across many jobs, not accuracy rates and not a promise about any one process. The subgroup differences are borrowed from other meta-analyses, several drawn from non-applicant samples, and a subgroup difference is not adverse impact: impact depends on how a score is used, the selection ratio, and who applied 2. The corrected table is worth reading in full before a purchase, and the correction itself is more recent than most of the decks quoting it.
Which categories changed when an assistant could complete them?
The ones delivered as an unproctored artifact with a known right answer. A remote job knowledge test, an untimed situational judgment inventory and most take-home work samples now measure something different from what their validity evidence was collected on, because the thing being scored can be produced without the candidate. The category label on the invoice did not change. The claim it supports did.
Look at where the numbers came from. The .33 for work samples rests on 54 studies in which 53 tested people already doing the job 2. Those participants were not applicants, had no incentive to outsource the task, and had no assistant to outsource it to. The estimate still holds for the setting it came from, and that setting is not an unsupervised take-home.
Two things did not change. A live conversation where every candidate is asked the same scripted questions and pushed with follow-ups still measures the candidate, because the follow-up is unrehearsed. And an exercise where an assistant is explicitly allowed, and the working rather than the artifact is what gets read, gets more informative when the assistant is present.
The category that changed most is the one nobody lists: the unscored screening chat. Federal selection law never treated it as outside the rules. The Uniform Guidelines define a selection procedure to cover the full range of assessment techniques through informal or casual interviews and unscored application forms 3. Swapping a scored test for a conversation does not move the step outside the Guidelines, and it removes the record that would have defended it. When the shortlist has come down to three instruments, that comparison is worked through here.
How do you pick one for this role?
Two moves. Name the two or three job-content decisions the role gets wrong most expensively, then pick the cheapest category that puts a candidate in front of one of them. Everything else is a tiebreaker, and tiebreakers rarely need to be bought. A role that fails on judgment under ambiguity is not diagnosed by a reasoning test, however well that test predicts on average.
The base rate is worth knowing before the meeting. In SHRM's benchmarking survey of member organisations, 34% used structured interviews for executive roles, 37% for middle management and 36% for individual contributors, while work sample interviews ran at 12%, 11% and 9% of organisations 4. Self-reported, collected in 2021, and organisations claim more structure than they run, so treat those as an upper bound. Even as an upper bound, the job-shaped categories are a minority practice.
Do not read this as a technical-hiring question. A warehouse operations role, a claims desk and a clinic front office each have two or three decisions worth putting in front of a candidate, and the category map does not change because the work is not knowledge work. What changes is how easy the work sample is to build, and the answer is usually easier than the vendor implies.
For a first hire out of school, where none of the categories have much history to work with, the early-career version of this choice is a different problem with a different shortlist. And if a personality inventory has reached the final two, read what that category can support before it goes in front of anyone.
Common questions
Is a situational judgment test a work sample?
No. A situational judgment test asks what a person would do; a work sample watches what they do. The first measures knowledge of a good answer, which is learnable from a study guide and is not the same thing as producing the answer under real conditions. The second produces an artifact that a reviewer can read and argue with. They are often priced similarly and sold from the same catalogue, which is why the distinction has to be made by the buyer rather than the seller.
Does a validation study have to come before using one of these?
Not automatically. Under the 1978 federal Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607), the validation burden attaches where a selection procedure shows adverse impact, and the agencies' own guidance says that with no adverse impact there is no validation requirement. That is a narrower rule than "everything must be validated," and it is also thinner protection than it sounds: adverse impact is measured after the fact, on a pool you have already run the test on. Job-relatedness is worth being able to describe on day one, whether or not anyone asks, and where your own process sits is a question for counsel.
Which category is cheapest to build in-house?
A structured interview, by a wide margin. It costs question writing, a rating scale, and an hour of interviewer training, with no licence and no vendor. It is also the top-ranked single method in the corrected 2022 estimates, which makes it the rare case where the cheapest option to build is also the best evidenced. What it does cost is discipline: the same questions in the same order for every candidate, scored on a common scale, or it stops being the thing the evidence describes.
What about assessment centres and simulations?
Assessment centres and simulations are composites rather than a seventh category, usually a work sample and a structured interview run together over a day with observers. A day of one does not out-predict the structured interview inside it, which stays top ranked in the corrected 2022 estimates 2, and almost nobody runs them anyway: the 2021 SHRM benchmarking data puts assessment centre use in the low single digits of organisations 4. A compressed version, one exercise plus a scored debrief, is the part worth keeping.
Can one assessment cover several roles?
A cognitive or personality instrument travels across roles by design, which is exactly why it says little about any one of them. A job-knowledge or work-sample instrument does not travel, and that is the source of its value. The practical middle is one instrument per occupation rather than one per job title: the decisions a financial analyst faces are shared across analyst roles, and the exercise can be too.
Where does an AI-skills assessment fit in this map?
An AI-skills assessment is a work sample with a specific content choice, and it sits in that column of the map. It puts job content in front of the candidate and reads what they produced, so it inherits the work-sample column's strengths and its limits. What makes it a distinct purchase is the content: the decisions being observed are delegation, verification and rejection of a confident wrong answer, which older work samples had no reason to include.
References
- 1. Griggs v. Duke Power Co., 401 U.S. 424 (1971) law.cornell.edu Supports the claim that a hiring test must be tied to the specific job rather than measuring the person in the abstract.
- 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the corrected validity estimates per category, the paired Black-White subgroup differences, and the concurrent-sample base of the work sample figure.
- 3. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) govinfo.gov Supports the claim that an informal or unscored interview is a selection procedure under the same rules as a scored test.
- 4. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) shrm.org Supports the adoption rates showing that structured interviews and work samples remain minority practice.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.