Assessment design
The Validity Table Everyone Quotes Was Corrected in 2022
Of the standard hiring methods, structured interviews predict job performance best, at a corrected validity of .42. The table most vendor decks still reproduce comes from 1998; the 2022 re-analysis found that prior meta-analyses had systematically overcorrected for range restriction, cutting mean estimates by .10 to .20 points. Job knowledge tests now land at .40, work samples at .33, cognitive ability at .31, and an unstructured conversation at .19. The spread is narrow, so combining two job-related methods beats hunting for the single best one.
The takeA deck that still shows work samples at .54 and cognitive ability at .51 is quoting a 1998 summary and has not read the 2022 correction that replaced it. That is worth saying out loud in the room, because the same slide is usually load-bearing for the purchase. The correction does not say the methods stopped working. It says the gaps between them are small enough that the design of your round matters more than which product gets signed.
Where Olive fits
Open a role and see what the work shows
Olive supplies the second kind of evidence beside a structured round: one role-grounded assignment worked with an AI assistant, returned as six findings, each carrying the timestamped excerpt it rests on. A human reviewer writes every word, and outcomes read as demonstrated, partly demonstrated or not demonstrated.
Rank your shortlistWhat changed in the 2022 re-analysis?
The numbers, not the rank order. Sackett, Zhang, Berry and Lievens found that prior meta-analyses had systematically overcorrected for range restriction, and re-estimated the literature without that error. Most methods kept their relative position, but mean validity estimates fell by .10 to .20 points. Structured interviews came out top ranked at .42, ahead of job knowledge tests at .40, empirically keyed biodata at .38, work samples at .33, cognitive ability at .31 and unstructured interviews at .19 1.
Those are corrected correlations between a method and supervisor ratings of job performance, pooled across many jobs. They are not accuracy rates, not percentages, and not a promise about any one hiring process. The .42 for structured interviews carries an 80% credibility interval of .18 to .66, so a structured interview in a particular setting can land anywhere in that band 2.
The most useful comparison in the paper sits inside a single method: structured interviews at .42 against unstructured ones at .19, from a sample-size-weighted combination of two earlier meta-analyses 1. The difference between a scripted, scored interview and a conversation is larger than the difference between most pairs of methods on the list.
Structured means something specific, and it is not a product feature. The US Office of Personnel Management's guide defines it by three properties: the same questions in the same order for every candidate, a common rating scale, and agreement in advance on what an acceptable answer looks like 3. Any interview with those three properties is the thing the evidence describes. Any interview without them is the .19.
Why is the 1998 table still on the slide?
Because it is a good slide and nothing replaced it. The older summary gave a single ranked list with big, memorable numbers: work samples at .54, cognitive ability at .51. It travelled into textbooks, certification syllabi and vendor decks, where it has been reproduced for a quarter of a century. The 2022 paper replaced the numbers without replacing the slide, and a corrected estimate is a much harder thing to circulate than a clean ranking.
Note whose numbers moved. The .54 for work samples traces back to a 1974 narrative review, and Schmidt's own unpublished 2016 update had already dropped it to .33 on the strength of a newer meta-analysis, before the 2022 re-analysis reached the same figure 1. The estimate still being quoted was abandoned by its own lineage years before most of the decks quoting it were built.
The less comfortable reason the old table survives is commercial. Its numbers support buying something. A .54 for work samples reads as an argument for an assessment platform. A .51 for cognitive ability reads as an argument for a test. A .42 for structured interviews reads as an argument for spending three weeks writing questions and training interviewers, which nobody sells and no line item covers.
The correction is to a statistical artifact rather than to the practice, and the authors say plainly that selection procedures remain useful 1. What it changes is the size of the claim any single method supports, and therefore how much of a purchase decision should be allowed to rest on one coefficient.
Which two methods should you combine?
A scored conversation and a piece of the work, in that order of confidence. In the authors' applied follow-up, a mechanically weighted composite of predictors reaches about .61, and giving cognitive ability a weight of zero inside that composite costs .05, dropping it to .56 2. The gain comes from combining different kinds of evidence rather than from finding a better single instrument.
That number carries a condition worth repeating in the room. A composite assumes the predictors are combined mechanically with sensible weights and that the assumed intercorrelations hold 2. It is not what a team gets from stacking four unscored conversations, which mostly measure the same thing several times and then average the impressions into confidence.
Practice is a long way behind the evidence. In SHRM's benchmarking survey, structured interviews were used to assess 34% of executive, 37% of middle management and 36% of individual contributor candidates, while work sample interviews ran at 12%, 11% and 9% 4. Self-reported by HR respondents about 2021 practice, against SHRM's own definition, and organisations claim more structure than they run, so read those as an upper bound on a minority practice. The cheapest available improvement for most teams is not a purchase.
AI changes what goes inside the two methods rather than how many to run. Every estimate here was collected on job content a candidate produced alone, and most of the underlying studies tested people already in the job. If the content of a work sample is something an assistant completes for free, the method's number does not transfer to it, and the algorithm-style coding screen is the clearest case.
Read a validity coefficient for what it is
It is a correlation with a supervisor's rating, averaged over many jobs and many years, describing a method rather than a hire. A .42 does not mean forty-two percent of anything. It means that across the pooled studies, people who scored higher on the method tended to be rated higher by their managers, with a great deal of scatter inside that tendency and no guarantee about the next candidate.
Three habits keep the number honest in a meeting. Never mix the two sets: quoting .54 for work samples in the same sentence as .42 for structured interviews compares a 1998 figure with a 2022 one and produces a ranking that exists in neither paper. Quote the interval alongside the point estimate when a decision is close. And keep the criterion in view, since supervisor ratings carry their own biases, so a method that predicts them well is predicting a rating rather than performance itself.
The corrected numbers also make a demographic point the old table hid. In the same 2022 table, the mean US Black-White subgroup difference is .23 for structured interviews, .67 for work samples and .79 for cognitive ability tests 1. A subgroup difference is not adverse impact, which depends on how the score is used, the selection ratio and who applied. But the pairing is why most predictive and least disparate no longer point at the same product.
For a decision that has to survive a lawyer rather than a budget meeting, the legal comparison of a work sample against an AI screen is the version that matters. For the narrower head-to-head, work sample against structured interview puts the two estimates side by side.
Common questions
Does the correction mean cognitive ability tests are not worth running?
No. Sackett and colleagues, whose 2022 re-analysis cut the ability estimate from .51 to .31, are explicit that selection procedures remain useful. Their applied follow-up shows that ability contributes less inside a combination than its solo number suggests: giving it a weight of zero in a composite costs .05. Set against the largest Black-White subgroup difference of any method in that table, at .79, the small marginal contribution is what makes ability tests the hardest case to justify, rather than any claim that they measure nothing.
Is a higher validity coefficient always the better buy?
No, because cost, candidate experience and demographic consequences are not in the coefficient. A method predicting at .40 that takes ten minutes and applies to every applicant can be worth more in practice than one at .42 that consumes a panel day. The coefficient answers one question: how well did this method track supervisor ratings across the pooled studies. Budget, elapsed time and subgroup differences are separate questions, and the 2022 re-analysis deliberately pairs the last of those with each validity estimate.
How much does structure actually add to an interview?
In Sackett et al.'s corrected 2022 estimates, structure roughly doubles what an interview predicts, from .19 unstructured to .42 structured, drawn from a sample-size-weighted combination of two earlier meta-analyses. Before any correction, the raw observed figures were .13 and .32. The unstructured estimate's lower credibility bound sits at about zero, meaning some studies find nothing at all. Structure here means a research coding of format, not a product: same questions, same order, common rating scale.
Do these numbers say anything about AI-assisted candidates?
Nothing at all, and that limit is worth stating before anyone quotes them in an AI conversation. The underlying studies are overwhelmingly concurrent, meaning they tested people already doing the job, and none of them involve interviews conducted by AI or candidates working with an assistant. The estimates describe methods applied to content a person produced alone. Whether a given method still measures what it used to depends on whether its content survives an assistant, which is a design question rather than a validity question.
What is the cheapest way to raise the quality of a hiring round?
Add structure to the interviews already being run. It costs question writing, a rating scale, and agreement in advance on what a good answer contains, with no licence, no vendor and no extra stage in the loop. In the corrected estimates it is the single largest available move, larger than switching between most pairs of methods. SHRM's 2021 benchmarking data suggests a large share of employers have not made it, so for most teams it is still on the table.
References
- 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the corrected validity ranking, the .54 work sample lineage, the structured against unstructured gap, and the paired subgroup differences.
- 2. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors cambridge.org Supports the composite reaching about .61, the .05 cost of dropping cognitive ability, and the credibility interval around the structured interview estimate.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the three-property definition of a structured interview, from a government source rather than a vendor.
- 4. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) shrm.org Supports the gap between what the validity evidence recommends and what employers reported running in 2021.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.