Assessment design

Vendor Validity Evidence Transfers Only If the Jobs Match

A vendor's validation study usually doesn't cover your roles on its own. Borrowing it has conditions attached: the study has to meet the technical standards, and job analyses on both jobs have to show that the people in the studied job and the people in yours perform substantially the same major work behaviours. Those analyses are what make the claim transferable, and they rarely arrive. Ask for them before treating validated as an answer, and where the jobs differ, check the tool against your own current employees.

The takeValidity and fairness are two claims resting on two bodies of evidence, and a buying committee that hears one answer for both has been sold the wrong reassurance. A tool can predict performance genuinely well and still select at very different rates across groups. A tool can produce even selection rates and predict nothing at all. Validated answers the first question only. Ask the second one separately, out loud, and notice which vendors have a document for it.

Where Olive fits

Open a role and see what the work shows

Olive's twelve item banks are each grounded in one occupation and its SOC code, and a session comes back as six findings, each anchored to a timestamped moment, with the candidate granted the identical report free. Olive applies this article's test to itself: there is no criterion-related validity study to hand over, and no bias audit has been performed, because attempt volume is too low for a four-fifths ratio to mean anything, which olive.is states plainly.

Rank your shortlist

What does 'validated' actually claim?

That a score from the procedure related to some measure of job performance, in some jobs, at some point. It is a claim about a study rather than a property of the product, which is exactly the reading a certification-shaped word invites. Three variables decide whether the claim reaches your roles at all, and a deck rarely names any of them: the criterion, the jobs, and the year.

The numbers themselves moved recently enough that material built before 2022 is quoting a superseded table. Sackett, Zhang, Berry and Lievens examined the five approaches used to build range-restriction corrections in selection meta-analyses, concluded each one often produces substantial overcorrection, and re-derived the estimates: most procedures keep their rank order, with mean validity reduced by .10 to .20 points, and structured interviews rather than cognitive ability came out top 1. The mean across the old top five was .49; across the revised top five it is .37 1. Schmidt and Hunter's 1998 figures, still the most quoted table in hiring, are the ones being revised 2.

A structural caveat sits under every coefficient in either table and is almost never stated. The overwhelming majority of selection validation research is concurrent, measured on people already hired and doing the job rather than on applicants: 98% of the studies in one work sample meta-analysis, 95% in a situational judgment test meta-analysis, 74% in an interview meta-analysis 1. Those people all cleared somebody's hiring bar before they were measured, so the study is not direct evidence about how the method sorts an applicant pool. The methods still predict. The evidence behind them is more indirect than a column of decimals suggests, and that is why your own pipeline numbers will not reproduce the published ones.

Name the criterion the tool was validated against

The criterion is what the tool actually predicts, and four vendors can all say validated while predicting four different things. Supervisor ratings, first-year tenure, sales numbers and ticket throughput are not interchangeable. A tool validated against turnover predicts who stays, which is a legitimate thing to want and is not the same as who does the work well.

Ask what the assessment is about, too, because content relevance carries more of the validity than format does. In the job knowledge meta-analysis behind the 2022 revision, all 164 studies together produced a mean observed validity of .22, while the 59 studies using knowledge tests built for the job in question produced .31, rising to .40 once corrected for unreliable performance ratings 1. That is a comparison between subsets of one meta-analysis rather than a controlled experiment, and job knowledge tests presuppose candidates who already have the knowledge, so it is evidence about hiring experienced people. The direction still travels: an assessment about this job beats a general instrument about people.

One further question about the number itself: ask what it is made of. The .31 reported for integrity tests in the 2022 revision is a sample-size-weighted average of two meta-analyses that disagree by more than a factor of two, .44 against .18, and a dedicated attempt to reconcile them failed 1. A single point estimate can be a compromise between two studies that cannot be reconciled. A vendor quoting one has usually not looked, and that is worth knowing before the number goes into a board paper.

When does someone else's study transfer to your role?

When three conditions hold, and they are written down. The Uniform Guidelines, the 1978 US federal regulation at 29 CFR 1607.7, let another employer's criterion-related study be used when the evidence clearly demonstrates the procedure is valid, when appropriate job analyses on both jobs show the incumbents perform substantially the same major work behaviors, and when the studies include a study of test fairness for each race, sex and ethnic group significant in the borrowing user's relevant labor market 3.

The job analysis is the load-bearing document and the one that rarely arrives. It is what turns a study about somebody else's contact-centre agents into evidence about your hybrid analyst role, and without it a buyer is accepting a claim about jobs nobody in the room has seen. Ask which jobs were in the study and ask for the analysis on that side, then get one done on your own role, because the regulation asks for both and the buyer's half is the half that gets skipped. The third condition has an exit written into it: where a study meets the first two, carries no investigation of test fairness, and it is not technically feasible for the borrowing employer to run one, the study may be used until evidence from elsewhere shows unfairness 3. The same section adds a limit worth quoting back: where there are variables in the other study likely to affect validity significantly, the user may not rely on it 3. That is public legal fact and not advice, and it names one jurisdiction; whether a particular purchase clears it belongs in front of counsel.

Three more questions fit in one email: the sample size, the year, and whether the model has been retrained since. A retrained model is not the model that was validated, and a retrain date is rarely volunteered. This bites harder for machine-scored formats, where the published evidence is thinner. A peer-reviewed study of one vendor's machine-scored video interviews reports an uncorrected .24 with job performance across five organisational samples totalling 1,124 people, four of its five authors employed by that vendor, against an uncorrected .32 for human-rated structured interviews in the same comparison 4. Both figures are uncorrected, which makes the comparison internally fair, and neither is the corrected .42 usually quoted for structured interviews. That figure rests on five samples, against the 105 effect sizes behind the structured-interview estimate it is set beside 4.

What to do when transport is weak

Run a small concurrent check against people already doing the job rather than commissioning a full study. Give the assessment to a dozen or two current employees whose performance you can already characterise, and look at whether the results order them in a way anyone recognises. It is weak evidence by design and it is local, which is precisely the trade being made.

Two things to hold steady while doing it. A check on a dozen people is not a validation study and should never be described as one, inside the company or outside it. And the range-restriction problem that produced the 2022 revision applies to your own pipeline data as well, so your numbers will sit below the published ones for reasons that have nothing to do with the tool 1.

Where transport stays weak, the honest conclusion is usually about what job the tool should be given. An instrument that is one input into a human decision needs less evidence behind it than a gate that stands alone and ends applications, because the failure mode is different and recoverable. That is also the cheapest configuration change available, and it is reversible in an afternoon. How to run the tool before it becomes a gate is the pilot design question; the wider purchase checklist, integrations and pricing included, is how to evaluate an AI-skills assessment before buying one.

On Monday, add three lines to the vendor questionnaire: what was the criterion, which jobs were in the study and can the job analysis be shared, and when was the model last retrained. Write down which vendor answered which, because that table is the most useful artifact the evaluation will produce.

See what gets scored

Common questions

Is a meta-analysis good enough evidence on its own?

It is evidence about a method in general, not about the product in front of you or about your roles. A vendor citing a meta-analysis is saying that assessments of this broad type have predicted performance somewhere, which is a weaker claim than it sounds. Ask what evidence exists for this tool, on which jobs, against which criterion. If the answer is only the meta-analysis, what you are buying is a category.

What is a job analysis, and who writes it?

A structured description of what the job actually consists of: the major work behaviours, the knowledge and skills each one requires, and how often each occurs. It is written by someone who observes and interviews people doing the work, usually with the manager and one or two incumbents. Vendors sometimes have one for the jobs in their study, which is the document that would make their evidence transferable. If neither side has one, that gap is the honest finding of the evaluation.

Does 'validated' mean the tool is legal to use?

No. Under US federal selection law, validity evidence is one part of a job-relatedness argument, and job-relatedness only becomes the question once a selection procedure shows adverse impact. Legality also depends on how the tool is configured, how candidates with disabilities are accommodated, and what notices apply where the candidate sits. Treat a vendor's validity claim as evidence you may need rather than as a compliance conclusion, and take the compliance question to counsel in the relevant jurisdiction.

Can I ask a vendor for the study itself?

Yes, and the response is informative whatever it is. Ask for the technical report rather than the summary slide: the sample, the criterion, the jobs, the year, the correlations before and after any corrections, and the job analysis. Some vendors will send it under an agreement, some will send a redacted version, and some will offer only the deck. Sorting vendors by what they will hand over separates them more sharply than any feature comparison.

What if the role did not exist when the study ran?

Then the vendor's study does not transport, by definition, and the honest move is to say so in writing before anyone stretches the analogy. Roles reshaped by AI-assisted work are the common case here: the tasks moved even where the title did not. Compare work behaviours rather than titles, keep the tool as one input into a human decision, and run a small concurrent check against current employees so at least some of your evidence is about the job as it exists now.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology, 107(11), 2040-2068 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2022. static1.squarespace.com Supports the 2022 downward revision and its size, the concurrent-study share behind published coefficients, the job-specific knowledge test comparison, and the integrity-test estimate averaging two irreconcilable meta-analyses.
  2. 2. The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings Psychological Bulletin, 124(2), 262-274 (American Psychological Association); full text opened at homepages.se.edu, 1998. homepages.se.edu Named as the superseded table a pre-2022 vendor deck is likely to be quoting, rather than cited for any current estimate.
  3. 3. 29 CFR 1607.7 - Use of other validity studies Code of Federal Regulations, via Cornell Legal Information Institute, 1978. law.cornell.edu Supports the three conditions for borrowing another user's criterion-related study, the exception where no fairness investigation exists and an internal one is not technically feasible, and the limit that a study with variables likely to affect validity significantly may not be relied on.
  4. 4. Psychometric Properties of Automated Video Interview Competency Assessments Journal of Applied Psychology, 109(6), 921-948 (Liff, Mondragon, Gardner, Hartwell and Bradshaw; four of five authors employed by HireVue); accepted manuscript hosted by HireVue, 2024. hirevue.com Supports the uncorrected .24 job-performance figure for one vendor's machine-scored video interviews across five samples of 1,124 people, its vendor authorship, the uncorrected .32 comparison, and the five-studies-against-105-effect-sizes gap the paper states itself.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.