Teams
Do Assessment Scores Still Predict Performance With AI in the Workflow?
Whether assessment scores still predict performance with AI in the workflow splits by role. A validity coefficient is measured against one job's performance data and travels no further. Where the criterion was production speed, AI flattened it: one support deployment raised issues resolved per hour by 14% on average and 34% for novices, with minimal effect on experienced agents [2]. Where it was accuracy or judgment, the old number probably holds, unless a model now passes the test unaided. Re-validate role by role, and check the measure still varies.
The takeHere is the part the vendor conversation skips: most of these coefficients were never yours. They arrived in a technical report computed on somebody else's job, and nothing in the years since asked them to prove themselves on your roles. AI did not degrade that evidence so much as make its age legible, by moving one criterion loudly enough that somebody finally looked. The audit will probably turn up more tests that never predicted anything on your roles than tests AI broke. The re-validation everyone is dreading is mostly a first validation.
Where Olive fits
Open a role and see what the work shows
A coefficient describes how a test tracked the job as it used to be done, and no re-analysis of it says what a person does when an assistant hands them a confident draft. Olive reads that from a real session instead: a 40-to-60-minute occupational assignment with an AI assistant, returned as six findings a human reviewer writes with the timestamped excerpt behind each one, and the candidate is granted the same report.
Rank your shortlistDo assessment scores still predict performance with AI in the workflow?
Split the question by role, because the answer splits there too. A coefficient is measured against one job's performance data and carries no further than that job. Where the performance data still rewards judgment (catching a wrong figure, refusing a bad framing), the old coefficient probably holds. Where it rewarded production speed, AI moved the criterion out from under the test.
The general validity table is not what broke. The revised meta-analytic estimates put structured interviews at .42 and general mental ability at .31, down from the .51 the field quoted for decades, and the predictors at the top of that list are the ones specific to individual jobs rather than general measures of a candidate's attributes 1. Job-specific is the operative phrase, and it cuts both ways: specificity is why those measures predict well, and why one of them predicting well tells you nothing about the next role.
Inside one org chart, the answer already forks. A support team whose performance data is tickets resolved per hour and a research team whose performance data is whether a recommendation survived the quarter are two different validity studies, and only one of them has had its criterion rearranged. Before touching a coefficient, be specific about what AI actually does inside each role. The answer is a task list, not a headcount.
Why did the tests lose their signal?
Because many of them measured production, and AI made production cheap. Across 5,179 customer-support agents, an AI assistant raised issues resolved per hour by 14% on average, 34% for novice and low-skilled agents, and minimal for experienced and highly skilled ones 2. A test that sorted fast workers from slow ones was reading a gap the tool has since closed.
The experimental evidence points the same way and names the mechanism. In a preregistered experiment with 444 college-educated professionals on occupation-specific writing tasks, time taken fell by 0.8 standard deviations and output quality rose by 0.4, and inequality between workers decreased because the tool compressed the productivity distribution by benefiting low-ability workers more 3. The same paper found the tasks themselves restructured toward idea-generation and editing and away from rough-drafting 3.
Statistically, that is range restriction arriving on the criterion side rather than the predictor side. A correlation needs both things to vary. When output rate flattens across a team, less performance variance is left for any test to explain, and a coefficient computed on the old spread overstates what the test can do now. Nothing got worse about the test; the outcome got flatter.
Read both studies at their stated scope. One is a single firm's support organization, the other an online experiment on writing tasks, and neither licenses a claim about all knowledge work 23. What they establish is direction and mechanism: the gains concentrate at the bottom of the skill distribution on production tasks, which is precisely where a production-based test used to earn its coefficient. Whatever variance remains has moved toward judgment rather than speed.
Which assessments still hold, and which are dead?
Hold: anything validated against error rate, defensibility, or a judgment the assistant cannot settle. Suspect first: timed production tests, first-draft writing samples, formulaic case math, and any online assessment a model now answers in seconds. The question is not whether AI can do the task. It is whether the performance measure the test was validated against still varies across your people.
The Uniform Guidelines are unusually concrete here, because they name the criteria a criterion-related study may use (production rate, error rate, tardiness, absenteeism, and length of service) under the standing instruction that whatever criteria are used represent important or critical work behaviors or work outcomes 4. Production rate is on that list. So is error rate. AI moved one of them hard and the other barely at all.
Three tests to run over your own battery, in the order they cost you least:
- Was the criterion a rate? Tickets per hour, tasks shipped, drafts per week. Treat the coefficient as stale until it is re-measured.
- Could an assistant produce a passing answer unaided? If so, the score now measures tool access, which is the case for an online assessment models have already solved.
- Was it a structured interview about past behavior, a hands-on work sample, or a licensure-adjacent knowledge test? Those sit near the top of the revised estimates, and none of them was validated against typing speed 1.
Watch the word "still" here. A test that predicts nothing today may have predicted nothing in 2019 either, and never been checked. Re-validation regularly turns up an instrument bought on a vendor's coefficient, applied to a role the study never covered, and left in place for years. That failure has nothing to do with AI, and the same audit finds both.
How do you re-validate an assessment against your own performance data?
Six steps, and the first one is not statistical. Name the performance measure the test was validated against, then check whether that measure still varies among your people. If it does not, no re-analysis rescues the test. If it does, correlate current scores against current performance one role at a time, and date the result so the next person knows when it goes stale.
1. Find the sentence that names the criterion. In the vendor's technical report or your own study, one sentence says what the coefficient was computed against. If nobody can produce that sentence, the test was never validated for this use, and AI is the second problem rather than the first. 2. Check that criterion for variance. Pull six to twelve months of the actual measure. If the spread collapsed after the team adopted assistants, a test cannot predict what no longer differs between people. 3. Run it per role, not per company. The Guidelines call for criteria relevant to the job or group of jobs in question 4. One coefficient over a mixed population averages away exactly the split this article is about. 4. Add an accuracy-side criterion. If output rate flattened, whatever separates people now sits on the error side: rework, escalations, corrections caught in review, recommendations withdrawn. Test against that alongside the old measure. 5. Be honest about sample size. The Guidelines put the feasibility judgment on the employer, including how many people a meaningful criterion-related study needs 4. At thirty hires a year no coefficient will be stable; take the content route instead: a job analysis, then assessment content drawn from the tasks it names. 6. Date it and set the review. Changes in the relevant labor market and in the job are named as reasons a validity study becomes outdated 4. Write the review date beside the number.
Keep the adverse-impact analysis on its own track while this runs. A re-validation answers whether the test predicts; it does not answer whether it selects unevenly, and the two studies read different data. That is the distinction behind a validated assessment versus a bias-audited one.
What do you do while the re-validation runs?
Keep the instrument, drop the cutoff. The Guidelines permit interim use of a procedure not fully supported by current evidence on two conditions: substantial evidence of validity already exists, and a study designed to produce the rest is in progress 4. They also note that evidence sufficient for pass/fail use may be insufficient for ranking 4. A score you cannot currently defend belongs behind the human read, not in front of it.
Three moves that cost a week rather than a quarter:
- Move the test after a work sample instead of before it. A screen-out at the top of the funnel is the use that needs the most evidence and, right now, has the least.
- Add one accuracy-side exercise. Hand over a deliberately flawed artifact (a draft with a wrong figure, an analysis resting on a broken assumption) and read what the candidate does with it. That is a direct measure of the behavior that still varies, and the case for hiring against verification rather than production.
- Tell the panel what the score no longer means. A stale coefficient does more damage when three interviewers still treat the number as settled.
Then close the loop where it actually breaks. Track people hired under the old bar against how they performed six months in, which is the only evidence that separates a test that stopped predicting from a candidate who aced the interview and struggled in the first quarter for reasons no instrument would have caught. Whatever you keep, the standard does not move: a selection procedure with adverse impact must be job-related and consistent with business necessity, which the EEOC frames as evaluating skills as related to the particular job rather than skill in general 5.
Common questions
Does a vendor's validity coefficient still apply to your team?
Only if the study's job and criterion match yours. A coefficient is computed against one job's performance data, and the measures that validate best are the ones built for a specific job, which is exactly why they do not transfer freely. Ask which occupation the study covered, what performance measure it used, and when the data was collected. If that measure was a production rate gathered before the team adopted assistants, treat the number as a hypothesis to test locally rather than as evidence about your hires.
How much data do you need to re-validate an assessment?
More than most teams have, which is the honest answer. A criterion-related study needs enough people in one job, and a performance measure you actually trust, for a correlation to be stable; the Uniform Guidelines put that feasibility judgment on the employer rather than naming a number. Below that threshold the defensible route is content validity: a job analysis, then assessment content drawn from the tasks it names, documented. That path produces no coefficient, and it is stronger than a coefficient computed on thirty people.
Should you stop using assessment scores until you re-validate?
Stop using them as a cutoff; keep them as one input a person reads. The Uniform Guidelines allow interim use when substantial validity evidence already exists and a new study is in progress, and they note that evidence good enough for a pass/fail decision may not support ranking. In practice: move the test behind a work sample, take the automatic screen-out off it, and tell the panel the number is provisional. Removing the instrument outright usually just relocates the same judgment somewhere less documented.
Which criterion should replace production speed?
An accuracy measure the role already produces. Rework, escalations, corrections caught in review, withdrawn recommendations, defects found downstream: pick one the team already records rather than inventing a metric for the study. Two conditions. It has to represent an important work behavior or outcome, and it has to still vary across people. If error rates are identical too, the behavior that differentiates sits upstream of anything you record, and what you need is a work sample rather than a new criterion.
Does AI in the workflow change adverse impact obligations?
No. A selection procedure that screens people out carries the same Title VII obligation whether or not AI is in the job: where it has adverse impact, it must be job-related and consistent with business necessity, judged against the particular job rather than skill in general. What changes is the evidence behind that claim. A validity study whose criterion no longer varies is a weaker defense than it was a year ago, and an unexamined cutoff inherited with a vendor contract is the hardest version to defend.
References
- 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors doi.org Revised operational validity estimates, including structured interviews at .42 and general mental ability at .31 against the long-quoted .51; the highest-validity predictors are those specific to individual jobs rather than general measures of a candidate's attributes.
- 2. Generative AI at Work nber.org Across 5,179 customer-support agents, AI assistance raised issues resolved per hour by 14% on average, 34% for novice and low-skilled workers, with minimal impact on experienced and highly skilled workers.
- 3. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence economics.mit.edu 444 college-educated professionals on occupation-specific writing tasks: time taken fell 0.8 SD, quality rose 0.4 SD, and inequality between workers decreased as the tool compressed the productivity distribution by benefiting low-ability workers more; tasks shifted toward idea-generation and editing, away from rough-drafting.
- 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR Part 1607 govinfo.gov Section 1607.14B names production rate and error rate among usable criteria and requires criteria relevant to the job or group of jobs in question; 1607.5J permits interim use with substantial evidence plus a study in progress; 1607.5G notes ranking use needs more evidence than pass/fail; 1607.5K names changes in the labor market and the job as reasons a validity study becomes outdated.
- 5. Employment Tests and Selection Procedures eeoc.gov Title VII standard: a selection procedure with adverse impact must be job-related and consistent with business necessity, evaluating skills as related to the particular job rather than skill in general.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.