Assessment design
Is Your Assessment Broken After the Latest Model Release?
A pass rate that jumps on a work sample right after a model release usually means the task got easier, not that the assessment broke. If the sample scores the artifact and AI is allowed, a better model lowers what the task costs to finish, so the same cut score passes more people at the same level of ability. Check three cheaper explanations first: pool composition, case leakage, grader drift. Then re-anchor the cut score per occupation against a freshly marked reference set, and record what changed and when.
The takeNotice which direction the surprise came from. A pass rate that jumps gets a meeting by Friday. One that sags after a quiet rubric edit can sit for a year, and I doubt many programs have ever escalated a bar for being too strict. A trigger only fires in the direction someone is watching, which is the entire argument for putting a date on the calendar. Then be plain about what that cadence is. An output-scored work sample rents its difficulty from a release schedule nobody at your company sets, and the rent goes up.
Where Olive fits
Open a role and see what the work shows
Olive returns six separately evidenced findings from one candidate's session (problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification), each anchored to a moment in the record rather than to the polish of the artifact. Every released report carries its rubric, scorer and bank versions, and the candidate is granted the same document.
Rank your shortlistDid the assessment break, or did the task get easier?
The task got easier. A work sample scored on the quality of what comes back is measuring the candidate and the assistant together, and only one of those changed last month. When the assistant's contribution rises, the joint output rises with it, and a cut score frozen at last year's floor now passes people who would have failed in March. The instrument is intact. Its anchor is stale.
Benchmarks show how big a step can be. On SWE-bench, a coding benchmark introduced in 2023, AI systems could solve just 4.4% of coding problems in 2023, a figure that jumped to 71.7% in 2024 3. Two other benchmarks published the same year moved 18.8 and 48.9 percentage points in the same window 3. Those are research benchmarks rather than hiring tasks, and the exact figures are not the point. The point is that the difficulty floor of an AI-assisted task moves in steps, on a release calendar nobody at your company controls.
METR measured the same thing as duration. Across six years of models, on a suite of software and reasoning tasks, the length of task an agent completes with a 50% success rate has been doubling roughly every seven months 4. Read that as a maintenance fact rather than a forecast: the share of your two-hour assignment a candidate can hand over completely is larger this quarter than it was last quarter, and it will be larger again.
Nothing about your candidates had to change to produce the jump you are looking at. Whether assessment scores still predict performance with AI in the workflow is a validity question and deserves its own answer. This one is narrower and more urgent: the number you compare against was set against a task that no longer costs what it cost.
What else could have moved the rate that week?
Start with three suspects that are not the release. Your pool changed: a new job board, a referral push, a rewritten posting. Your case leaked: twelve weeks in the wild is enough for a packet to be posted somewhere. Your graders drifted: a new reviewer, an edited rubric, calibration that decayed quietly. Separate them before you touch the bar, because each one has a different fix.
The cheapest discriminator is a blind re-score. Pull 20 archived submissions from the quarter before the release, strip the dates and the names, and put them through today's rubric with today's reviewers. If the old work now scores higher, your graders moved and the model is innocent. If it scores the same, the submissions themselves changed, and you have narrowed the problem to the task.
Then split the live rate two ways. By source, because a rate that rose in one channel only is a pipeline story. By week, because a release is a step and a hiring push is a ramp. A step that lands within days of a release, holds, and shows up across every source at once is the signature worth acting on.
Read the shape of the distribution, not just the headline number. A release that lifts everyone shifts the whole curve right and thickens the middle: submissions that used to land just under the line now land just over it. A leaked case looks different. It produces a cluster of near-identical strong answers carrying the same structure and the same omissions, and it usually leaves the weakest band untouched.
How often should you re-anchor the cut score?
On a calendar, plus a trigger. Quarterly is the shortest cadence a small program can actually sustain for an AI-allowed task, and it is short enough that no single release sits unexamined for long. The trigger is a rate outside its historical band for two consecutive intake windows. Anything faster becomes a rewrite treadmill; anything slower lets the bar drift through a whole hiring season.
Re-anchoring is not nudging the number until the rate looks familiar. Under the federal Uniform Guidelines on Employee Selection Procedures, in force since 1978, cutoff scores "should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force" 1. That sentence is the whole job. Acceptable proficiency is a property of the work, and if the work now includes an assistant that produces a competent first draft, last year's cut no longer describes it.
So anchor against work, not against last quarter's percentage. Take 25 to 40 recent submissions per occupation, have two reviewers mark them independently against the answer key, and put the cut where acceptable proficiency actually sits in that marked set. Hold the set out of whatever tuning follows, or you are fitting the bar to the sample. The same discipline applies the first time you are setting a defensible bar for good enough at AI.
The Guidelines also treat validity as something that stands once shown, until the study is subject to review for currency 1. A model release is exactly the kind of event that makes a review due, and writing that down converts an awkward Tuesday into a scheduled item with an owner's name on it.
Which parts of the score don't move when the model does?
The candidate's own decisions. What they framed before generating anything, what evidence they demanded and actually opened, what they kept rather than handed over, what they refused and on what grounds, what they tested against something outside the conversation. A better model writes a better draft. It does not make a candidate go and check the figure the recommendation rests on.
What stops separating people, and what still does, by the work you are hiring for:
- Financial analysis. Every memo comes back clean, so writing quality separates nobody. What still separates: whether the growth rate was reconciled to the filing it came from, and what the recommendation did when it would not reconcile.
- Software engineering. The patch satisfies the suite it was handed, whichever model wrote it. What still separates: whether a failing test was written before the fix, and whether the candidate named the case the suite never covered.
- Marketing. Positioning briefs read publishable across the board. What still separates: whether the headline market figure was opened, and what changed when it turned out to be a vendor poll rather than an estimate.
- Healthcare revenue cycle. The appeal letter reads persuasively from any assistant. What still separates: whether the denial that should be conceded was conceded, and on what stated grounds.
- Data and analytics. The interpretation is fluent in every submission. What still separates: whether anyone recomputed the result the interpretation rests on.
Scoring that column is a change to your rubric, not to your task, and it holds its meaning across releases because none of it is a property of the output. The catch is that it needs a record of the work, and an unwatched take-home has none: you get the deliverable and not the decisions. That is the same constraint behind what a work sample should test now that AI can produce it, and it is why a short verification note in the brief buys more than another rubric row about polish.
Write the maintenance schedule before the next release
Put five lines in the program doc and the next release becomes a scheduled task instead of an emergency. Name the owner. Name the metric and its baseline, per occupation rather than blended. Keep a marked reference set. Set the calendar cadence and the out-of-band trigger. Record every change to a cut score with its date, its evidence, and the person who signed it.
The record is not paperwork. The Uniform Guidelines say each user should maintain records disclosing the impact its selection procedures have on identifiable race, sex and ethnic groups, and a selection rate for any such group below four-fifths of the highest group's rate is generally regarded by the federal enforcement agencies as evidence of adverse impact 2. Moving a cut score moves every one of those ratios. Recompute them on the marked set before the new bar goes live, not after the quarter closes.
Resist the urge to patch the task after every release. A task edit breaks comparability with your own history, and that history is the only baseline you own: change the case and last quarter's numbers stop meaning anything, which is the one thing you cannot buy back. Change the anchor first, and treat a task rewrite as its own project with a pilot before it becomes a gate.
Then be honest with the hiring managers about what maintenance buys. Re-anchoring keeps the bar describing acceptable proficiency. It does not make an output-scored task hold still across releases, and the next one will move it again. The only part of the design that stops drifting is the part that reads the candidate's decisions rather than the artifact, which is a rubric change worth scheduling in the same quarter as the re-anchor.
Common questions
Should you raise the cut score right after a model release?
Not by feel, and not until the rate looks like last quarter's. Re-anchor instead: have two reviewers mark 25 to 40 recent submissions per occupation against the answer key, then set the cut where acceptable proficiency sits in that marked work. A number chosen to restore a familiar pass rate is a target, not a standard, and it will not survive the first candidate who asks what the bar measures. Log the old value, the new value, the date and the evidence.
Does a higher pass rate mean the assessment is no longer valid?
No. Those are different questions. Validity is about whether the inference from the work to job performance holds; a cut score is only where you draw the line. A release can lift every submission at once, which moves the rate without changing what the assessment is evidence of. What it does undermine is a bar anchored to an older version of the task, and that is repaired by re-anchoring rather than by replacing the instrument. Check the graders in the same pass, since a rubric edit produces the same symptom.
How do you tell a model release from a change in your candidate pool?
Re-score old work blind. Pull 20 submissions from the quarter before the release, strip dates and names, and run them through today's rubric with today's reviewers. Scores that come out higher point at the graders. Scores that come out the same point at the submissions. Then split the live rate by source and by week: a release shows up as a step within days and across every channel at once, while a sourcing change shows up in one channel and ramps.
Should you ban AI on the work sample to keep the bar stable?
A ban does hold the rate still, and it holds it still by measuring something the job no longer asks for. If the role uses an assistant daily, an unassisted sample answers a question nobody has. The stable thing to score is not the absence of the tool but the decisions made inside it: what the candidate framed, what evidence they demanded, what they refused, what they tested against something outside the conversation. Those do not move when the model does.
What belongs in the record when you move a cut score?
The old value, the new value, the effective date, the marked reference set it was derived from, the two reviewers who marked it, and the subgroup selection rates recomputed at the new bar before it goes live. Add the reason in one sentence, naming the release or the scheduled review that triggered it. That record is what turns a defensible decision into a demonstrable one, and it costs about ten minutes at a moment when you are already doing the work.
Which roles drift fastest when a new model ships?
The ones whose assignments sit closest to what current models do well: code, structured writing, and summarizing a supplied packet. A software task scored on whether the patch passes tends to move first and furthest. Work whose difficulty lives in reconciling sources that disagree, or in deciding what to concede, moves more slowly, because an assistant will still produce a confident answer the material does not support. Put the fast ones on a shorter re-anchoring clock than the rest.
References
- 1. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.5 (General standards for validity studies) ✓ law.cornell.edu Paragraph (H): where cutoff scores are used they should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force. Paragraph (K), review of validity studies for currency: additional studies need not be performed until the validity study is subject to review.
- 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 (Information on impact) ✓ law.cornell.edu Paragraph (A): each user should maintain and have available for inspection records disclosing the impact its tests and other selection procedures have on employment opportunities by identifiable race, sex or ethnic group. Paragraph (D): a selection rate below four-fifths of the rate for the highest group is generally regarded by the federal enforcement agencies as evidence of adverse impact.
- 3. The 2025 AI Index Report, Chapter 2: Technical Performance ✓ hai.stanford.edu On SWE-bench, AI systems could solve just 4.4% of coding problems in 2023, a figure that jumped to 71.7% in 2024; gains of 18.8 and 48.9 percentage points on MMMU and GPQA over the same period.
- 4. Measuring AI Ability to Complete Long Tasks ✓ metr.org The length of task AI agents complete at a 50% success rate has increased exponentially over the past six years, with a doubling time of around seven months.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.