Teams

Seniority Prices Judgment, and Judgment Just Got Scarcer

AI makes senior experience worth more and less at once, because seniority was always two things bundled together: how fast you produce, and how well you catch what is wrong before it ships. The measured gains land on the first half, and they concentrate in less experienced workers. Judgment moves the other way: in one field trial, an assistant left people less likely to reach the right answer on a task outside its strength. What decides your position now is which half of your record you can point to.

The takeThe uncomfortable part of this is that throughput was carrying more of a lot of senior reputations than anyone wanted to admit. Being fast used to look identical to being good, because slow work rarely got the chance to prove it was also careful. Now that a junior with an assistant can match your pace on a first draft, the two have come apart, and some experienced people are going to discover that speed was most of what they were selling. That is not a reason for panic. It is a reason to find out, honestly, which half of your own value was ever the judgment half.

Where Olive fits

Open a role and see what the work shows

The six dimensions Olive scores name what capable AI work looks like: framing before generating, sourcing the claim that matters, keeping the judgment you shouldn't delegate, and verifying against something outside the conversation. They are a usable study list whether or not you ever meet the assessment.

Rank your shortlist

Name the Two Halves of What You're Paid For

Seniority has always bundled two different things into one number on an offer letter: how much you produce, and how well you catch what is wrong with it before someone else has to. For most of a career the two moved together closely enough that nobody separated them, because the person trusted to move fast was usually the one asked to check the work. Seniority never guaranteed the second half. It just rarely got priced on its own.

AI breaks that bundling apart by attacking only one half of it directly. It is very good at producing a plausible first draft quickly, which is the throughput half. It has no comparable track record of knowing when that draft is confidently wrong, which is the judgment half. Watching a junior colleague produce your output in a fraction of the time is real and worth noticing. It is evidence about the throughput half of your value, not a verdict on the whole thing.

The reason this feels so personal is that most performance reviews never separated the two halves either. "Fast and reliable" reads as one compliment, not two, until something arrives that is fast without being reliable, and suddenly the two words need to be pulled apart to mean anything at all.

Which Half Is Actually Compressing?

The throughput half is compressing measurably, and the compression concentrates on people with less experience rather than more. Across three pooled company-run trials covering 4,867 software developers, an AI coding assistant raised completed tasks by 26.08%, a figure carrying a standard error of 10.3% that no single one of the three experiments was precise enough to establish on its own, and less experienced developers both adopted it faster and gained more from it than senior colleagues on the same tool 1.

A separate study of 5,179 customer support agents at one software firm, identified off a staggered rollout rather than individual randomization, found the assistant lifted issues resolved per hour by 34% for novice and low-skilled agents while having minimal impact on experienced, highly skilled ones 2. Read together, these say the same thing from two angles: AI is closing the production gap between junior and senior, which is exactly the compression a senior person feels when a newer colleague suddenly moves at their pace.

The judgment half tells a different story. In a field experiment with 758 consultants at one firm, those given GPT-4 on 18 tasks inside its capability finished 25.1% faster and scored more than 40% higher on grader-rated quality 3. On one task chosen deliberately to sit outside that capability, the same tool left them 19 percentage points less likely to reach the correct answer, averaged across the two AI conditions 4. Telling those two situations apart, recognizing when a task has quietly moved outside what the tool handles well, is a judgment call, and it is the exact skill that does not get easier just because drafts arrive faster.

Why Judgment Is Getting Scarcer, Not Safer

Judgment gets scarcer for a plain supply reason: the volume of plausible-looking output rose sharply while the number of people who reliably catch what is wrong with it did not. Hold that loosely, because it is a mechanism rather than a promise, and employment data is a lagging indicator that looks like safety right up until it stops. It is still the direction the available evidence points.

One separate finding, drawn from a payroll panel rather than a survey, points the same way at the aggregate level: occupations built more around tacit, situational judgment showed faster employment growth for mid-career and senior workers over the same window that codified, formally-taught work showed slower entry-level employment growth 6. The paper's own authors call this a raw, non-causal pattern rather than a proven mechanism, but it is consistent with the same story: the parts of work built on years of pattern-matching against real situations are not the parts currently getting cheaper to produce.

Experienced people can also be wrong about their own speed, in direction rather than only in size. In a randomized trial, 16 experienced open-source developers working on repositories they had maintained for years finished 19% slower with early-2025 AI tools, after forecasting a 24% speedup beforehand and still estimating a 20% gain afterward 5. Sixteen developers in one setting is not a general productivity result, and the paper says so. What travels is the gap between belief and measurement: if people with years of practice can misjudge their own AI-assisted output from the inside, a check from outside the work is the thing that keeps holding its value.

Change What Your Record Shows

The concrete move is to shift the evidence you carry from what you produced toward what you decided and caught. A senior resume built entirely around volume, lines shipped, tickets closed, decks turned around overnight, is now competing directly against a junior with an assistant, and throughput alone is the weakest ground to defend there.

A record built around specific judgment calls does not have that competition, because no assistant is producing that record on its own. The time you rejected a plausible-looking recommendation and why, the assumption you caught before it reached a client, the moment a number looked right and you checked it anyway and found it wasn't: these are the entries a resume of pure output cannot contain, and they are exactly the entries a fast junior with an assistant has no equivalent for yet.

Whether a fast junior with AI should be hired over an experienced senior is answered on the employer side by the field rather than by years served, and the advice there includes confirming that the senior verifies anything at all, because seniority on its own does not supply that. Read from your chair, that is the instruction: make the checking legible instead of assuming tenure implies it. The real risk here is not that judgment stopped being valuable. It is leaving a record built entirely around output while the tasks inside your own job change shape.

See the benchmarks

Common questions

Is AI actually replacing senior workers, or just junior ones?

The measured evidence points the other way. The clearest gains concentrate in less experienced workers on tasks with a checkable right answer, and in one field experiment an assistant left people less likely to reach the correct answer on a task outside its capability. That is compression of one part of senior value, not replacement of the whole role.

How do I prove judgment rather than just claiming it?

Specifics beat description. A record of a particular moment you caught a wrong assumption, what you checked it against, and what changed as a result is concrete evidence an interviewer can question. A general claim to have good judgment is not.

Should I stop emphasizing speed and volume on my resume entirely?

Not entirely, since some roles genuinely reward fast, checkable output. The shift is about balance: if every line on your resume is about volume and none is about a decision you made or a mistake you caught, that resume is competing on the ground AI compresses fastest.

Does this mean younger workers with AI skills have an advantage over me?

On raw throughput, in some tasks yes: the trials that measured production gains from AI assistance found them concentrated in less experienced workers. Those trials measured speed and volume on work with a checkable answer, not judgment on ambiguous calls, so they say nothing about the second gap. The two are not the same skill, and a resume can show one without showing the other.

What if my current role really is mostly production work?

That is worth taking seriously rather than dismissing. If a role is genuinely built around throughput a model now matches, the honest move is toward roles built around checking and deciding, not toward hoping the market reverses on its own.

References

  1. 1. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers MIT Department of Economics (working paper; later Management Science), 2025. economics.mit.edu Supports that AI coding assistance raises completed-task volume, with less experienced developers adopting faster and gaining more than senior colleagues.
  2. 2. Generative AI at Work (NBER Working Paper 31161) National Bureau of Economic Research, 2023. nber.org Supports that measured AI gains concentrate in novice and low-skilled workers with minimal effect on experienced ones, backing the throughput-compression claim.
  3. 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the speed and quality gain on tasks inside an assistant's strength, used here as the throughput half of the split.
  4. 4. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the accuracy drop on a task outside an assistant's strength, used here as evidence the judgment half is not compressing the same way.
  5. 5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports that even experienced practitioners misjudge their own AI-assisted speed, evidence for why an external, skeptical check stays valuable.
  6. 6. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence Stanford Digital Economy Lab (Erik Brynjolfsson, Bharat Chandar, Ruyu Chen), 2026. digitaleconomy.stanford.edu Supports the codified-versus-tacit-knowledge pattern showing faster employment growth for mid-career and senior workers in tacit-knowledge-heavy occupations.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.