Teams

Promotion Criteria That Survive AI-Assisted Output

Promote on decisions, not deliverables. AI raised output volume across the team at once, so shipped surface area no longer separates a senior from a strong mid-level. What still separates them is what the person chose not to build, which claim they refused to pass along, and the check they ran that changed a recommendation. Make the promo packet carry one or two artifacts with that reasoning attached instead of a list of things shipped.

The takeBolting "uses AI tools" onto the lower rungs of a ladder is the cheapest available response and the least useful one. It adds a requirement nobody can fail and leaves the rung boundaries exactly where they were, resting on a volume proxy that has stopped working. If a ladder revision does not change what evidence a promotion committee reads, it has not changed anything. Rewrite the evidence first, and most of the rung wording will follow on its own.

Where Olive fits

Open a role and see what the work shows

Olive's six dimensions describe the same markers a promotion committee is reading for: keeping the judgment that should not be handed over, turning down an output that does not hold up, and testing a claim against something outside the conversation. Every finding in the report is written by a person and carries the moment in the session it rests on.

Rank your shortlist

Why volume stopped separating people

Because the measured gains from AI concentrate at the bottom of the distribution, which compresses exactly the spread a ladder used to read. In a pre-registered experiment, 444 professionals doing occupation-specific writing tasks finished 37% faster with ChatGPT and scored 0.45 standard deviations higher, with the largest gains going to the weakest writers 1.

The consulting version is larger and points the same way. Across 758 Boston Consulting Group consultants working 18 realistic tasks chosen to sit inside the model's capability, the group given GPT-4 completed 12.2% more tasks and produced more than 40% higher quality by grader score; consultants below the average performance threshold gained 43% against their own baseline, those above it 17% 2. On the one task deliberately placed outside that capability, the same tool left consultants 19 percentage points less likely to reach the correct answer 2. Both studies are set-piece tasks done once and graded by evaluators, on 2023 models. Neither is a claim about a year of real work.

What travels is the shape. Production got cheaper for everybody, and cheapest fastest for the people who had been slowest at it. A promotion packet listing what someone shipped is now reading a number that moved for reasons unrelated to the person, and the mid-level engineer's list looks a great deal like the staff engineer's list.

That is not an argument that levels collapsed. It is an argument that the proxy collapsed. Scope and impact were always the criteria; volume was the cheap way to infer them without reading the work. The cheap way is gone, and two options are left: read the work, or promote on impressions.

What still separates a senior from a strong mid-level?

The judgment applied to the output, and it shows up most clearly in a study where AI widened the gap instead of narrowing it. Among 640 Kenyan small-business owners given a GPT-4 assistant over WhatsApp, there was no detectable average effect: high performers at baseline gained just over 15% while low performers did about 8% worse, a spread driven by which advice owners chose to act on rather than by differences in the advice itself 3.

Small businesses in Kenya are not a promotion committee, and the outcome measured there is revenue rather than task quality. The mechanism is what carries. When everyone can obtain a plausible answer, the difference between people is which plausible answers they accept.

Three level markers follow from that, and all three predate AI:

  • The scope call. What the person decided not to build, and the reasoning that got them there. Cutting the right half of a project is the most senior item on this list and the least visible in any count of shipped things.
  • The refusal. A claim, a figure or a recommendation the person declined to pass along because the support behind it did not hold. Name the claim and name what was missing.
  • The verification that changed something. A check that produced a different answer, and the decision that moved as a result. A check confirming what everybody already believed is fine work and weak evidence.

None of those needs a new competency. They need somebody to read a piece of work closely enough to see them, which is the cost the volume proxy existed to avoid. It is also why a test result alone cannot settle a promotion: do assessment scores still predict how someone performs once AI is in the workflow.

Rewrite the promo packet around two artifacts

Ask for two artifacts and the reasoning behind them, not a list. One should be something the candidate decided not to do, with the thinking that led there. The other should be a deliverable where something a model produced was wrong or unsupported and the candidate caught it. Both are short, and both are checkable against people who were in the room.

A workable packet fits on three pages:

1. One paragraph on the scope of the work over the period, in the candidate's own words. 2. Artifact A, the decision not taken: the document, thread or memo, plus two sentences on what the alternative would have cost. 3. Artifact B, the catch: what arrived, what was wrong with it, how it was found, what changed downstream. 4. Two names who can confirm each artifact independently.

Cut three things from the packet: the list of shipped items, the self-rating, and the tool inventory. The tool inventory is the one that looks most like evidence and carries the least. Knowing which assistants somebody holds an account for says nothing about the level they are working at, which is also why hiring for tool fluency so often disappoints (the AI-skills hire who works exactly like the rest of the team).

The packet is short enough that a manager cannot pad it, and that is a large part of why it works.

How to run a committee that can read for judgment

Give the committee less to read and a rule about what counts. A packet of forty shipped items is not evidence a group can weigh, so it gets weighed by whoever wrote the most confident summary. Two artifacts of a page each, plus two named people who can confirm the decision happened, fits inside a meeting and survives a challenge from the candidate afterwards.

Then hold the room to three habits. Read the artifact before the summary, because the summary is the part written to persuade. Ask one follow-up per artifact that only somebody who did the work could answer, which is the cheapest authorship check available and the only one that behaves well under pressure. And write the decision down next to the artifact it rested on, so the next cycle calibrates against something firmer than memory.

Name the two failure modes out loud. The first is promoting the person who talks about AI most, which is a presentation skill and predicts nothing. The second is the mirror image: marking down heavy AI use as a proxy for lightweight work, which punishes the method rather than the result and quietly turns a ladder into a tool policy. The review cycle one level down has the same problem and the same fix: how AI use belongs in a performance review, and how it doesn't.

See the benchmarks

Common questions

Should a career ladder name AI tools on any rung?

Name behaviors, not products. A rung that says "familiarity with a named assistant" expires when the tool does, and it is satisfied by anyone with a login. A rung that says "can state which parts of a deliverable were generated and why each was safe to accept" describes something a reviewer can check and something that stays true across tool changes. If procurement needs a tool inventory, keep it in the tooling doc where it belongs.

How do I level someone whose output tripled but whose judgment has not been tested?

Give them a decision to own before the cycle closes, then read what happens. Volume without a tested call is a mid-level pattern, however impressive the throughput, and promoting on it produces a senior title over a person who has never had to say no to anything. A scoped decision with a real cost attached, made over a few weeks with a named outcome, is enough evidence for one rung. Waiting a whole cycle to find out is the expensive version of the same test.

Does this mean scope and impact stop being the criteria?

No. Scope and impact stay exactly where they were; what changed is the evidence used to infer them. Committees rarely measured scope directly, because reading the work is expensive, so they read the volume and the surface area instead and treated those as a stand-in. That stand-in has stopped tracking the thing it stood for. The criteria are unchanged and the evidence has to be rebuilt.

What if the strongest artifact is confidential?

Write the decision without the confidential content. A one-page account of what the call was, what the alternatives cost, and who confirmed it does the work of the artifact without reproducing it, and confidentiality is a normal condition for senior work rather than an exception to the process. The committee is assessing the reasoning and the confirmation, not reading the underlying document. Set the redaction rule once so candidates are not guessing at it each cycle.

Can this run alongside the ladder already in place?

It can, and running it alongside is usually the cheaper path. Leave the rung descriptions untouched for a cycle and change only the packet: two artifacts, the reasoning, two confirmations. Committees then discover in the room which rung wording is doing no work, which is far better information than a ladder rewrite drafted in advance. Rewrite the wording after one cycle, against what the committee actually argued about.

References

  1. 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) MIT Department of Economics, 2023. economics.mit.edu Supports the claim that measured AI writing gains are largest for the weakest performers, at 37% faster and 0.45 standard deviations higher across 444 professionals.
  2. 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that AI compresses the performance distribution on tasks inside its capability, with below-average consultants gaining 43% against their baseline and above-average consultants 17%, and the same experiment's single out-of-frontier task, where the AI groups were 19 percentage points less likely to be correct.
  3. 3. The Uneven Impact of Generative AI on Entrepreneurial Performance eScholarship, University of California (Berkeley Haas / Harvard Business School), 2024. escholarship.org Supports the claim that the remaining spread between people comes from which AI output they act on, with high performers up just over 15% and low performers about 8% worse.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.