Teams

How AI Use Belongs in a Performance Review, and How It Doesn't

Rate the work, not the tool. An AI-use line item in a performance review measures throughput and penalizes the person who correctly decided not to use AI on a task it would have got wrong. Where the job genuinely changed, change the criteria there: what evidence the person demanded, what they verified, what they refused to ship. If you cannot point to it inside a real deliverable, it does not belong in the review.

The takeTie a 2026 rating to AI adoption and you have built a volume incentive and called it a standard. Adoption is not a result. The moment a rating depends on how much AI someone used, the fastest route to a high mark is to run more work through a model and check less of it, which is the exact behavior a review exists to price down. The tool choice is the least interesting thing on the page, and a rating that notices it has stopped rating the work.

Where Olive fits

Open a role and see what the work shows

The same behaviors read the same way outside a review cycle: framing the problem before generating, demanding a source for the claim that carries the decision, and testing an answer against something outside the conversation. Olive reads those from a real work session rather than from a self-assessment.

Rank your shortlist

Why an AI-usage rating goes wrong

Because the only thing an AI-usage rating can measure cheaply is self-report, and self-report about AI is unreliable in a way that has been measured. Sixteen experienced developers in a randomized trial forecast a 24% speedup from AI tools, believed afterwards they had gained 20%, and were in fact 19% slower across 246 real tasks 1.

Two things follow. The first is that a rating built on what people say about their own AI use is measuring something other than what the person can do. In a study of 288 teachers in Taiwan who sat both a self-report and a knowledge-based test of AI literacy built on the same framework, correlations between the self-reported and objective factors ran from r = 0.07 to r = 0.24, and the profiles found included people who overestimated themselves and people who underestimated themselves 2. Teachers are not employees, and a weak correlation is not proof that anyone is bluffing. It is still a poor foundation for a number that follows someone into a compensation conversation.

The second is that the prize is smaller than the line item implies. A nationally representative US survey asked people who use generative AI at work how much longer the previous week would have taken without it: mean self-reported savings came to 5.4% of work hours among users, which the authors put at 1.4% of hours across all workers 3. Those are self-reports carrying the same caveat. An average that small, measured that softly, should not be carrying a rating.

A separate cost lands in the work itself. Price a method and people optimize the method. The person who tried a model on a task it handles badly, noticed, and went back to doing it by hand has done the job right, and an AI-usage scale marks them below the colleague who shipped the model's answer unread.

Where the job actually changed, and what to write there

The change is not that people use a tool. It is that a confident, plausible wrong answer now arrives inside otherwise good work, and somebody has to be the one who catches it. That is a judgment behavior, it sits on competencies most frameworks already carry, and it is the only part of AI use worth a rating at all.

So write those behaviors into the rows the framework already has:

  • Judgment. Names which part of a deliverable came from a model and says why that part was safe to accept.
  • Quality. Sends work back when the source behind a load-bearing claim cannot be produced.
  • Ownership. Takes the consequence of an error that arrived through a tool instead of attributing it to the tool.
  • Communication. Leaves enough of a trail that a reviewer can tell which decisions were made and which were generated.

Each of those is gradeable from artifacts a manager already reads. None of them asks anyone to declare how much AI they used, which is the disclosure a review should never require: it is unverifiable, it gets answered strategically the moment a rating depends on it, and it says nothing about the work. Whether those behaviors deserve a row of their own is a question for the framework, and the default answer is the rows you already have: should AI be its own competency, or folded into the ones you already have.

Write the behaviors as things a manager can point at

Every behavior in the review needs a place a manager could point to inside real work: a document, a pull request, a ticket thread, a decision memo. Assessment research pushes the same direction. Researchers building an AI literacy measure for a US Navy robotics programme reported that a scenario task simulating AI use on the job outperformed the abstract tests they had adopted from prior work or written themselves 4.

The rewrite is mechanical. Take the line item you were about to add, and ask what a manager would attach to it.

Before: *Uses AI tools effectively in day-to-day work. Rate 1 to 5.*

After: *Flagged the vendor-spend figure in the Q3 board deck as unsourced, traced it to a superseded contract, and the corrected number changed the recommendation. Demonstrated.*

The second version is longer, and that is the point. It is a claim with a receipt, so it survives a calibration meeting and a disagreement with the employee. The first version survives neither, and in calibration it collapses into whichever manager talks about AI most.

Two habits keep this honest. Collect the examples during the cycle instead of reconstructing them in review week, because reconstruction is where a tool-use anecdote quietly replaces an outcome. And treat course completion as attendance, never as evidence: a certificate says somebody sat through a session, and whether the behavior changed afterwards is a separate question with a separate answer (is AI training enough, or do you have to check the work afterwards).

What about the person who uses no AI at all?

Rate their work. If the output is on time, correct and at the level expected, there is nothing to mark down, and marking it down turns a review into a tool-adoption campaign. The question worth asking is narrower: is there a task where the team's standard method now runs through a model, and is this person's route producing a better result or only a slower one?

That question has a checkable answer, and it resolves one of three ways. The method is genuinely better and the person is behind, so coach it and note the gap the way you would note any other skill gap. The method is worse and they were right to refuse it, which belongs in the review as good judgment rather than as resistance. Or the difference does not show up in the output at all, in which case it is a preference and a review has nothing to say about it.

The reverse case deserves the same treatment. Somebody whose volume tripled this year has not automatically had a strong year, because volume is the easiest number on the page to raise without raising the quality behind it. What that does to the level above is its own question: what a promotion should be based on when AI produces a lot of the output.

See the benchmarks

Common questions

Can I require employees to disclose how much AI they used on a deliverable?

You can ask, and your own client contracts or audit rules may already require it. Just keep it out of the rating. A disclosure is useful for provenance, audit and client commitments; it is useless as a performance measure, because the answer is unverifiable and becomes strategic the moment a score depends on it. Ask for the disclosure at the point where it matters, which is when the work is handed over, and keep the review focused on whether the deliverable was correct and whether the person could stand behind the parts of it a model wrote.

The company has mandated AI adoption. How do I review against that without rewarding volume?

Convert the mandate into a task-level expectation, never a person-level score. Name the specific workflows where the tool is now the standard method, then rate whether the work coming out of those workflows meets the bar. That gives leadership a real adoption picture, because you can count workflows, and it protects the person who correctly worked around the mandate on a task the model handles badly. If leadership wants a number, give them the share of workflows converted.

Should AI use appear in a performance improvement plan?

Only as a named behavior with a named deliverable attached. A plan saying use AI more is unmeasurable and unfair, because there is no defined finish line. A plan saying draft the weekly client summary in the shared template, verify every figure against the source system before sending, and cut turnaround to one day is measurable, and it happens to be achievable with or without a model. Write the outcome, let the person pick the method, and check the outcome.

How do I review a manager when I cannot see how their team uses AI?

Review the artifacts of management rather than the tooling. Did their team ship work that held up, and when something wrong got through, did the manager find it before the customer did? Ask whether the manager has a stated point in the workflow where a deliverable can be sent back, and whether it has actually been used in the last quarter. That is observable from the record, unlike anything about how anyone prompts.

Does an AI certificate belong in the review?

It belongs in the development section, never in the rating. A certificate records participation and a passed test. It says nothing about what the person did with the training afterwards, and vendor credentials vary enormously in what they even claim to cover. If someone completed a course, note it as development, then look for the behavior in the work over the next quarter. That second step is the one that carries evidence, and it is the one most review cycles skip.

References

  1. 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that self-reported AI speedups are unreliable evidence, including the gap between a forecast 24% speedup, a believed 20% speedup and a measured 19% slowdown.
  2. 2. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the claim that self-rated AI skill and demonstrated AI skill are weakly related, cited here for the gap between the two instruments rather than for any share of workers.
  3. 3. The Rapid Adoption of Generative AI (NBER Working Paper 32966) National Bureau of Economic Research, 2025. nber.org Supports the claim that self-reported time savings from generative AI are small in aggregate, at 5.4% of hours among users and 1.4% across all workers.
  4. 4. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arXiv (Bogart, Warrier, Agarwal, Higashi, Zhang, Flot, Savelka, Burte, Sakr), 2025. arxiv.org Supports the claim that a realistic work scenario measures applied AI skill better than an abstract knowledge test, stated as the authors' own comparison.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.