Teams

What Separates Heavy AI Use From Good AI Use

Read the work, not the dashboard. Good AI use leaves four traces: a framing or criteria list written before anything was generated, a source opened for the one claim the decision rested on, model output thrown away with a stated reason, and a check run against something outside the chat. Ask for those on two recent deliverables and heavy but unexamined use shows itself in about ten minutes each. Seat counts and prompt volume cannot show any of it.

The takeThe adoption dashboard is not a weak measure of AI capability. It is a measure of something else that happens to be easy to collect. The ability to say what a model got wrong is scarce and hard to fake. The ability to send it a prompt is neither. A team that keeps reporting the easy number upward as capability will keep being surprised by the work that ships.

Where Olive fits

Open a role and see what the work shows

Olive's six dimensions describe the same ground: framing before generating, sourcing the claim that carries the decision, keeping the judgment that should not be handed over, rejecting output that should not stand, and testing a claim against something outside the conversation. Olive reads those from a real working session rather than from a self-assessment.

Rank your shortlist

Why does the usage dashboard say nothing about quality?

Because it counts actions, and quality lives in the judgment around them. Microsoft's own admin report calls someone an active Copilot user when they take one intentional action inside the chosen window, and a single submitted prompt in twenty-eight days qualifies 1. That number moves when a person presses a button. It does not move when someone reads an output and decides it is wrong.

Self-report is not a fallback either. In a randomized trial, 16 experienced developers working on repositories they had known for years forecast that AI tools would cut completion time by 24%, still believed afterwards that it had cut it by 20%, and were measured 19% slower across 246 real tasks 2. Sixteen developers in one setting is not a general productivity finding, and the authors do not offer it as one. What travels is the gap: people were wrong about the direction, not only the size.

Asking the team to rate its own AI skill has the same problem in a different coat. When 288 teachers took both a self-rating and a knowledge test built on the same four-part framework, correlations between what they claimed and what they demonstrated ran between 0.07 and 0.24 3. That is a study of teachers rather than employees, and it does not show that people flatter themselves; the profiles included underestimators. It shows the two instruments measure different things, which is enough to stop using one as a proxy for the other.

The three easiest signals to collect, then, are seat activity, self-reported speedup and self-rated skill. They are also the three that carry the least.

What four traces does good AI work leave?

Four, and each one is visible in the artifact rather than in a tool. A frame written before anything was generated. A source actually opened for the claim the decision rested on. Something the model produced that was thrown away, with a reason attached. A check run against something outside the conversation: a query, a test, a person who would know.

  • A frame written first. Criteria, constraints, what a good answer would have to contain. It can be four lines in a scratch file. Its absence looks like a deliverable that answers a question nobody quite asked.
  • One source opened. Not every claim. One: the claim the decision rested on. The trace is a link, a page number, a screenshot, or an account of what the source actually said next to what the model said it said.
  • Something rejected. A draft, a paragraph, an approach, a whole first attempt, with a sentence about why it went. A body of work with no rejections in it was accepted rather than edited.
  • A check against the outside. A number confirmed against the system that holds the number. The cited case read. The claim put to someone who would know if it were false.

None of the four requires anyone to use AI less. All of them are compatible with running every step through a model, and none of them is produced by running every step through a model. That is what makes them worth asking for: the answer cannot be inferred from how much the tool was used, in either direction.

Run the ten-minute review on two recent deliverables

Pick two things that shipped in the last month, sit down with whoever wrote them, and ask four questions about each. It takes about ten minutes per deliverable and needs no tooling, no dashboard and no new policy. The questions are about the work rather than about AI, which is what keeps the conversation from turning into a defence.

Ask what the criteria were before drafting started. Ask which single claim the decision rested on, and what was opened to check it. Ask what got thrown away, and why. Ask what would have had to be true for the conclusion to be wrong, and how that was tested.

A good answer is specific and slightly boring: the number came out of the model's summary, it was checked against the export, it was off by a rounding rule, here is the corrected version. An empty answer generalises. It talks about being careful, about always double-checking, about the tool being useful for first drafts. Careful is a disposition. A check is an event with a time on it. Hearing the difference between the two takes practice on the reviewer's side as well, which is the case for managers putting in hands-on hours of their own.

Run this on the strongest person on the team first, so the pattern being looked for is visible before it is being judged. These four questions are also what a good interviewer is reaching for when they try to see what good AI use looks like in a candidate, and the same distinction sits under the argument about measuring speed or measuring judgment.

What about the person who barely uses it?

They might be the strongest signal on the team, or a genuine holdout, and the same ten minutes tell you which. The distinction is whether the person can name the step where the model was the wrong instrument and say what it would have got wrong. That is an account. Discomfort with an unfamiliar tool produces no account of any particular step, only a preference.

Low use among strong performers is what the evidence predicts. In a rollout to 5,179 customer support agents, a conversational assistant raised issues resolved per hour by 14% on average, with a 34% improvement for novice and low-skilled agents and almost nothing for experienced, highly skilled ones 4. One firm, one occupation, one assistant trained on that firm's own transcripts, and the outcome measured was issues resolved rather than quality of judgment. The direction still holds: the people with least to gain from a given assistant are sometimes the ones who know most.

The costlier failure runs the other way. In the field experiment with Boston Consulting Group consultants, on one task chosen to sit outside what the model could do, 84.5% of the control group reached the correct answer against 60% and 70% in the two AI conditions 5. Consultants could not tell which side of the line the task sat on, and the group given a prompt-engineering overview did worse than the group given none. One task, one 2023 experiment, and the frontier has moved with every model release since.

An artifact shows both of those patterns, and a counter shows neither. No dashboard can show you the person who used the model on a step where it should not have been used. A ten-minute conversation about one deliverable can, and it works the same way on the person who used it least.

See the benchmarks

Common questions

Is there any usage metric worth watching at all?

Two are worth a glance, neither as a capability measure. Whether use is spreading past the early adopters tells you something about access and permission. Whether a team's use dropped after a policy change tells you the policy landed badly. Both are questions about conditions rather than skill. Anything framed as prompts per head invites people to produce prompts, which is the one behaviour available for free.

Should the team fill in a self-assessment of AI skill?

Use it to find out what people believe they are allowed to do, not what they can do. Self-ratings and demonstrated ability come apart, and a form asking someone to rate their own AI skill mostly measures confidence, exposure and question wording. If the goal is a training plan, ask instead for one recent deliverable and four questions about it: what the criteria were before drafting, which claim the decision rested on and what was opened to check it, what got thrown away, and how the conclusion was tested. Ten minutes of that beats a full page of survey responses.

Does heavy use ever indicate good work?

It indicates a tool is available and permitted, which is worth knowing in the first weeks of a rollout. It stops being informative almost immediately after. Two people with identical usage counts can differ completely in whether anything they produced was checked, and the one who ran fewer sessions may have run them on the steps that carried the decision.

What if the manager cannot judge the output themselves?

Then the review is theatre, and the fix is hours in the tool rather than a course. A manager needs enough time working with a model to recognize a fabricated citation and a plausible wrong number on sight, because those are exactly what a fluent deliverable hides. Until that exists, pair the review with someone who does the work daily and have them ask what the criteria were, what was checked, what got thrown away and how, while the manager listens.

How often is this worth running?

Once a quarter per person is plenty, and running it once at the start is what actually changes behaviour. The first round teaches the team what will be asked about, which is most of the value: people who expect to be asked what they threw away begin keeping track of what they threw away. Keep it away from the rating cycle, or the answers turn into performances.

References

  1. 1. Microsoft Copilot usage report, Microsoft 365 admin center Microsoft Learn, 2026. learn.microsoft.com Supports the claim that an adoption report counts a user as active on one intentional action inside the selected 7, 28, 90 or 180 day window.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that self-reported speedups can be wrong in direction, so a team's own account of its AI gains is not a measurement.
  3. 3. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the claim that a self-rating of AI skill and a demonstrated one measure different things.
  4. 4. Generative AI at Work (NBER Working Paper 31161) National Bureau of Economic Research, 2023. nber.org Supports the claim that measured gains concentrate in less experienced workers and can be near zero for experienced ones.
  5. 5. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that unexamined use fails on tasks that look like the ones a model handles well.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.