Teams

Four Levels of AI Skill, Written So Two Managers Agree

Levels of AI skill hold up only when each one names a behavior under a real task, not a state of mind. At rung one, a person hands the task to an assistant and accepts what comes back. At two, they supply context and edit the result. At three, they check the output against something outside the model and catch an error. At four, they decide what should not be delegated at all. Three and four are the rungs worth hiring against, and neither shows up in a self-rating.

The takeMost competency ladders are built to be filled in, which is why they end up measuring enthusiasm. Someone who uses an assistant constantly and checks nothing will place themselves high on every published scale, and their manager, who has seen only the finished output, will often agree. The ladder that would catch this asks a narrower question: what did they verify, and what did they decline to hand over. Build that one, even if it fits on half a page and grades fewer people than you hoped.

Where Olive fits

Open a role and see what the work shows

The same behaviors describe capable AI work on a team: framing before generating, demanding a source for the claim that matters, keeping the judgment that should not be handed over, and testing an answer against something outside the conversation. Olive reads those from one working session rather than from a self-assessment.

Rank your shortlist

Why the usual four levels stop separating people

Because they grade attitude and tool count rather than behavior. Aware, applied, integrated, strategic: none of those names a thing a person could be seen failing to do, so the placement falls back on how often someone talks about AI and how many products they have open. Everyone using an assistant daily lands in the middle, the distribution collapses, and the scale stops working exactly where you needed it.

The frameworks people borrow from are usually honest about this, and the borrowing is where it goes wrong. The AI Fluency Framework, in its current version, sets three named sub-competencies under each of its four competencies, twelve in all, against three modalities of human-AI interaction 1. It is a taxonomy: no levels, no scoring anchors, no proficiency thresholds published with it. Turning a taxonomy into a five-point scale by adding adverbs produces something that looks like a rubric and behaves like a survey.

A subtler failure sits underneath. A ladder written in adjectives gets applied to the person, so it absorbs everything a manager already believes about them. A better adjective will not fix that. Move the unit of judgment from the person to a task, and the question becomes what happened in this piece of work.

Accept one consequence up front: a behavioral ladder places fewer people confidently than an attitude ladder does, because half your team has never been observed doing the thing. That is not a defect in the ladder. It is the information the old one was hiding.

What each rung looks like under a real task

Take one piece of work someone actually did with an assistant and ask four questions in order. Did they hand it over and accept what came back. Did they supply context and edit the result. Did they check a claim against something outside the model and catch something. Did they decide part of the task should not be delegated at all. The highest question answered yes is the rung.

Spelled out, with the failure that defines each boundary:

  • One, accepts. The task goes over intact, the output comes back, it ships roughly as written. Fails at the first factual error nobody looked for.
  • Two, edits. Context supplied up front, output reworked before it goes out. Better prose, same exposure: editing improves what the reader sees without testing whether it is true.
  • Three, checks. At least one claim gets traced to something outside the conversation, and at some point that check catches an error. This is the first rung with evidence attached rather than effort.
  • Four, refuses. Some part of the work is kept back on purpose, with a reason: the judgment belongs to a person, the source material cannot go into a prompt, the model has no way to know what matters here.

Two details make this hold up in practice. The rungs are cumulative, so a person at four still does the checking at three. And the boundary between two and three is the one that matters most and gets blurred most, because heavy editing feels like verification and is not. If nobody left the conversation to confirm anything, the work is at rung two however polished it reads.

Check that two managers place the same person on the same rung

Run the ladder twice on the same work before anyone's level goes in a system. Give two managers the same three pieces of output, without names, and ask each to place them. Where the placements differ, the disagreement is almost always about rung three, and the fix is to write down what counts as a check in your work rather than adding another adjective to the definition.

The research on this is thinner than anyone would like, and it points one way. A team building an AI literacy assessment for a US Navy robotics programme reported that a scenario task simulating AI use on the job outperformed the knowledge tests they had adopted from prior work or written themselves, and argued that prevailing assessments overweight technical foundations against practical judgment 2. One programme, one context, a preprint with no effect size in its abstract: it supports a design choice rather than a benchmark. The design choice is to anchor the rung to a task rather than to a description.

What calibration usually surfaces is not a scoring disagreement but a missing definition. One manager counts "asked a colleague" as a check; another requires a document. Both are defensible and the pair has to pick one, in writing, per role. That single sentence does more for consistency than the rest of the framework combined, and it is the piece most competency models never write down.

A rubric two reviewers apply the same way is its own problem with its own mechanics, worked through in how to write a rubric for AI use that two reviewers score the same way.

Should a self-rating ever set the level?

Use it to find out what people think, never to place them. Self-report and demonstrated ability come apart in the measured cases, and they come apart in both directions, so a self-rating is not even reliably inflated. It is a different quantity, useful for spotting where confidence and evidence disagree, and unusable as the input to a level that affects pay, promotion or a hiring decision.

The cleanest evidence sits outside the workplace. Among 288 Taiwanese teachers who took both a self-report and a knowledge-based test of AI literacy built on the same framework, correlations between the objective and self-reported factors ran from 0.07 to 0.24, and the analysis found six distinct profiles including both overestimators and underestimators 3. Teachers in Taiwan are not your engineering team, and a weak correlation means the two instruments measure different things, which is not the same as saying people flatter themselves. Both readings support the same operational rule: the two numbers are not interchangeable.

Belief about one's own AI use is unreliable even among experts working in their own domain. In a randomized trial, sixteen experienced open-source developers working on repositories they had maintained for years forecast that AI tools would cut their completion time by 24%, estimated afterwards that it had cut it by 20%, and were measured 19% slower 4. Small sample, one setting, and the magnitude does not transfer. What transfers is that the estimate was wrong in direction, not only in size.

So run the survey if it is useful for planning, and keep it out of the ladder. Where the two disagree is where to look first, and what the minimum every employee needs looks like sets the floor the whole ladder sits on. Whether AI should be a competency of its own at all is a prior question, handled in should AI be its own competency or folded into the ones you have.

See the benchmarks

Common questions

How many levels should a competency model have?

Four is enough for AI work and five is usually one adjective too many. The number is not the design decision; the boundary is. Each gap between rungs has to name a specific behavior that appears on one side and not the other, and there are three such boundaries worth drawing: accepting versus editing, editing versus checking, checking versus refusing. Add a fifth rung only when you can say what someone does there that a four cannot.

Can someone be at different levels in different tasks?

Routinely, and a model that cannot express it will produce arguments. People check carefully in the domain they know and accept freely in the one they do not, which is expected rather than inconsistent. Place the rung per task family rather than per person: analysis, client writing, code, research. The pattern across those placements is more useful than any single number, and it tells you where to put the training.

What does rung four look like for a junior?

Smaller in scope and still real. A junior at rung four does not redesign the process; they hold back one piece of a task and say why, usually because the source material is sensitive or because the call is the manager's to make. The refusal is the same behavior at a different altitude. Expecting juniors to be incapable of it is how teams end up hiring for output speed and being surprised later.

Do levels belong in a performance review?

Only with the same care any behavioral rating gets, and never as a standalone score. A rung is a description of what happened in a piece of work, which is useful evidence in a conversation and dangerous as a number attached to a person. If it feeds compensation, expect the ladder to be gamed within one cycle: people will perform the visible parts of checking without doing the checking.

How do you place someone who barely uses AI at all?

Off the ladder rather than at rung one. Rung one describes handing work over and accepting the result, which is a behavior with a failure mode. Not using an assistant is not a rung; it may be a deliberate and correct choice for the work, and grading it as the bottom of a scale quietly makes usage the thing being measured. Record it as unobserved and decide separately whether the role needs it.

References

  1. 1. Framework for AI Fluency Ringling College of Art and Design (Rick Dakan and Joseph Feller), 2025. ringling.libguides.com Supports the claim that the widely borrowed fluency framework is a taxonomy with no levels, scoring anchors or proficiency thresholds attached.
  2. 2. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arXiv (Bogart, Warrier, Agarwal, Higashi, Zhang, Flot, Savelka, Burte, Sakr), 2025. arxiv.org Supports the design argument that a scenario task measured applied AI literacy better than the knowledge tests the same researchers used.
  3. 3. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the claim that self-reported and demonstrated AI literacy correlate weakly and diverge in both directions.
  4. 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that people can be wrong about the direction of AI's effect on their own work, not only about its size.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.