Teams

Should AI Be Its Own Competency in the Framework?

Give AI its own competency only where it changed the deliverable; where it changed just the method, fold it into the competency you already have and rewrite that competency's behavioral anchors. The test runs one competency at a time, on what AI changed there: the method, or the deliverable. If what gets produced, approved or shipped is now a different object, it needs its own row, because no existing definition was written about it. Most companies end up with both, split function by function.

The takeThe framework question is standing in for a harder one: whether the thing people get promoted for is still the thing they do. Adding a row is the cheap move. It costs one approval and concedes nothing. Rewriting the anchors under a competency you have carried for years means saying out loud that the old ones described work that has moved, and that some strong ratings were given against them. Nobody has counted how that choice splits across companies, but the incentive points one way: the new row gets approved and the old anchors stay standing. That is the tell. The framework changed and no rating did.

Where Olive fits

Open a role and see what the work shows

The six dimensions Olive reads are the behavioral anchors a competency framework has to write for itself: framing before generating, demanding a source for the claim that matters, keeping the judgment that should not be handed over, building something between the brief and the answer, refusing output on substance, and testing a claim against something outside the conversation. Olive reads those from one recorded occupational session rather than from a self-assessment or a manager's recollection.

Rank your shortlist

What is the split test, and how do you run it?

One question, asked of one competency at a time: did AI change the method, or did it change the deliverable? If the output a person is accountable for is the same document, decision or build, and only the route to it moved, the competency you already have is the right home. Rewrite its anchors. If what gets produced or approved is now a different thing, no existing definition covers it, and it needs its own row.

Run it with the job description in front of you rather than the framework:

1. Name the deliverable the person is accountable for today. Not the task list: the object that leaves their hands and that somebody else acts on. 2. Name what that object was two years ago. Same object by a different route, or a different object? 3. Same object: stay inside the existing competency. Rewrite its behavioral anchors so they describe how the work is done now, and leave the competency name alone. 4. Different object: write a new competency. None of the definitions you have was written about it, and stretching one to fit produces an anchor nobody can rate.

Both blanket answers fail, in opposite directions. Bolting one company-wide AI competency onto every scorecard produces a row that reads "uses AI tools effectively", which is unratable by a manager, undisputable by an employee, and quietly ignored by the second review cycle. Folding it in everywhere hides the two or three functions where the job genuinely changed, because the old anchors still look satisfiable and nobody is forced to rewrite them.

The test needs one input the framework does not contain: which roles the change actually reached. Which roles need AI skills is the list to settle first, because a framework decision taken before it exists is a guess applied uniformly to people the guess does not fit.

Rewrite one competency both ways

Take Analysis and Problem Solving, which nearly every framework already carries. Below it is written twice on the same job (a financial analyst producing an investment memo), first folded in, then standalone. The competency name is identical in both. What differs is where the AI behavior is described, and therefore what happens to the rest of the framework.

Folded in (Analysis and Problem Solving, anchors revised):

LevelAnchor
DevelopingStates what the analysis has to settle, and what would make an answer wrong, before producing anything, including before the first prompt.
ProficientDemands a source for the claim the recommendation rests on, and opens it. Distinguishes a source that was retrieved from one that was described.
AdvancedRe-derives the load-bearing figure outside the tool that produced it, and changes the recommendation when it does not tie out.

Standalone (Analysis and Problem Solving untouched, plus a new row):

CompetencyAnchor
Working with AI · DevelopingFrames the problem in own terms before generating; names a constraint in the same breath.
Working with AI · ProficientKeeps a defined piece of the work by hand, and refuses model output on substance with the reason stated.
Working with AI · AdvancedTests a claim against something outside the conversation, and the result changes a number, a recommendation or a stated limit.

For this analyst, the folded-in version wins. The memo is still the memo, the reader still acts on it the same way, and the anchors now describe how it gets made. The standalone version creates a second place to write "checked the number," and a manager rating both rows will double-count one act or split it arbitrarily.

What moved for that analyst is the method, and the direction of the move is measurable. A survey of 319 knowledge workers describing 936 real tasks found effort shifting from gathering information toward verifying it, from problem-solving toward integrating a partial response, and from execution toward stewarding a task the model is doing part of 1. Verification, integration and stewardship are not new competencies. They are the old ones with the weight redistributed, which is exactly what an anchor rewrite is for.

Now change the job. A support lead who no longer writes replies but approves drafted ones has a different object leaving their hands: not a reply, but a decision about somebody else's reasoning, taken sixty times an hour. Analysis and Problem Solving was never written about that, and no anchor rewrite reaches it. That role gets the standalone row, and the standalone row is the one that survives review, because you can point at the act it names. What good AI use looks like in practice is the vocabulary both versions are drawing on.

Why does the answer differ function by function?

Because AI reaches different fractions of different jobs, and the fraction decides the verdict. Anthropic's index of roughly a million assistant conversations mapped to occupational tasks found about 36% of jobs showing AI use across at least a quarter of their tasks, but only around 4% across at least three-quarters 2. A framework decision taken once, at company level, is applied to both ends of that distribution.

The same index splits usage between augmentation, where the person stays in the loop, and automation, where the task is handed over: roughly 57% against 43% 2. That split is the split test in aggregate. Augmentation is a method change and belongs inside existing competencies; automation moves the deliverable and is what a new competency is for. Read the figures as what people brought to one assistant rather than as what your firm does. The mapping measures conversations, not staffed work.

Even inside one function the change is uneven. In a study of 5,179 customer support agents, access to a conversational assistant raised issues resolved per hour by 14% on average, 34% among novice and low-skilled agents, and close to nothing among the most experienced 3. A single framework row rated on a single scale is being asked to describe a capability that moved four times as much for one cohort as for another.

So run the test per function and expect a mixed answer:

FunctionWhat usually movedWhere it goes
Financial analysisMethod: the memo is still the memoFolded into the existing competency
Customer support, tier oneDeliverable: approving a draft, not writing a replyIts own competency
Software engineeringContested: depends whether reviewing generated code became the primary actTest it per level, not per function
Marketing contentDeliverable at volume: the brief became the artifactIts own competency
Legal operationsBoth: method for research, deliverable for first-pass reviewSplit by task, not by title

A framework that forces one answer across that table will be wrong in half the organization, and wrong in the direction nobody notices: the folded-in roles look fine either way, and the changed-deliverable roles get a row that describes a job they stopped doing. What the AI actually does in the role is the evidence this table has to be rebuilt from, function by function, with your own job descriptions.

Where does an AI competency sit in the framework?

In a tiered model, position settles most of the argument. The Department of Labor's Building Blocks Model runs six tiers: personal effectiveness, academic, workplace, industry-wide technical, industry-sector technical, and occupation-specific or managerial 4. Method changes belong at tier three, where workplace competencies apply to everyone. Deliverable changes belong at tiers five and six, where competencies are allowed to differ by occupation and are expected to.

That placement resolves the objection most often raised against a standalone row, which is that it will be generic. A tier-three competency is generic by design, and "uses an assistant within the disclosure rules and states when output was model-produced" is a perfectly good tier-three anchor. The failure is writing an occupation-specific behavior at tier three, where it applies to people it does not fit, or writing a generic one at tier six and calling it rigorous.

Three tests before a proposed AI competency goes into the framework:

  • Can you write an anchor a manager could observe on an ordinary Tuesday? "Demanded a source for the revenue claim and opened it" is observable. "Uses AI tools effectively" is a value statement wearing a competency's clothes.
  • Can anyone fail it? A row nobody can be rated below on is decoration, and it costs a real slot in a review conversation that is already too short.
  • Would deleting it change any rating? If the same person receives the same rating with and without the row, it duplicates something already in the framework, and the honest fix is the anchor rewrite rather than the new row.

The wording carries into hiring whether you intend it to or not, because the competency framework is where job requirements are written from. Writing AI skills into job requirements is the downstream surface, and a vague internal anchor becomes a vague external requirement roughly one quarter later.

What breaks when the framework decides pay or promotion?

The competency stops being a description and becomes a selection procedure. The moment a rating moves someone (a promotion, a band, a reassignment onto an AI-heavy team), it carries the same expectations as a hiring test: job-related, applied consistently, and reviewed if the pattern of who passes turns out uneven 5. That is an argument for observable anchors, not an argument against the row.

Three things follow, and none of them is expensive at design time.

  • A generic anchor is the exposure. "Uses AI effectively" cannot be shown to be job-related, because nobody can state what it required. The specific anchor is both the fairer instrument and the defensible one.
  • A new competency has no rating history. The first cycle produces numbers with nothing to compare them against. Announce that cycle as a diagnostic before it runs, and keep it out of pay until there is a distribution to read.
  • Self-assessment is the weakest possible input, and it is the default one for a brand-new row. Sixteen experienced developers working real issues in their own repositories took 19% longer with AI tools allowed and believed afterwards they had been 20% faster 6. People misread their own AI-assisted throughput by that margin in work they know intimately.

Then the limit worth stating plainly, because it is the one a framework cannot fix. A competency framework says what good looks like. It does not observe anybody doing it. The rating still comes from a manager reconstructing work they mostly did not watch, on a behavior (what was refused, what was checked, what was kept) that is visible only while it happens and invisible in the finished deliverable. That gap is why a defensible AI proficiency bar has to name evidence rather than adjectives, and why assessment scores and actual performance part company when the instrument grades output instead of acts.

So: run the split test per competency, write anchors somebody could be marked down on, place them by tier, and decide separately, and later, whether any of it touches money.

See the benchmarks

Common questions

Can one AI competency cover the whole company?

At the workplace tier, yes, and it should be short: disclosure, data handling, and stating when output was model-produced. Below that tier it cannot. A single company-wide row has to be written loosely enough to fit an underwriter and a designer, and loose anchors are the ones managers skip. The usual shape that survives is one generic workplace-tier competency plus occupation-specific anchors inside the technical competencies each function already has.

How many rating levels should a new AI competency have?

The same number as every other competency in your framework. Inventing a separate five-point scale for one row means two calibration conversations, two distributions to explain, and a permanent argument about how the scales map. If the existing levels genuinely cannot express the behavior, the problem is the anchors rather than the scale, and rewriting the anchors is a smaller change than introducing a second measurement system into a review process.

The framework was refreshed last year. Is reopening it worth it?

Usually not the whole framework. Anchors are the cheap edit: they live below the competency names, they rarely require re-approval, and rewriting three of them inside an existing competency is a normal maintenance change. Reserve the expensive move (a new row, new calibration, new distribution) for the functions where the split test says the deliverable itself changed. That is typically a minority of functions, which is what makes the anchor-first sequence affordable.

Should AI appear in the leadership competencies too?

Yes, and as decisions rather than tool use. A manager's observable acts are different: setting the verification standard for work the team ships, deciding which judgments the team may not delegate, and rating other people's AI-assisted work without treating polish as evidence. Writing "uses AI tools" into a leadership competency measures the wrong person's behavior. Writing "states what the team must verify before shipping" measures something a peer could confirm or dispute.

How do you rate a competency nobody has been rated on before?

Evidence first, distribution later. Ask each rater for one named instance behind the rating: a moment where a source was demanded, a direction refused, a figure re-derived. Ratings supported by no instance are a signal about the anchor, not about the person. Run the first cycle as a diagnostic, publish what the anchors turned out to mean in practice, and let the second cycle be the one with consequences attached.

Where does an assessment fit alongside the framework?

A framework defines the behavior; an assessment observes it. Olive is employer-purchased and built for hiring rather than internal review, so it answers the question one step earlier: whether a candidate already works this way. The session is a 40-to-60-minute occupational assignment with an AI assistant available, and a human reviewer writes six findings, each attached to a timestamped moment. Outcomes are demonstrated, partly demonstrated or not demonstrated, never a number, and the candidate receives the identical report.

References

  1. 1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks and Wilson (Microsoft Research and Carnegie Mellon, CHI 2025), 2025. microsoft.com 319 knowledge workers describing 936 real tasks: effort shifted from information gathering toward verification, from problem-solving toward response integration, and from execution toward task stewardship.
  2. 2. The Anthropic Economic Index Anthropic, 2025. anthropic.com Roughly a million conversations mapped to occupational tasks: about 36% of jobs showed AI use across at least 25% of tasks and around 4% across at least 75%, with usage split about 57% augmentation to 43% automation.
  3. 3. Generative AI at Work (NBER Working Paper 31161) Brynjolfsson, Li and Raymond, National Bureau of Economic Research, 2023. nber.org 5,179 customer support agents: assistant access raised issues resolved per hour by 14% on average and 34% among novice and low-skilled workers, with minimal effect on the most experienced.
  4. 4. Building Blocks Competency Model Competency Model Clearinghouse, U.S. Department of Labor Employment and Training Administration, 2026. careeronestop.org Six tiers: personal effectiveness, academic, workplace, industry-wide technical, industry-sector technical, and occupation-specific or management competencies.
  5. 5. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2024. eeoc.gov Any step used to decide who advances is a selection procedure: job-related, applied consistently, and reviewed for adverse impact.
  6. 6. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers across 246 randomized real issues took 19% longer with AI tools allowed and believed afterwards they had been 20% faster.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.