Teams
Which Roles Actually Need AI Skills Right Now?
A role needs AI skills right now only when three things hold at once: a model can already draft much of the occupation's published task list, that work runs weekly rather than yearly, and a plausible wrong answer costs money or credibility before anyone downstream catches it. In most companies that is three to six roles. A role meeting one condition gets tool access and a one-page policy. A role meeting all three needs a written bar and someone who reads the output before it ships.
The takeScore roles against those three conditions and the same uncomfortable finding keeps turning up. It rarely survives into the deck. The roles that worry you are not short of skill, they are short of a reader. Someone used to check that work, a reorganization took them out, and a course costs less than a headcount and photographs better. So the capability gap gets funded and the review gap does not. Most mandates train one side of a problem the org chart created, and the completion dashboard is why nobody notices.
Where Olive fits
Open a role and see what the work shows
The six dimensions Olive reads are a working description of what capable AI work looks like in an exposed role: framing before generating, demanding a source for the claim the decision rests on, keeping the judgment that should not be handed over, building something between the brief and the answer, refusing output with a reason, and testing a claim against something outside the conversation. Olive is employer-purchased and built for hiring, so it settles those for a candidate on an assignment written for their occupation, read by a human reviewer, rather than for people already on the team.
Rank your shortlistWhich roles actually need AI skills right now?
Any role where a model can already draft a large share of its published tasks, those tasks run weekly rather than yearly, and a wrong answer costs money or credibility before anyone downstream catches it. Three conditions, all true at once. A role that meets one of them needs tool access and a short policy. A role that meets all three needs a written bar and something that checks the work against it.
The filters, in the order that saves the most time:
- Task exposure. How much of the role's real task list can an assistant do a competent first pass on? The task list, not the job description. A financial analyst's comparable-company screen and first-draft memo are exposed; the same analyst's conversation with a lender is not.
- Frequency. An exposed task performed twice a year is a problem you solve twice a year, with a checklist. An exposed task performed forty times a week is where a capability gap compounds into a quarter of bad output.
- Cost of a wrong answer. A model produces a plausible number, a plausible clause, a plausible summary. In some roles the next person downstream catches it as a matter of course. In others it goes into a board pack, a filing, a patient record or a price.
Only the third filter is genuinely about your company, and it is the one most rollouts never write down. Two firms with identical org charts rank the same role differently because one has a review step after it and the other removed that step in a reorganization three years ago.
Run all three and the list usually comes to three to six roles, not a department and not a headcount plan. Naming what AI actually does inside each role is the input; a training catalogue is not.
Why is task exposure a fact about the occupation, not about your company?
Because exposure is measured against what an occupation does, and occupations are defined outside your walls. An exposure study covering the whole U.S. workforce found around 80% of workers could have at least 10% of their work tasks affected by large language models, while about 19% could see at least half of their tasks affected 1. Both figures describe the same labor market. The distance between them is the decision you are making.
Usage data says the same thing from the other direction. A Microsoft Research team classified 200,000 anonymized Copilot conversations against O*NET work activities and found the most common and most successful AI-assisted activities are information work (creating, processing and communicating information), which is why applicability cuts across sectors instead of concentrating in one 2. Occupations differ less in whether AI touches them at all than in which part it touches, and in whether that part is the part carrying the risk.
Within an occupation, the share of tasks involved stays modest. Anthropic's usage index found roughly 36% of occupations showed AI use across at least a quarter of their tasks, and only about 4% across at least three-quarters 3. Broad and shallow, in other words: most roles have some exposed work, very few roles are mostly exposed work, and a company-wide mandate prices every role as though it were the second kind.
In practice, a reconciliation-heavy finance role and a drafting-heavy marketing role change on different timelines, in different directions, for reasons that have nothing to do with which manager is more enthusiastic. Your company contributes exactly one thing to the calculation, and it is the cost column.
Run the triage in an afternoon
Four columns, one row per role, filled from documents you already own. Pull the occupation's task statements, mark the ones a model can draft, note how often each runs, and write what a wrong answer costs before someone catches it. O*NET publishes a task list for every SOC code (26 tasks for financial and investment analysts, downloadable as a spreadsheet 4), so the first column is a copy-and-paste rather than a workshop.
1. Map each internal title to a SOC code. Roughly is fine. Your "Senior Insights Manager" is a market research analyst or a data scientist; pick one and move on. The point of the external list is that it was not written by the person who wants the training budget. 2. Mark exposed tasks. For each task statement, ask whether a competent first draft could come out of a model given the material your team already has. Not whether it would be right. Whether it would be plausible enough to ship if nobody checked. 3. Weight by frequency. Ask the manager how many times a week the exposed tasks happen. This is the column that separates a real gap from an interesting one. 4. Write the cost sentence. One line per role: what a plausible wrong answer does before it is caught, and who catches it today. If the honest answer is "the client", the role is on the list.
| Role | Exposed tasks | How often | Cost before someone catches it | Verdict |
|---|---|---|---|---|
| Financial analyst | Comparables, first-draft memo, variance commentary | Daily | A wrong figure reaches an investment committee | On the list |
| Marketing associate | Positioning drafts, research summaries, ad copy | Daily | A fabricated statistic goes out under the brand | On the list |
| Contracts paralegal | Clause comparison, first-pass redlines, summaries | Weekly | A position taken against the signed precedent | On the list |
| Enterprise account executive | Call notes, follow-up email, proposal boilerplate | Daily | A colleague notices before the customer does | Tool access |
| Facilities manager | Vendor emails, occasional policy drafting | Monthly | Rework, caught in the ordinary approval path | Tool access |
The verdict column is the entire output. Two values, not five: a role either gets a written capability bar and a check, or it gets access and a one-page policy on what may not be pasted into a chat window. Anything more elaborate turns a triage into a program, and programs are how the original question stops being asked.
Do not let the exercise become a competency framework in the same sitting. Designing an AI competency framework is a real piece of work with its own failure modes, and it should be done for three roles that earned it rather than for every role at once.
Which roles are you adding AI skills to for no reason?
The ones where the exposed tasks are rare, cheap to correct, or not really in the job. Four patterns account for most of the waste, and each is visible in the triage table the moment you fill the frequency and cost columns. None of these roles needs a capability bar. Each needs tool access, a policy about what data may go into a chat window, and no further ceremony.
- The title-level mandate. Someone adds "AI fluency" to a job family because the family name sounds technical. Field service, IT support and implementation roles frequently sit here: the task list is diagnostic and relational, and the exposed part is the status email.
- Training the task nobody performs. The occupation's task list says report writing; on your team, a shared template handles it and the person spends their week on customer calls. Exposure at the occupation level is a starting point, and the frequency column is what corrects it for your company.
- Counting usage as capability. Seat licences, prompt counts and completion rates measure activity. Someone who ran forty prompts and shipped every answer untouched used more AI than a colleague who ran four and threw out three, and the second person is doing the job you are trying to hire for. The relevant split is between work a model augments and work it does outright 3; only the second kind moves a capability bar.
- The enthusiasm ranking. The roles that end up in an AI program are often the roles whose leaders asked first. That is a fact about your leadership meeting, not about the work, and it is the single most common way the short list goes wrong.
The uncomfortable case is a role with high exposure, high frequency and a wrong answer nobody catches, because the honest reading is not that the role needs training. It needs a review step it does not have. Training a person to check their own output is worth doing; it is not a substitute for the second pair of eyes that used to exist, and treating training as sufficient without verifying the work is how a mandate produces a green dashboard and worse output.
What changes once the short list is set?
Three things, in order: the bar, the check, and the job requirements. Write what good looks like for each listed role in observable terms, decide role by role whether to train the capability or hire it, and put a check after the training so completion is not the only evidence you own. A short list is worth having only if something downstream changes because a role is on it.
- Write the bar in acts, not adjectives. "Demanded a source for the revenue claim and opened it" is observable. "Shows strong AI judgment" is a memory of how a conversation felt. Setting a defensible bar for a specific role is a half-day per role and it is what makes every later decision arguable in a good way.
- Decide train or hire per role, not per company. A role with three exposed tasks and a patient manager is a training problem. A role where the exposed work is most of the job and the market has people already doing it is a hiring decision with its own arithmetic.
- Check the work afterwards. Course completion is an attendance record. A short task on your own material, with the assistant available and a rubric written beforehand, is the only thing that tells you whether exposure became capability.
- Change the requisition last. Once the bar exists, writing AI skills into the job requirements is a small edit. Doing it first produces a requirement no interviewer can assess and every candidate can claim.
Revisit the table when the work changes, not on a calendar. A new tool in the finance stack, a review step removed, a team that absorbs a function. Each of those moves a row. Nothing else does, and an annual refresh of a document nobody consults is the thing this exercise was meant to avoid.
Common questions
How is this different from asking managers which teams want AI training?
It starts outside the building. A manager's request measures interest and confidence, which correlate with neither exposure nor risk. The occupation's task list was written by someone with no stake in your budget, so it gives you a first column nobody in the room can argue with. Manager input still matters, but for one thing only: how often the exposed tasks actually happen on that team. Ask that question specifically, rather than asking whether the team wants training, and the answers stop being a popularity contest.
Does every role on the short list need the same training?
No, and a shared curriculum is the most common way a good short list produces nothing. The acts that matter are the same everywhere (framing before generating, demanding a source, keeping the judgment you should not hand over, testing a claim against something outside the conversation), but the object differs by occupation. An analyst ties a number back to the filing. A paralegal pulls the precedent. A marketer opens the survey the statistic came from. Teach the act against that role's own material, or it will not transfer.
What about roles that will be exposed in two years but aren't now?
Leave them off and write down why. A role you list early spends budget on a capability the work does not yet reward, and the training expires before the exposure arrives. What is worth doing now is naming the trigger: the tool, the integration, or the removed review step that would move the role onto the list. Then the next revision takes ten minutes rather than restarting the exercise. A triage that tries to be a forecast stops being useful as a triage.
Should the short list drive hiring requirements too?
Yes, but after the bar exists, not before. A requisition that says "AI proficient" without a written standard produces a claim every candidate makes and no interviewer can check. Once a role has an observable bar, the requirement writes itself from it, and the interview round has something to assess against. Roles that stayed off the short list should not gain an AI requirement at all. Adding one narrows the applicant pool for a capability the job does not exercise.
How many roles should end up on the list?
In most companies under a thousand people, three to six. If the list has twenty roles on it, the frequency and cost columns were not filled in honestly, or exposure was marked at the job-family level rather than the task level. If the list is empty, check whether anyone applied the third filter: nearly every organization has at least one role where a plausible wrong answer reaches a customer, a regulator or a board with no intermediate reader.
Can the same assessment be used across every role on the list?
The structure travels; the material does not. What is being observed (framing, evidence, delegation, refusal, verification) holds across occupations, which is why one rubric shape can cover several roles. The task cannot be shared, because a plausible wrong answer in finance looks nothing like a plausible wrong answer in legal operations, and a generic exercise grades neither. Expect one common frame and a different assignment per occupation family.
References
- 1. GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models ✓ arxiv.org Around 80% of the U.S. workforce could have at least 10% of their work tasks affected by LLMs, and approximately 19% of workers may see at least 50% of their tasks impacted; exposure is scored at the occupation level.
- 2. Working with AI: Measuring the Applicability of Generative AI to Occupations ✓ arxiv.org 200,000 anonymized Copilot conversations classified against O*NET work activities: the most common and successful AI-assisted activities are information work, and applicability is widespread across sectors because most occupations have information-work components.
- 3. The Anthropic Economic Index ✓ anthropic.com Roughly 36% of occupations showed AI use for at least 25% of their tasks, only about 4% for at least 75% of tasks, and usage split 57% augmentation to 43% automation.
- 4. Financial and Investment Analysts (13-2051.00) ✓ onetonline.org The occupation carries 26 published task statements, downloadable as a spreadsheet, which is the unit a task-exposure triage is run against. Accessed 24 August 2026.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.