Teams
The AI Skills Gap Is a Specification Problem First
An AI skills gap is probably a specification problem first, and it is cheap to find out. Write the skill as observable behavior for two or three real roles, then read recent work against it. Most teams find a distribution rather than a gap: a few people already working that way, unnoticed because nobody had named it. Whatever is left after that is a genuine shortage, and it is small enough to hire or train against deliberately.
The takeA company-wide programme is the most expensive available answer to a question nobody has asked precisely. It is also the easiest to approve, because it converts an uncomfortable specification problem into a procurement decision with a vendor attached and a date on it. That is the real appeal, and it is worth saying out loud before the budget is signed. A gap you have not defined is not a gap you have measured; it is a discomfort you have funded.
Where Olive fits
Open a role and see what the work shows
Olive's six dimensions name the same behavior a specification is reaching for: how the problem gets framed, which claims get a source, what stays with the person, and what gets checked against something outside the conversation. Olive reads those from one role-grounded session per person and returns six findings, each with the timestamped moment it rests on, as an input to a decision a person still makes.
Rank your shortlistWhere does the gap number come from?
Almost always from a survey asking managers or employees to rate a capability that the survey never defined. The respondent supplies their own definition, which drifts with how much AI coverage they read that month, and the aggregate is a mood rather than a measurement. That is not a reason to dismiss it. A widespread feeling that the organisation is behind is worth investigating, but it is where the work starts.
The tell is that the same question produces incompatible answers depending on who is counted. A Federal Reserve note comparing the three main US adoption surveys found they disagree by construction: about 18% of firms had adopted AI at the end of 2025, about 41% of the workforce reported using generative AI at work in November 2025, and an employment-weighted estimate put 78% of the labor force at firms that have adopted AI 1. Those are three different units of analysis and none of them converts into another. The 78% describes whose employer has adopted AI, which is a different question from who uses it.
So the first question to ask about any gap figure is what it counted. A firm-level number and a worker-level number for the same period can differ by twenty points and both be correct, because most firms are small and most workers are not at small firms. The Census Bureau's biweekly survey put AI use among US businesses at 19.8% in May 2026, against 37% of firms with at least 250 employees 2.
None of that answers whether your company has a gap. It establishes something more useful: a published gap figure is not evidence about your team, and there is no external number you can substitute for looking.
Why an average is the wrong shape for this
Because it collapses the distribution, and the distribution is the finding. Inside a company the same ambiguity that splits the published surveys operates on your own numbers: an internal survey and a manager's estimate of one team can differ by half the headcount without either party being careless. An average across that spread describes nobody on the team.
One study measured that spread directly. In the staggered rollout of an AI assistant to 5,179 customer support agents at one software firm, issues resolved per hour rose 14% on average, with a 34% improvement for novice and low-skilled agents and almost nothing for the experienced ones 3. One firm, one occupation, a pre-ChatGPT-era assistant on problems with correct answers: the numbers do not transfer to your team. The shape does. An average movement of 14% concealed two groups moving completely differently, and a programme designed against that average would have reached neither.
Aggregate effects are also smaller than the narrative suggests, which is worth knowing before the budget conversation. In a nationally representative US survey, people using generative AI at work reported mean time savings of 5.4% of their work hours, which works out to 1.4% across all workers including non-users 4. Self-reported, late 2024, and time saved is not value created unless the freed hours go somewhere. It is the honest counterweight to a deck implying the whole company is a step behind.
What survives every caveat here is a method. Count something specific, in your own company, against a definition you wrote down first.
Define the behavior for two or three roles first
Pick the roles where AI work is already happening and write, in a paragraph each, what good looks like in that job specifically. Not a competency framework and not a policy: four or five sentences naming decisions someone could be seen making. What goes to the assistant, what context it gets, what gets checked and against what, what stays with the person. Two weeks of calendar time, mostly other people's.
The drafting has to be done with the people doing the job, and it goes wrong the same way every time when it is not. A definition written in a leadership offsite describes work as leadership imagines it, so the resulting gap is between the team and an imagined role. The correction is cheap: two practitioners per role, an hour each, asked what they actually hand over and what they would never hand over. Their disagreements are the specification.
Then sample the work itself. Take five recent pieces of output per role and read them against the paragraph. Was anything traced to a source. Did anyone keep back part of the task and say why. This is slower than sending a form, and it is the only step in the exercise that produces evidence at all. Self-report and observed work are known to diverge in both directions.
Doing this without a survey nobody answers honestly is its own small problem, worked through in how to find out what the team can already do with AI. If the paragraphs are hard to start, four levels of AI skill written as observable behavior gives a ladder to borrow and edit.
What to do with whatever gap is left
Sort what you found into three piles. People already working the way the definition describes, who need naming and a role in teaching the rest. People a fortnight of supervised practice would move, which is a training item with a date. And capability nobody has at all, the only genuine hiring question. That third pile is usually the smallest, and it is the only one the gap figure claimed to measure.
The first pile is worth dwelling on, because it is the one specification reveals and surveys hide. Teams routinely have two or three people whose verification habits are exactly what the company is about to pay a vendor to install, and nobody had noticed because the behavior had no name and no line in a review. Naming it costs nothing and does most of the work: the practice spreads by imitation once it is visible and legible, and the internal teacher is more credible than the external course.
The second pile sets the size of any programme you buy, which is the number the whole exercise exists to produce. A programme sized to the second pile is a fraction of one sized to the whole company, and it is aimed at people who have a specific thing to learn rather than at everyone.
The third pile is a hiring decision with a real bar attached, and by this point the requirement writes itself out of the definition you already drafted. Whether to fill it by hiring or by training is the standing question, and whether to hire for AI skills or train the team you have takes it up. Which roles should have been in the exercise at all is answered in which roles actually need AI skills right now.
Common questions
How long does the specification exercise take?
About two weeks of calendar time and a few hours of anyone's actual attention. Two practitioners per role for an hour each, an afternoon to draft the paragraphs, a week for people to react, and a couple of hours to read a sample of real work against the result. The main cost is scheduling. If it is taking a quarter, it has turned into a competency framework project, which is a different and much larger thing.
What if leadership already committed to a company-wide programme?
Run the specification anyway and let it size the rollout. The exercise does not have to cancel a programme to be worth doing; it decides who goes first, what the material covers and what counts as having landed. A programme aimed at a definition drawn from your own roles is a better version of the same purchase, and the evidence it produces is what any later review of the spend will ask for.
Is a skills gap ever the right diagnosis?
Sometimes, and the real one is narrower than the framing suggests. Genuine shortages show up in specific work rather than across a company: nobody who can evaluate a model's output in a regulated domain, nobody senior enough to decide what should not be delegated in a new product area. Those are real and they are hiring problems. The company-wide version almost never survives contact with a written definition.
Should the assessment cover managers too?
Especially managers, and they are the group most often skipped. A manager who cannot tell whether a report's work was checked will accept polished output and reward it, which teaches the team that polish is the standard. That failure spreads faster than any individual's, because it operates on everyone the manager reviews. Sample their judgments about other people's work rather than their own use of an assistant.
How do you know the specification is any good?
Hand it to two managers with the same three pieces of work and see whether they reach the same conclusion. Agreement means the paragraph names something observable. Disagreement means it names an attitude, and the fix is in the wording rather than in the raters. Running that check once, before anything is measured or bought, is what separates a specification from a mission statement.
References
- 1. Monitoring AI Adoption in the US Economy (FEDS Notes) federalreserve.gov Supports the claim that the main US adoption surveys disagree by construction because they count different units.
- 2. Large Firms With at Least 20 Employees Biggest AI Users census.gov Supports the firm-level adoption rate and the size gradient between large and small employers.
- 3. Generative AI at Work (NBER Working Paper 31161) nber.org Supports the claim that an average effect concealed two groups moving very differently, which is the distribution a gap figure destroys.
- 4. The Rapid Adoption of Generative AI (NBER Working Paper 32966) nber.org Supports the self-reported time-savings figures used as a counterweight to claims that AI has already transformed output.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.