Teams
How Many People Do You Need If AI Takes Part of the Workload?
How many people you need next year, if AI takes part of the workload, comes out as a band per occupational family. Compute it from hours on tasks an assistant can draft, times a gain measured on your own work, minus the review time drafts add. Two things cancel a saving: where a backlog rations the work, a faster team clears it and hires anyway, and a gain at the level you were about to hire into changes who you hire, not how many. Unmeasured, the gain is zero.
The takeThe two ways this plan can be wrong do not cost the same. Hire a few too many and you carry the mistake for a year. Close the entry level and you meet it years later, when the senior bench you did not train is the bench you are trying to hire from. The payroll data cited here shows the thinning has already started in the most exposed occupations, and on what is public so far that is a budgeting decision dressed up as a technology effect. If a number in your plan has to be wrong, make it the one a single budget cycle can undo.
Where Olive fits
Open a role and see what the work shows
A headcount range is an assumption about throughput, and throughput with an assistant depends on whether a person frames the problem before generating, demands a source for the claim the decision rests on, keeps the judgment that should not be handed over, and tests a claim against something outside the conversation. Olive reads those six things from an occupational assignment a human reviewer writes up finding by finding, which is evidence about one candidate you are about to hire rather than a productivity multiplier for the plan.
Rank your shortlistHow many people do you actually need next year?
Fewer in one or two occupational families, the same in most, and more in at least one. The company-wide answer is unusable because none of its inputs are company-wide. Run the arithmetic per family: hours spent on tasks an assistant can already draft, times a gain you have measured on your own work, minus the review time those drafts create. What comes out is a band, not a number.
The plan itself is old arithmetic. Headcount equals forecast work volume divided by throughput per person. AI moves one term, unevenly, and only where it lands:
- Exposed hours, not exposed tasks. A task an assistant drafts well in four minutes matters to the plan in proportion to the hours it currently eats. Most task lists carry two or three lines worth half the week and a dozen worth an afternoon between them.
- Realized gain on those hours. The figure from your own work, not from a vendor deck and not from the pilot team's enthusiasm. Until it has been measured, the honest entry is zero.
- Added review time. Every draft that ships needs someone to check the claim it rests on. That work is real, it lands on your most senior people, and it is missing from nearly every plan showing a saving.
Two further terms decide whether a throughput gain becomes a headcount change at all. Demand is the first: if the work is rationed by a backlog rather than by customer volume, a faster team clears the backlog and hires on the same curve. Skill mix is the second: the gain often appears at exactly the level you were about to hire into, which changes who you hire rather than how many.
Start from what AI actually does inside each role, then from the roles where exposure is real. A plan built on titles is a plan built on a job architecture that was written to set pay bands.
Why does one company-wide productivity number fail?
Because it is wrong in both directions at once. Exposure varies several-fold across occupations, so a blanket cut over-prunes the families where a model already drafts most of the work and leaves untouched the families where it drafts almost none. The same assumption then under-invests where the gain is real. One number, two opposite errors, and no way to see either from the top line.
An exposure study covering the whole U.S. workforce put around 80% of workers at 10% or more of their tasks affected by large language models, and about 19% of workers at half or more of their tasks 1. Both figures describe the same labor market. The distance between them is the spread your plan is pretending does not exist.
Depth follows the same shape. Usage data across occupations found roughly 36% showing AI use for at least a quarter of their tasks and only about 4% for at least three-quarters 2. Broad and shallow: most families have some exposed work, very few are mostly exposed work, and a company-wide mandate prices every family as though it were the second kind.
There is an accounting reason the average lies, too. A company-wide figure is dominated by the largest family, so the number you compute describes your biggest team and is then applied to a controller group, a field operations group and a research group whose task mixes have almost nothing in common with it. Two companies with identical org charts land on different answers because one has a review step after the exposed work and the other removed it in a reorganization three years ago.
Measure the realized gain before you spend it
The gain in your plan has to come from your own work, because the published measurements disagree by sign. Controlled studies have found large gains for new workers, close to nothing for experienced ones, and a measured slowdown for experts working in code they already knew. None of those is a company-wide multiplier. A number you did not measure is a preference with a decimal point on it.
The two most useful measurements point in opposite directions, and both are real:
- Large gains, concentrated at the bottom of the experience curve. Across 5,179 customer support agents given an AI assistant, issues resolved per hour rose 14% on average, but about 34% for novice and low-skilled workers, with minimal effect on experienced and highly skilled ones 3. If your exposed family is mostly senior, the average in that headline is not your number.
- A measured slowdown among experts on familiar ground. In a randomized trial, 16 experienced open-source developers took 19% longer to complete 246 issues when allowed to use early-2025 AI tools, in repositories they already maintained. They had forecast a 24% speedup, and after finishing still believed they had been sped up by 20% 4.
That 39-point gap between what the developers experienced and what they believed is the single most important line for a planner. Self-report is the weakest instrument in the building, and it is what most headcount assumptions are built from. A follow-up run on late-2025 tools, with 57 developers across 800-plus tasks, still measured no speedup (roughly -18% for returning participants and -4% for new ones, both intervals wide enough to cross zero), and the researchers say selection effects make even that weak evidence 5.
So measure. Take one exposed family, split real work between an AI-allowed arm and an AI-disallowed arm for a month, and count completed units and rework rather than asking anyone how it went. Two arms, one month, real tickets. That gives you a low and a high for one row of the table, which is more than any published figure can give you.
Run the range per occupational family
Five columns, one row per family, filled in a morning from documents you already own. Current headcount, share of hours on tasks an assistant can draft, the low and high gain you measured, and the review time added. Compute the need at both ends, then write the one thing that would move the row. The output is a band per family and a date to look again.
1. Map families to occupations, roughly. Your "Senior Insights Manager" is a market research analyst or a data scientist. Pick one, take the published task list, and move on. The point of an external list is that nobody in the budget meeting wrote it. 2. Weight the exposed tasks by hours, not by count. Ask the manager which three tasks eat the week. Exposure on a task nobody spends time on is worth nothing to the plan. 3. Enter the gain as a range. Low end and high end from your own two-arm measurement. If you have not measured, the low end is zero and the high end is whatever the pilot claimed, which is a wide band on purpose. 4. Subtract the review time. Percentage points of senior capacity spent checking drafts that would not otherwise exist. Nobody enjoys this column; every plan that omitted it has been wrong. 5. Divide forecast volume by the result, at both ends.
A worked example, four families of a hundred people each, flat volume:
| Family | Hours on drafted tasks | Measured gain | Review time added | Need next year |
|---|---|---|---|---|
| Customer support | 55% | 10–30% | 4 points | 89–99 |
| Financial analysis | 30% | 0–15% | 3 points | 99–103 |
| Software engineering | 40% | -5–20% | 5 points | 97–108 |
| Field operations | 10% | 0–10% | 1 point | 100–101 |
Four hundred people today; the honest total for next year is 385 to 411. No single percentage would have produced either end, and the two families that move most move in opposite directions. That total is not a hedge. It is the finding. It tells you which family to measure first, because narrowing software engineering is worth eleven people and narrowing field operations is worth one.
What do you do with a range instead of a number?
Budget the low end, pre-approve the difference as requisitions that open on evidence, and attach a named trigger to every row. A range is a commitment device: it says what you would have to see before moving, in advance, while nobody is arguing. The trigger is a measurement (completed units per person, rework rate), never a seat-licence count or a training completion rate.
Four things the range changes that a single number cannot:
- The entry level gets planned deliberately. Payroll data covering millions of U.S. workers shows employment of 22-to-25-year-olds in the most AI-exposed occupations sitting 19% below where it would be had it tracked their less-exposed peers, with no comparable gap for experienced workers in the same occupations 6. Your senior bench is a stock that juniors refill; cutting the intake is a decision about 2030 taken in a 2027 budget meeting. Whether a junior with an assistant substitutes for a senior is the question to answer before the requisition closes, not after.
- Train-or-hire is decided per family. A family with three exposed tasks and a patient manager is a training problem. A family where the exposed work is most of the job is a hiring decision with its own arithmetic.
- The review column gets an owner. Training people to check their own output is worth doing and is not the same as having a second reader, and treating training as a substitute for verification is how a plan books a saving that shows up later as rework.
- The hiring bar gets written before the requisition. If the plan assumes people who work well with an assistant, the loop has to be able to tell. That means a work sample with the assistant present, a rubric written beforehand, and knowing what an assessment result does and does not predict before it carries any weight.
Revisit rows when a trigger fires, not on a calendar. A tool that lands in the finance stack, a review step removed, a team absorbing a function: each moves one row. An annual refresh of a document nobody consults is the thing this exercise exists to avoid.
Common questions
What productivity assumption should we use if we have measured nothing yet?
Zero, with a measurement scheduled inside the quarter. A zero assumption is not pessimism. It is the only entry that does not commit money to a number nobody has seen. Then run one two-arm comparison on a single exposed family and replace the zero with a real low and high. Planning on a vendor figure or a pilot's self-report costs more than waiting a quarter, because the requisitions you did not open are harder to recover than the ones you did.
Does a productivity gain automatically mean fewer people?
No, and assuming it does is the most common error in these plans. Headcount is work volume divided by throughput, so a throughput gain reduces headcount only if volume is fixed. In most functions it is not: a support queue clears faster and the backlog shrinks, an analyst team answers more of the questions it used to decline, a product team ships more of a roadmap that was already too long. Ask what happens to the extra capacity before booking it as a saving.
How do we measure realized gain without running a formal study?
Split real work into two arms for a month. Same team, same task type, half the work done with the assistant and half without, assigned rather than chosen, because self-selection is what breaks these measurements. Count completed units and count rework separately. Do not ask anyone how much faster they felt: in a randomized developer trial, participants believed they had been sped up 20% while measurably running 19% slower 4. Feelings and throughput came apart by 39 points among people doing the work.
Should we stop hiring entry-level roles?
Not as a blanket decision, and not before checking what the juniors currently absorb. The exposed tasks in many families are the ones juniors learn on, so cutting the intake removes the training path that produces your future seniors, while the review work AI creates lands on the seniors you already have. Where the evidence points is toward planning the entry level explicitly with a defined path, rather than letting it fall out of a productivity assumption applied to the whole company.
What do we do with a vendor's industry productivity figure?
A vendor's figure is a hypothesis about a different task mix. A published gain was measured on a specific set of tasks, with a specific experience mix, under conditions you cannot see. The support-agent study that found a 14% average also found roughly 34% for novices and minimal effect for experienced workers 3: same tool, same firm, one number that describes neither group. Use an outside figure to set the top of your range, then measure your own low end.
How often should the plan be revisited?
When a trigger fires, which is usually two or three times a year and never on a fixed date. Write the trigger into the row when you build it: a tool reaching general availability inside a workflow, a review step removed, a team absorbing a function, or a measurement crossing the threshold you named. Triggers keep the plan honest because they are written before anyone has a stake in the answer. A calendar refresh produces a document that is updated and unread.
References
- 1. GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models ✓ arxiv.org Around 80% of the U.S. workforce could have at least 10% of their work tasks affected by LLMs, and approximately 19% of workers may see at least 50% of their tasks affected; exposure is scored at the occupation level.
- 2. The Anthropic Economic Index ✓ anthropic.com Roughly 36% of occupations showed AI use for at least 25% of their associated tasks, and only about 4% for at least 75% of tasks.
- 3. Generative AI at Work ✓ nber.org Across 5,179 customer support agents, access to an AI assistant raised issues resolved per hour by 14% on average, about 34% for novice and low-skilled workers, with minimal effect on experienced and highly skilled workers.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org Randomized trial: 16 experienced developers on 246 issues in repositories they maintain took 19% longer with AI tools allowed, having forecast a 24% speedup and believing afterwards they had been sped up 20%.
- 5. We are Changing our Developer Productivity Experiment Design ✓ metr.org Follow-up on late-2025 tools with 57 developers across 800-plus tasks measured roughly -18% speedup for returning participants and -4% for new recruits, both with intervals crossing zero, and the authors describe the evidence as weak because of selection effects.
- 6. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence ✓ digitaleconomy.stanford.edu ADP payroll data covering millions of U.S. workers through June 2026: employment of 22-to-25-year-olds in the most AI-exposed occupations stands 19% below where it would be had it kept pace with less-exposed peers, with no comparable gap for experienced workers.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.