Teams
What Do You Hire For After AI Didn't Absorb the Work?
After cutting a team betting AI would absorb the work, hire for the work that came back. The tasks that landed on whoever was left are the specification: exceptions the tool couldn't close, escalations it created, and checking nobody was assigned. Pull four weeks of it, sort by what each item cost when it went wrong, and write the requirements from the top. Expect fewer people at a higher band, plus a junior rung so somebody learns the exception set. Reposting the old req re-hires the volume AI absorbed.
The takeThe cut wasn't wrong about the tool. It was wrong about who was holding the parts the tool never touched, and nobody had written that down, which is why the bet looked clean on a spreadsheet. What the reversal buys you is an inventory of your own exception set that no planning exercise would have produced, paid for in people. Spend it properly. Then look at the bill after this one: a company that buys judgment at senior rates every time it runs short, while the rungs that used to grow it sit empty. I'd expect that one to arrive late and large.
Where Olive fits
Open a role and see what the work shows
The same behaviors describe the work that came back: framing before generating, demanding a source for the claim that matters, keeping the judgment that shouldn't be handed over, and testing a claim against something outside the conversation. Olive assesses those in a 40-to-60-minute occupational assignment and returns six separately evidenced findings written by a human reviewer, with the candidate granted the same report.
Rank your shortlistWhat actually came back after the cut?
The residual. Not the whole job you removed, but the slice the tool couldn't close: the odd case, the angry escalation, the number that had to be checked before it went out, and the coordination nobody automated because it lives in three systems. That slice used to be distributed across a team. After the cut it lands on whoever is still there, undocumented and unowned.
The shape is predictable, because AI usage lands on a narrow band of an occupation's tasks rather than spreading evenly across it. Anthropic's Economic Index, which mapped roughly a million Claude.ai conversations onto the roughly 20,000 work tasks in the U.S. Department of Labor's occupational database, found that "only approximately 4% of jobs used AI for at least 75% of tasks," while "roughly 36% of jobs had some use of AI for at least 25% of their tasks," with 57% of tasks augmented against 43% automated 1. A cut priced against a whole role removes the tasks the tool never touched along with the ones it did.
The lift also concentrated where the work was routine. In a staggered rollout of a generative AI assistant to 5,179 customer support agents, Brynjolfsson, Li and Raymond measured issues resolved per hour rising 14% on average and 34% among novice and low-skilled workers, with minimal effect on experienced and highly skilled ones 2. The cases the tool did not lift are the cases that come back, and they were already being handled by the people a headcount model reads as most expensive.
Read across four functions and the returned work has the same shape:
- Customer operations. The refund that falls outside written policy, and the account the assistant already answered once with the wrong clause.
- Finance and accounting. The reconciliation where two systems report the same figure differently, and the close item nobody can source.
- Engineering. The change that passes the tests it was handed, and the incident that starts with a merged patch nobody read.
- Legal operations. The markup a summary called standard, where the indemnity cap moved.
- Marketing and research. The on-message statistic that turns out to be a vendor's own customer survey, written up as market-wide.
All of them are judgment on an exception, and what AI actually does in a role is the same question asked before a cut rather than after one.
Why reposting the old job description re-hires the wrong half
Because the old req described volume, and volume is the part that got absorbed. It listed the throughput the role produced: tickets closed, decks built, models updated, copy shipped. Hire against that text and you select for people who are fast at the thing your tool is now adequate at, and you never test the exceptions that broke the bet in the first place.
A quieter version of the same mistake: adding "experience with AI tools" to the same req. It reads as a correction and changes nothing, because it selects for people who list tools rather than people who have caught a tool being wrong. Write the requirement as a task rather than a tool name. The sentence that works names the decision the person has to make and the evidence they have to produce for it.
The reason the first bet failed is also the reason it looked fine going in. METR ran 16 experienced open-source developers through 246 issues on repositories they already maintained, and found that when they were allowed to use AI tools they took 19% longer, having forecast a 24% speedup and still believing afterwards that they had been sped up by 20% 3. Self-reported velocity moved the opposite way from measured velocity. If the business case for the cut rested on how fast the pilot felt, it rested on that gap.
And the returned work has a price, most of it checking. In Stack Overflow's 2025 developer survey, 45.2% of the 31,476 respondents who use AI tools and answered its frustrations question selected "debugging AI-generated code is more time-consuming" 4. That is not a complaint about quality, it is a line item. The work moved from producing to verifying, and verifying was nobody's role on the org chart you cut.
Read four weeks of the returned work, then write the requirements
Open the queues and read what a person actually touched since the cut. Four weeks is enough. For each item, record what came in, why the tool didn't finish it, what the human did, and what it would have cost if nobody had. That table is the job description, written from evidence rather than from memory of the role you removed.
Four columns, one row per item:
- What arrived. The request, ticket, exception or question, in the words it came in.
- Why the tool stopped. Wrong policy clause, missing context, two sources disagreeing, a judgment the model is not permitted to make.
- What the person did. The act, not the outcome: opened the contract, called the vendor, recomputed the figure.
- What it would have cost. The refund, the restatement, the outage, the client call. A number where you have one, a sentence where you don't.
Sort by the last column. The top of that sorted list is the role. Everything below the line where the cost stops mattering is a candidate for staying automated, and saying so in writing is what keeps the rehire from becoming a quiet reversal of the whole decision.
Then size it from the same log rather than from the org chart you used to have. Count the hours the returned work consumed and the hours it sat in a queue, and carry a range instead of a point, because the boundary moves when a model changes and moves again when a process migrates. The general form of that is to plan headcount as a range rather than a number. If several roles are in play at once, triage which roles need AI skills by task before writing any of the reqs.
Plan for fewer people at a higher band than the ones you cut. Exception handling is judgment under ambiguity with a cost attached, and it prices accordingly. Put that in the plan in the same sentence as the savings, because a rehire that reverses the cut at a higher unit cost is a number your board will find on its own.
What should the second hiring pass assess for?
Three behaviors, tested on the returned work itself. Does the candidate frame the problem before generating anything. Do they demand a source for the claim the decision rests on. Do they check the model's output against something outside the conversation, and change the answer when the check disagrees. Those are the acts the automation could not perform, which is why the work came back.
Build the exercise out of the log. Take one real item, strip the identifiers, hand it over with an assistant available and 45 minutes on the clock, and ask for the decision plus the reason it holds. The assistant should be willing to produce a confident answer that is wrong in a way only checking catches, because that is what the returned work does every day.
Five versions of that exercise, one per field:
- Customer operations. A refund the assistant already denied, citing a clause that doesn't apply. Ask what to send and what the first answer got wrong.
- Financial analysis. Two systems reporting the same revenue figure differently, with a model happy to reconcile them in prose. Ask which is right and what settled it.
- Software engineering. A patch that passes every test it was handed. Ask what it breaks, and where the fixture agrees with the bug.
- Legal operations. A summary calling a markup standard. Ask what actually changed, and against which signed precedent.
- Market research. The most quotable statistic in the packet, sourced to a vendor's own 40-person customer survey. Ask whether it can carry the recommendation.
Grade the acts, not the volume of assistant use. Someone who judged the model was the wrong instrument for a step and did it by hand has demonstrated exactly the thing you're short of. This is the same exercise as hiring for verification rather than production, aimed at a role you already know is missing it.
Give every candidate the same item and the same rubric, including the people you're recalling, and tell them in advance what is being read. Two graders should be able to fill that rubric separately and agree, which "caught that the clause didn't apply" allows and "showed good judgment" doesn't.
Don't staff the returned work with seniors alone
The returned work is senior-shaped, so the rehire is usually fewer people at a higher band. Staffing it entirely at that band still leaves nobody learning the exception set, and the junior rungs in AI-exposed occupations are thinning at once. Stanford's Digital Economy Lab, reading ADP payroll records through June 2026, puts employment of 22-to-25-year-olds in AI-exposed occupations 19% below where it would be had it kept pace with their less-exposed peers 5.
The same paper splits the effect the way this problem splits. Declines concentrate in occupations where AI usage primarily substitutes for human tasks; where usage primarily complements workers, employment is flat or rising, especially for experienced workers, and the adjustment runs through reduced hiring rather than through separations 5. The authors call these descriptive indicators rather than causal estimates, and that hedge is worth keeping. What it means for your plan is narrower: the people who would have grown into the exception work are being hired well below trend across AI-exposed occupations, so the bench does not arrive on its own.
Some of the returned work is trainable and some isn't. Reading a contract against a system record is teachable in a quarter. Knowing which customer's escalation is really a churn signal takes a year of that customer. Split the log on that line and decide hire against train on the two gaps separately, then train the first half in the people you kept and hire for the second.
Two limits on all of this, worth stating before the plan goes out. The log is four weeks of one team's queue, so it captures the exceptions that arrive in four weeks and not the annual ones: re-read it after the close, the renewal cycle, or whatever your seasonal spike is. And the boundary it describes will move, in both directions, the next time a model or a process changes. Re-read it quarterly, and price the next automation decision on measured throughput rather than on how fast the pilot felt, which is the instrument that failed the first time.
Common questions
Should you rehire the people you laid off?
Often yes, and it's usually the cheapest option available: they already know the exception set, which is the part that came back. Two conditions make it work. Put the offer at the band the returned work actually sits at rather than the one they left, because the job changed. And run them through the same exercise every external candidate gets, on the same rubric, so the decision rests on evidence rather than on who's available. Expect some to decline. A recall at the old title reads as an admission that nothing was reconsidered.
How many people do you need back?
Fewer than you cut, almost always, and the log will size it. Count the hours a person spent on returned work across four weeks, add the hours the work sat waiting in a queue, and divide by a realistic week. Then hold a range instead of a number, because the tool will absorb more of the queue next quarter and less of it during the next migration. Commit to the bottom of the range in permanent headcount and the top of it in contract or overtime capacity.
What do you put in the job description this time?
The returned tasks, named as tasks, with the decision attached to each. "Resolves refund exceptions that fall outside written policy and states the clause the decision rests on" beats "experience with AI tools." A tool name selects for people who list tools. A task with a decision in it selects for people who have made that decision. Keep one honest line on what the assistant does in the role, so candidates can tell the job wasn't quietly re-scoped back to production volume.
How do you test judgment over model output without building a whole assessment?
Take one real item from the returned-work log, remove the identifiers, and hand it to the candidate with an assistant open and 45 minutes. Ask for the decision and the reason, not a finished artifact. Read for three things: whether the problem got framed before anything was generated, whether a source was demanded for the claim the decision rests on, and whether anything was checked outside the model. Give every candidate the same item and the same rubric, and say in advance what is being read.
Should the returned work go to contractors instead?
For a burst of volume, yes. For the exception set, rarely. The returned work is where institutional knowledge lives: which customers escalate, which clause the last dispute turned on, which report nobody trusts. A contractor closes tickets and takes that knowledge at the end of the engagement. The honest split is contract capacity for volume that arrives in spikes, permanent staff for judgment that has to accumulate. The four-week log tells you which is which: recurring exception classes are staff, one-off backlogs are contract.
What if the AI investment starts working in six months?
Then the log changes and you read it again. That's the argument for quarterly re-reading rather than for waiting: the boundary between what the tool finishes and what comes back moves, and it moves in both directions when a model changes or a process migrates. Hire for the judgment that survives the boundary moving, and keep contract or overtime capacity for the part that doesn't. A hiring plan pinned to one quarter's automation boundary is the same bet that already failed, aimed at a different number.
References
- 1. The Anthropic Economic Index ✓ anthropic.com Roughly a million Claude.ai conversations mapped to O*NET tasks: only approximately 4% of jobs used AI for at least 75% of tasks, roughly 36% had some use for at least 25% of tasks, and 57% of tasks were augmented against 43% automated.
- 2. Generative AI at Work (NBER Working Paper 31161) ✓ nber.org Staggered rollout of a generative AI assistant to 5,179 customer support agents: issues resolved per hour rose 14% on average and 34% among novice and low-skilled workers, with minimal effect on experienced and highly skilled workers.
- 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced open-source developers across 246 issues took 19% longer when allowed to use AI tools, after forecasting a 24% speedup and still believing afterwards they had been sped up by 20%.
- 4. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co 45.2% of the 31,476 respondents who use AI tools and answered the frustrations question selected "debugging AI-generated code is more time-consuming".
- 5. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence ✓ digitaleconomy.stanford.edu ADP payroll records through June 2026: employment of workers aged 22 to 25 in AI-exposed occupations sits 19% below where it would be had it kept pace with less-exposed peers; declines concentrate where AI substitutes for human tasks and run through reduced hiring rather than separations. The authors present these as descriptive indicators, not causal estimates.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.