Teams
A Central AI Team Buys Demos; Embedded Capability Buys Work
Keep a small central AI group for two jobs: buying and securing the tools, and setting the verification standard every team's deliverables have to meet. Give delivery to the people who own the work product, because the judgment being learned is occupational, meaning knowing which claim in this particular job is expensive to get wrong. Write the sunset condition down at the start, or the central group's headcount quietly becomes the goal.
The takePilots launched, use cases catalogued, people trained: three numbers a central group can move without anything changing in the functions that paid for it. Ask instead for one deliverable from an operating team that is demonstrably different from the version it shipped last year, and ask the person who owns that deliverable to describe the difference. If nobody can, the budget bought demos. That is a survivable answer in year one and an expensive one in year three.
Where Olive fits
Open a role and see what the work shows
The six dimensions Olive reports on are one usable statement of the standard a central group is trying to set: problem framing, evidence sourcing, delegation boundary, working structure, output rejection, and verification. Every finding is written by a person and carries the excerpt it rests on.
Rank your shortlistWhich two jobs belong in the center?
Tooling and the standard. Tooling means contracts, security review, data handling terms, access management and a single written statement of what may go into which system. The standard means one definition of what counts as a checked deliverable, applied identically across functions, so that a marketing brief and a claims decision are held to the same evidence rule even though the content shares nothing.
Both are genuinely central work because both are worse when duplicated. Five functions negotiating five contracts get five sets of terms and no bargaining position. Five functions writing five verification standards get five, which is the same as none, because nobody can then say across the company what a checked piece of work means.
Keep the group small enough that it cannot absorb delivery. Three to six people is enough at mid size: someone who owns vendor and security, someone who owns the standard and the training design, and one or two who can sit with a function for a fortnight and leave.
What does not belong in the center is the thing every consultancy proposal puts there, which is a portfolio of use cases the central group builds on behalf of the business. That arrangement makes the center the customer of its own work, and the measure of success becomes the portfolio.
Why does delivery belong to the function?
Because the expensive judgment is occupational and does not transfer. Knowing that a comparable-sales figure is the number a valuation turns on, or that a coding modifier decides whether a claim is paid, is domain knowledge held by the people doing the job. A central specialist can build the workflow and cannot tell you which line in it will be costly if the model gets it wrong.
The Boston Consulting Group field experiment is the cleanest illustration. On 18 realistic consulting tasks inside the model's capability, consultants using GPT-4 completed 12.2% more tasks, finished them 25.1% more quickly and produced work graders rated more than 40% higher 1. On one task deliberately placed outside that capability, the same tool made people 19 percentage points less likely to be right 2. One firm, one sitting, a 2023 model. The transferable part is that the boundary was invisible to capable people, and the only defence against an invisible boundary is someone who knows the domain well enough to be suspicious in the right places.
The same reasoning explains why the work reshapes rather than disappears. The Burning Glass Institute found skills exposed to automation were 16% more likely than baseline skills to see demand decline in postings, while skills exposed to augmentation were 7% more likely to see demand rise, with the most automated occupations also the most augmented 3. Those are relative likelihoods against a baseline rather than the size of any change, and posting text is employer language. The shape still matters for org design: capability that reshapes a job has to live in the job.
Work out where that applies before staffing anything, because it is rarely everywhere at once. Which roles actually need AI skills right now is the question that sizes this decision.
Write the sunset condition before the first hire
State in the founding document what has to be true for the central group to shrink, and put a date on the review. Something checkable: tool contracts signed and renewing on schedule, the verification standard published and in use by every function, and two consecutive quarters where operating teams ran their own reviews without central help. Then review it on the date whether or not anyone raises it.
Without that clause the group does what every permanent function does, which is to find more work. The failure is not bad faith. A central team measured on activity will produce activity, and the activity that is easiest to produce is another pilot in another function, which also happens to justify another hire.
Three clauses worth having in the founding document.
- A named end state. What the organization looks like when the center is three people instead of nine, described concretely enough to argue about.
- A review date, not a review trigger. Triggers are never pulled by the people they affect.
- A rule about delivery headcount. If the central group is building deliverables for a function after the first two quarters, either the function is not staffed for it or the center has absorbed the work.
If a mandate arrived from above rather than from a function's need, the sequencing is different and tighter, and what has to be built in the first ninety days is a shorter list than the announcement implied.
What should the exec team ask for at the quarterly review?
One artifact per operating function, held next to what that function produced a year earlier, with the change named by whoever is accountable for it. Not slides about pilots. A quarterly report, a claims decision file, a campaign brief, a forecast, and a specific difference in each: a brief written before generation, a claim checked against a source, a draft rejected for a stated reason.
Be careful with the adoption numbers that arrive alongside the request. Published US adoption rates disagree by construction rather than by error: about 18% of firms had adopted AI at the end of 2025 on the Census measure, about 41% of the workforce reported using it at work in the same period, and an employment-weighted estimate put 78% of the labor force at firms that had adopted something 4. Those are three different units, and the 78% describes employers rather than users. Any of them can be quoted to make a program look ahead or behind.
The firm-level picture is also strongly size-dependent. Census put AI use among US businesses at 19.8% as of May 2026, against 37% among firms with at least 250 employees 5. Useful for sizing expectations, useless as a target, and the survey asks about use in the past two weeks rather than about depth.
Then ask two questions that no dashboard answers. Whether the verification standard is being applied by people who do not report to the central group, and whether the functions could keep going if the group were halved next quarter. If the answer to the second is no after two years, the capability was never embedded, and the headcount plan is resting on an assumption nobody has tested. The program that survives this review is the one where the upskilling ran inside the functions rather than beside them.
Common questions
What size company actually needs a central AI group?
Roughly the size where tool contracts and security review stop being one person's side task, which in practice is a few hundred employees and up. Below that, a named owner with a day a week does both central jobs. Above it, the cost of every function negotiating and reviewing on its own starts to exceed the cost of two or three dedicated people. Neither threshold is about AI sophistication. Both are about procurement and risk load.
Where should the central group report?
To whoever owns the operating cost of getting work wrong, which is usually the COO or the CFO rather than the CTO. Placing it under technology makes tooling the natural output and turns the verification standard into an afterthought. Placing it under a transformation office makes announcements the natural output. Reporting into operations keeps the question focused on deliverables, which is where the value and the exposure both sit.
Isn't a center of excellence the standard model?
It is the standard proposal, which is not the same thing. The model comes from consultancies that staff the function they recommend, and it is usually judged on what the center itself produces rather than on what the functions ship. That output can grow steadily while the deliverables it was bought to improve stay exactly as they were. Keep the parts that centralize genuinely shared costs and refuse the part where the center owns delivery.
How do we stop each function buying its own tools?
Make the approved list short, current and genuinely faster to use than a procurement exception. Shadow purchasing is usually a response to a central process that takes weeks to say yes. Publish what is approved, what data may go into each system, and a route for exceptions that resolves in days. Enforcement matters less than latency here: functions route around slowness, not around rules.
What if the functions say they do not have time for this?
They are describing the real constraint, and it is the one to fund. Embedded capability costs protected hours from people who are already delivering, which is a harder budget line than a central hire because it shows up as reduced output this quarter. Naming it plainly beats pretending the work is free. Start in the function where a wrong answer is cheapest, prove the pattern, and use that evidence to buy the hours elsewhere.
Should the central group run the training itself?
It should design the training and hand delivery to the functions. Central design keeps one standard and one set of reading questions, which is what makes results comparable across the company. Central delivery makes the sessions generic, because a facilitator without occupational credibility cannot tell which claim in a given deliverable carries the decision. Design once, run locally, and have the center read a sample of the output for consistency.
References
- 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that AI raises output substantially on tasks inside the model's capability.
- 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that the capability boundary was invisible to capable people on one out-of-frontier task, which is the argument for domain knowledge close to the work.
- 3. Beyond the Binary: How Automation and Augmentation Are Combining to Reshape Work burningglassinstitute.org Supports the claim that AI is reshaping what roles do rather than removing them, stated as relative likelihoods against a baseline.
- 4. Monitoring AI Adoption in the US Economy (FEDS Notes) federalreserve.gov Supports the caution that published AI adoption rates measure different units and cannot be compared directly.
- 5. Large Firms With at Least 20 Employees Biggest AI Users census.gov Supports the firm-level adoption baseline and the size gradient between large and small employers.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.