Teams
Should You Hire for AI Skills or Train the Team You Have?
Hire for AI skills or train the team you have? Both, for different gaps. Train for the tool gap: weeks of supervised practice close it, fastest for the least experienced [1]. Hire for the verification gap: knowing when a fluent answer is wrong is domain judgment, and no course moves it fast. Last quarter's escaped errors say which gap you have. Clumsy work that shipped means train. Fluent, polished and wrong means hire, and whoever catches those errors is usually on the payroll. Both piles full: train first, then hire.
The takeThe market is pricing this backwards. Tool fluency is the cheap half and it is the half job ads are written for, while the expensive half sits with people who have already been wrong in your field and can still feel the shape of it. Nobody has priced the two separately, and if the pattern holds, firms will keep paying a premium for the cheap half, because it arrives with a curriculum, a completion rate and an invoice, and judgment arrives as an apprenticeship nobody can put on a purchase order. Scarcity does not follow what is easy to buy.
Where Olive fits
Open a role and see what the work shows
Where the answer comes back "hire," the verification gap is what the six dimensions describe: framing before generating, demanding a source for the claim that matters, keeping the judgment that should not be handed over, and testing a claim against something outside the conversation. Olive is employer-purchased for hiring and reads those from one 40-to-60-minute occupational assignment, written up as six findings by a human reviewer, with the candidate receiving the identical report.
Rank your shortlistShould you hire for AI skills or train the team you have?
Both, but for different gaps. Train for tool fluency: it closes in weeks of supervised practice, and the measured gains are largest for your least experienced people 1. Hire for verification judgment: knowing when a fluent answer is wrong is domain knowledge, and no course transfers it quickly. Most teams have one gap badly and the other barely, then buy the wrong remedy because the question arrives as one question.
The two gaps behave differently under time and money.
- The tool gap is not knowing what an assistant can do, how to ask for it, or where it fits into the working day. It is procedural, it is teachable, and the evidence is unusually clean: across 5,179 customer support agents, access to a conversational assistant raised issues resolved per hour by 14% on average and by 34% among novice and low-skilled workers, with minimal effect on the most experienced 1.
- The verification gap is not knowing when the fluent answer is wrong. That is domain judgment in a new costume, and it does not move on a training calendar. Sixteen experienced developers working real issues in repositories they maintained took 19% longer with AI tools allowed, and still believed afterwards that the tools had sped them up by 20% 2.
Notice where each finding lands. The tool gap closes fastest for the people who know the job least. The verification gap bites hardest exactly where the work is most consequential, because that is where a confident wrong answer travels furthest before anything stops it.
If you have not settled which roles this question is even about, triage the roles by task exposure first. A hire-or-train argument about a role where an assistant does nothing consequential is an argument about nothing.
Which gap does your team actually have?
Read last quarter's escaped errors, not this quarter's tool survey. Pull the things that reached a customer, a client, a regulator or production and should not have, then sort each one into two piles: work that was clumsy, slow or in the wrong shape, and work that was fluent, well-presented and wrong. The first pile is a training bill. The second is a hiring decision.
Four questions settle it faster than any assessment vendor's demo:
1. Who on the team already does the checking? Name them. If one person catches everything, you do not have a capable team with a training gap; you have a single point of failure with a headcount gap. 2. Can anyone here write the answer key? Training only works if someone can say what a good check looks like in your field, on your material. If nobody can, no vendor's curriculum will supply it, and the training problem was a hiring problem wearing a purchase order. 3. Were the failures about speed or about being wrong? A slow, ugly, correct memo is a tool gap. A fast, polished memo with a number that does not tie out is not. 4. Would more AI usage have helped? Volume is not the variable. Someone who ran forty prompts and shipped every answer untouched used more AI than a colleague who ran four and threw out three.
Do not run this off self-report. In a survey of 319 knowledge workers describing 936 real tasks, higher confidence in the tool went with less critical thinking, while higher confidence in one's own domain ability went with more 3, so the most enthusiastic person in the room is not reliably the one doing the checking. The developers above misread their own throughput by nearly forty percentage points on work they knew intimately 2. A show of hands measures mood.
When the answer comes back "train," it is still worth reading what a training program can and cannot close before the invoice goes out.
Why does the verification gap depend on your field?
Because fields differ in how fast and how cheaply a wrong answer comes back. A compiler rejects bad syntax in seconds for free. A test suite catches a broken function in minutes. A reviewer catches a weak argument in days. A client catches a wrong number in a quarter, in front of their own board. The later the catch, the more of the checking one person has to carry alone.
That latency is the real variable behind the hire-or-train call, because it decides how much judgment the job requires before anything external saves you:
| What catches the error | How long, and what it costs | What has to be in the person |
|---|---|---|
| Compiler, type checker, linter | Seconds, free | Very little; the loop teaches itself |
| Test suite, reconciliation, a schema | Minutes, cheap | The habit of running it before shipping |
| Code review, desk check, second reader | Days, cheap if the reviewer is good | A reviewer who knows what to look for |
| A client, a customer, an auditor | A quarter, expensive and public | Judgment carried alone, in advance |
| Nobody, until the loss appears | Years | Someone who has already been wrong this way |
One caveat keeps engineering teams honest here: "the compiler catches it" is a claim about a class of error, not about a field. In a controlled study, participants with an AI code assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure 4. A security flaw compiles, passes review at a glance, and behaves exactly like a wrong number in a client deck: it comes back late, from outside, at cost.
The same survey that found effort shifting also found where it goes: from gathering information toward verifying it, from problem-solving toward integrating a partial answer, and from execution toward stewarding a task the model is doing part of 3. In a fast-catch field, your tools absorb some of that shifted work. In a slow-catch field, a person absorbs all of it, and that person either has the judgment or does not.
Run the diagnostic on last quarter's work
Give it an afternoon and two people: whoever owns the quality of the output, and whoever reviews it. Take twelve to twenty real items that shipped last quarter (memos, tickets, decks, claims, drafts) and mark each one on two questions only. Could a better question to the assistant have produced this faster? Would someone with more domain judgment have caught what was wrong?
The procedure, in order:
1. Sample what shipped, not what was praised. Pull items at random from the queue rather than from anyone's memory. Memory selects for the dramatic failure and the flattering success, and you need the ordinary middle. 2. Mark independently, then compare. Two markers who agree without conferring have found something. Two who disagree have found that the standard was never written down, which is its own result. 3. Count the second column. Items where a wrong thing survived to shipping are the verification pile. Items that were merely slow or awkward are the tool pile. 4. Read the split. Mostly tool pile: train, and verify afterwards on real work rather than on a completion certificate. Mostly verification pile: hire, or free up the person already doing the checking and make it their job. Both piles full: train first, because it is cheaper and its result arrives within a quarter, then hire once the tool gap stops confusing the measurement. 5. Write down what you decided and why. In six months someone will ask whether the training worked, and a before-picture is the only thing that can answer.
If the split points at hiring, the size of the ask is a separate calculation from the shape of it, and what the headcount actually needs to be is not answered by "we have a verification gap," only narrowed by it.
What does the hire have to be, when the answer is hire?
Hire for the checking, not for the tool list. The tool half of the job is the half your existing team can be taught, so paying a premium for it buys the cheap gap twice. What you cannot train inside a quarter is someone who has been wrong in your field before and knows the shape of it: the analyst who has watched a number fail to tie out, the engineer who has shipped the bug that compiled cleanly.
That changes what the interview has to observe. A candidate can describe checking a confident claim without ever having checked one, and a tool list on a resume is a purchase history rather than a record of judgment. The observable acts are narrow and worth naming: what got framed before anything was generated, what evidence was demanded for the claim the answer rested on, what was kept rather than handed over, what got refused and on what grounds, and what was tested against something outside the conversation. Building the exercise that shows those acts is a day of work and replaces a month of guessing.
Check inside first. The person who already catches the errors is usually on the payroll, and moving them into the AI-heavy role costs less than a search and carries the domain context a new hire will spend two quarters acquiring. The risk runs the other way too: enthusiasm about tools reads as readiness, and it is not the same thing.
One compliance note, because this diagnostic has teeth. Once a result decides who gets trained, who moves, or who is promoted, it is a selection procedure under the Uniform Guidelines, which name training programs explicitly 5. Selection for training or transfer counts as an employment decision when it leads to a covered one: hiring, promotion, retention 6. Nothing about that argues for skipping the exercise. It argues for running the same one for everyone in scope, on the same material, with the standard written before the first person sits it, and keeping the record of what you saw.
Common questions
How long before AI training shows up in the work?
For procedural use, weeks. Support agents given a conversational assistant resolved 14% more issues per hour on average, and novice and low-skilled workers 34% more, with minimal change for the most experienced 1. Expect the same shape internally: the people furthest behind move first and most. For judgment-heavy work the honest answer is that a course may not show up at all, which is why the check afterwards runs on real output rather than on completion rates.
Can verification judgment be trained at all?
Yes, but not on a training calendar. It moves the way domain judgment has always moved: someone reviews real work beside a person who knows what wrong looks like, and the person is wrong a few times where it is safe to be wrong. That is apprenticeship, measured in quarters. A curriculum with slides can teach the vocabulary in an afternoon and cannot compress the rest. Budget for the reviewer's time, not for the course fee, and expect the first year to be the expensive one.
Should the AI hire be a new role or an existing one?
Usually an existing one, filled by someone with deeper domain judgment. A standalone AI role tends to concentrate tool fluency where the errors are not, and it gives every other team a reason to stop thinking about it. The exception is a genuine build: if somebody has to stand up shared tooling, evaluation, or data plumbing, that is engineering work with its own job description, separate from the question of who on the team can tell when an answer is wrong.
What if the team will not use the tools at all?
That is a third problem, and neither training nor hiring fixes it. Non-use is usually a rational read of local incentives: unclear rules about what is allowed, a policy nobody has written, or a review process that punishes a visible mistake more than an invisible slowdown. Fix the rules first. A team that is quietly using assistants without telling you has the same root cause and is harder to see, because their output looks like compliance.
Does an AI certificate close either gap?
It is weak evidence for the tool gap and none for the verification gap. A certificate shows someone sat the material and could recognise a correct answer among four options. The behaviour that matters happens when no options are listed and a plausible answer already exists on the screen. Multiple-choice items cannot observe a refusal, a source being opened, or a figure being recomputed, because those are acts rather than answers. Treat certificates as a floor for tool access and vocabulary, not as a result.
References
- 1. Generative AI at Work (NBER Working Paper 31161) ✓ nber.org Across 5,179 customer support agents, access to a conversational assistant raised issues resolved per hour by 14% on average, 34% for novice and low-skilled workers, with minimal effect on the most experienced.
- 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues in repositories they maintained took 19% longer with AI tools allowed, and afterwards still believed the tools had sped them up by 20%.
- 3. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org 319 knowledge workers describing 936 real tasks: higher confidence in the tool is associated with less critical thinking, higher confidence in one's own ability with more, and the remaining effort shifts toward information verification, response integration and task stewardship.
- 4. Do Users Write More Insecure Code with AI Assistants? ✓ arxiv.org Participants with access to an AI code assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure.
- 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.16 (Definitions) ✓ ecfr.gov A selection procedure is any measure used as a basis for an employment decision, explicitly including training programs.
- 6. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.2 (Scope) ✓ ecfr.gov Section 1607.2(B): other selection decisions, such as selection for training or transfer, may also be considered employment decisions if they lead to hiring, promotion, retention or another listed decision.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.