Teams
A Manager Who Has Never Used AI Cannot Review AI Work
Managers who lead people working with AI need hands-on use themselves, and the requirement is hours rather than a credential. A manager reviewing AI-assisted work has to recognize three failure shapes on sight: the fabricated citation, the plausible number, and the fluent answer to a question nobody asked. That recognition comes from having been confidently misled a few times. Roughly two hours a week on their own real work, for a month, is the dose. Certification tests vocabulary and does not substitute.
The takeThe advice sold to executives, that leaders supply vision and change management while other people supply hands, was correct when a manager's job was approving a budget. It is wrong for this. The judgment now being asked of managers is whether a fluent deliverable was ever checked, and that judgment is not available secondhand. Any leadership programme promising AI capability without anyone opening the tool is selling the comfortable half of the job.
Where Olive fits
Open a role and see what the work shows
Verification is one of the six dimensions an Olive report returns a written finding on, and each finding quotes the moment in the session it came from. A manager who has spent hours in the tool reads an excerpt like that differently from one who has not.
Rank your shortlistWhat does a manager actually need to be able to do?
Recognize three failure shapes on sight, and ask for the trail behind a claim without the question landing as an accusation. That is the job. The three shapes are the fabricated citation that resolves to nothing, the plausible number that is wrong in the second decimal or the second order of magnitude, and the fluent answer to a question nobody actually asked.
The first is the easiest to catch and the least common in practice, because it fails the moment anyone opens the link. The second is the expensive one: a figure with the right units, the right rough size and the wrong value, sitting in a paragraph that reads like it was checked. The third is the hardest of all, because nothing in it is false. The brief asked what the renewal risk was and the deliverable explains, very well, what renewal risk is.
None of the three is visible from prose quality. All three are visible to someone who has produced them and noticed. The whole argument for hands-on hours sits there: the failure shapes are learned the way any other professional pattern is learned, by having been taken in and then working out how.
Why doesn't a certificate cover it?
Because a certificate exam tests recognition of vocabulary under time pressure, and the two guides quoted here set out no performance task. Google Cloud states plainly that its Generative AI Leader certification is for anyone in any job role, with or without hands-on technical experience 1. That is not a criticism of the exam. It describes what the exam is for, and it is not this.
The technical-sounding ones say the same thing in more detail. The AWS Certified AI Practitioner exam guide describes a target candidate with up to six months of exposure to AI on AWS who uses but does not necessarily build AI solutions, and lists developing or coding models as explicitly out of scope 2. The question types are multiple choice, multiple response, ordering, matching and case study. A pass evidences familiarity with concepts under time, which is a real thing to know and a different thing from judgment in use.
Even the regulator that pushed hardest declined to make a credential the answer. In the EU, the European Commission's own questions and answers on the AI Act's literacy obligation, which has applied since 2 February 2025, state that Article 4 does not entail an obligation to measure the knowledge of AI of employees, that there is no need for a certificate, and that no mandatory trainings are imposed 3. The Q&A is guidance and not statute. Its wording here is the August 2026 version of a page revised once already, so any compliance conclusion belongs with counsel. It cuts both ways too: no vendor's badge confers compliance, and the Commission says so about its own repository of practices.
So fund a course if the vocabulary is genuinely missing. Do not report it upward as capability, and do not accept it in place of the hours.
Give each manager two hours a week on their own work
Two hours a week, on work they were going to do anyway, for about a month. Not a sandbox, not a prompt library, not a workshop: their own real deliverables, where they already know what a right answer looks like and will therefore notice a wrong one. Nobody measured that dose; it is a judgment, and it is small enough that no calendar argument survives contact with it.
Three exercises fill the hours better than any curriculum:
- Ask for sources on a claim you already know cold, then open two of them. The gap between what the source says and what the summary said it says is the lesson, and it only lands the first time on a topic where you can see the gap.
- Give it something you know sits outside what it can do. Watch it answer anyway, at full confidence, in your own subject. Run this one first. A confident wrong answer inside your own expertise is the clearest version of what these hours are for.
- Take one output you nearly shipped and write down what would have gone wrong. Not that it was wrong. What it would have cost, and at which meeting someone would have found out.
Keep a running file of the things it got wrong in your own work. That file is the curriculum, it is specific to the domain, and it is the only artifact from this worth showing anyone else. Two managers comparing their files learn more in twenty minutes than either learns from a vendor deck.
How should a manager ask about the trail?
Use the vocabulary of the work itself. What did you check this against. What did the first version say that this one does not. Which part are you least sure about. All three are questions a decent manager asked long before AI existed, which is exactly why none of them reads as an accusation, and why all three are answerable by someone who did the work properly.
The question that does read as an accusation is whether AI was used, and it is also useless, because style is not evidence of anything. Asked to tell GPT-3 text from human writing across stories, news and recipes, non-expert evaluators performed at random chance, and three quick training methods lifted them only to about 55%, inconsistently across the three domains 4. The study is from 2021, the judges were crowdworkers and not hiring managers, and the passages were short. Models have improved since, which makes unaided guessing worse rather than better, though that is an inference and not a measurement in the paper.
Guessing has to give way to asking, which makes this a management problem rather than a detection one. It is the same underlying issue as managers who don't use AI having to judge AI-assisted work in interviews, and the same reason to stop rejecting people for sounding like AI.
Once a manager can ask the three questions and read the answers, the review structure is straightforward: two recent deliverables, ten minutes each, looking for the traces good AI work leaves in the artifact. Without the hours, the same review runs and finds nothing, because the person running it cannot tell a specific answer from a confident one.
Common questions
Does every manager need this, or only ones running technical teams?
Every manager who approves work. The failure shapes are the same in a marketing brief, a financial model and a pull request, and the non-technical cases are often worse because nothing fails loudly. A manager who never reviews a deliverable, only outcomes, has a weaker claim on the hours, though the number of managers genuinely in that position is smaller than the number who describe themselves that way.
Is a certificate ever worth funding?
When the vocabulary is genuinely missing and someone needs a structured way in, yes, at a modest price and with no expectation attached. What a foundational credential evidences is familiarity with concepts under time, which is worth something at the start. What it does not evidence is whether the holder can spot a confident error in their own subject, and no exam of that shape can, because it contains no performance task.
What if a manager's own work has nothing worth doing with a model?
Use the work anyway, precisely because the answer is already known. A manager who asks a model to summarise a market they know cold, or to draft a plan for a project they have run, is running the ideal exercise: the wrongness is visible and costs nothing. The exercise is diagnostic, not productive, and it works best on ground the manager already knows.
How do I know the hours actually happened?
Ask for the file of things the model got wrong in their own work. It cannot be produced without doing the hours, it takes thirty seconds to look at, and its contents are immediately informative: specific entries with a domain in them mean the hours happened, while generic entries about hallucination mean they did not. It is a better check than attendance and much better than a self-rating.
Does this apply to executives who no longer review work directly?
Partly, and for a different reason. An executive who has never seen a model produce a clean wrong answer will systematically misjudge what their organisation can safely automate, and will approve rollouts on the assumption that the output arrives checked. A shorter version of the same exercise, a few hours total on a subject they know well, is usually enough to correct that assumption.
References
- 1. Generative AI Leader certification cloud.google.com Supports the claim that a non-technical AI credential is designed for candidates with no hands-on experience.
- 2. AWS Certified AI Practitioner (AIF-C01) Exam Guide, Version 1.4 d1.awsstatic.com Supports the claim that a foundational AI credential tests concept recognition and contains no performance task.
- 3. AI Literacy - Questions & Answers digital-strategy.ec.europa.eu Supports the claim that even the jurisdiction with a live AI literacy duty requires no certificate and no testing of staff.
- 4. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text aclanthology.org Supports the claim that guessing whether text is AI-written is not a usable management method.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.