Teams

The Minimum AI Literacy Is Three Judgments, Not a Curriculum

The minimum every employee needs to know about AI is three judgments, and they fit on one page. Which categories of work may be handed to a model at all. What verification an output needs before it carries someone's name. What data must never go into a prompt. Everything else, including how the technology works, is optional context. Written that way the minimum is role-independent and checkable: ask someone to apply the page to a task from their own week.

The takeA literacy programme that opens with how transformers work has chosen the comfortable half of the problem. The technical explanation is easy to write, easy to assess with a quiz, and satisfies nobody's actual exposure. The boundary questions are harder because they force a company to say out loud which work it is willing to have a model do and which client material may never be pasted anywhere. It is that sentence, rather than the training, that most organisations are quietly avoiding.

Where Olive fits

Open a role and see what the work shows

Two of these three judgments are what Olive reads from a real session: where the delegation boundary is held, and what gets verified before an output goes out. A human reviewer writes all six findings against timestamped excerpts, and outcomes read as demonstrated, partly demonstrated or not demonstrated, never as a number.

Rank your shortlist

What are the three judgments?

Delegation, verification and disclosure of data. Which categories of work may go to a model, which must not, and who decides the ambiguous ones. What a given output needs before it goes out under someone's name, and against what. And which material never enters a prompt at all, named concretely enough that nobody has to interpret it: client files, candidate records, unreleased numbers, anything under an NDA.

Each one has to be written in your categories rather than in the abstract. "Use good judgment with confidential information" is not a boundary, because two reasonable people will draw it in different places and both will believe they complied. "Nothing from the client folder, no candidate names, no pre-announcement financials" is a boundary. It can be followed by someone with no interest in the topic, which is the actual test of a rule.

The verification judgment is the one worth spending the most time on, because it is the only one where the right answer varies by task. An internal summary needs almost nothing. A number in a board deck needs a traceable source. A claim about a person, a customer or a regulation needs to be confirmed outside the conversation before it goes anywhere. Write three tiers, give an example of each from real work, and stop.

That is the page. A single side, three headings, examples drawn from the last month of actual output. It takes an afternoon to draft and most of the argument happens in the room where the categories get chosen, which is the point of the exercise.

Why the technology explanation can wait

Because the workplace failure is not ignorance of the technology. It is a confident wrong answer accepted by someone who had no reason to suspect it. Understanding tokenisation does not prevent that, and a person who cannot define a model at all will still catch the error if they were told to check the claim against the source.

The failure mode has been measured directly. In a field experiment with Boston Consulting Group consultants, one task was chosen to sit outside what the model handled well, and consultants using GPT-4 were 19 percentage points less likely to reach the correct answer than the control group, who got it right 84.5% of the time 1. One task, one sitting, a 2023 model, so the size means little. What travels is the shape of the failure: nobody could tell which side of the line the task was on, and the output looked the same either way.

What would have helped in that room is a rule: a recommendation with a number in it gets traced to a source before it goes to the client. Nobody has to know how the model produced the number to apply it. That is the whole argument for putting the boundaries first and the mechanism second, or last, or in an optional session for people who want it.

The audience is the other argument for keeping the page short. Baseline training reaches people who did not ask for it and will read it once, so the material has to survive being skimmed in four minutes by someone with no interest in AI. In nationally representative US surveys in late 2024, 23% of employed respondents had used generative AI for work at least once in the previous week and 9% used it every work day 3, so a company-wide rollout is mostly speaking to people who have not formed a habit yet and will decide what the habit is from this page. Three judgments with an example under each fit inside that reading.

Does the law require a certificate?

Not in the EU, which is the jurisdiction with an explicit literacy duty, and this is worth checking before anyone buys a certification programme. The AI Act's Article 4 obligation is to take measures, not to test or certify staff. The European Commission's own guidance says the duty does not entail an obligation to measure employees' knowledge of AI and that there is no need for a certificate 2. Check the current position with counsel before relying on it.

Two details in that guidance cut against what most summaries still say. The Commission states that the sentence everyone quotes, requiring a sufficient level of AI literacy, is no longer the operative wording: amendments made by the Digital Omnibus on AI entered into force in mid-July 2026, and AI literacy remains an obligation for providers and deployers while no specific level is mandated 2. Secondary trackers and law-firm summaries were still reproducing the pre-amendment text at the time of writing, so a quote from one of them may be quoting a superseded rule.

The obligation itself did not go away, and a narrower one is on the way. Deployers of high-risk systems will have to ensure staff are trained for human oversight, and the Commission is explicit that handing people the instructions for use is not sufficient there 2. Recruitment and workforce uses are named as high-risk in Annex III of the Act itself, Regulation (EU) 2024/1689. Those requirements are not live: the Commission's implementation timeline puts them at 2 December 2027, deferred from 2 August 2026 by the Digital Omnibus 4. So the team running recruitment software has a heavier duty ahead of it, and most compliance checklists still print the superseded date.

For a US employer none of this binds directly, and it is still the best available description of what a regulator thinks a baseline is: measures taken, records kept, no exam. Keeping an internal record of what was rolled out and when is the practical takeaway.

Write the page, then check it against one real task

Draft the three judgments, then hand the page to four people from different teams with one instruction: apply this to something you did last week. What comes back is the only test that matters. If two of the four reach opposite conclusions about the same task, the boundary is not written yet, and no amount of training will close a gap that lives in the wording rather than in the people.

The pattern of disagreement tells you where to rewrite:

  • Disagreement about what counts as verification means the tiers are missing an example from that kind of work.
  • Disagreement about whether something is confidential means the data rule names a principle where it needs a list.
  • Agreement that the page does not apply to their work at all means that team has an AI question you have not asked about yet.

Then decide what happens after the rollout, because a baseline everyone has read and nobody applies is the common outcome. Checking is not the same activity as training and it needs its own design, which is the subject of whether AI training is enough or the work has to be checked afterwards.

Two places this baseline shows up immediately. New hires who reach for an assistant before they reach for a colleague are applying the delegation judgment on their own, well or badly, from their first week, and what to do when new grads ask AI before they ask anyone covers that. And the moment the floor is written, the question of what sits above it becomes answerable, which is four levels of AI skill written as observable behavior.

See the benchmarks

Common questions

How long should baseline AI training take?

An hour of reading and a short applied exercise, once, with the page kept somewhere findable afterwards. The three judgments do not take longer than that to explain, and anything past the first hour is either role-specific work or the mechanism of the technology, which most people will not retain and do not need. If a proposed programme runs to days for everyone, ask which of the three judgments needed the extra time.

Should the baseline be the same for every role?

The three judgments should be identical; the examples under them should not. A finance team and a support team need the same rule about what gets verified and completely different illustrations of what a verification looks like. Keep one page company-wide so nobody can claim a different standard applies to them, and attach a short role-specific appendix with examples drawn from that team's own recent work.

What if people already use AI without any rules?

That is the normal starting position and it makes the page more useful, not less. Existing use tells you what the categories should be, so gather it before writing anything: ask two people per team what they have handed to a model in the last fortnight. The answers usually reveal one practice you want to stop and two you want everyone to copy, and the page writes itself around them.

Does anyone need to know how a model actually works?

Somebody in the company does, and it is a smaller group than the whole company. Whoever selects tools, negotiates vendor terms or oversees a high-risk system needs a real technical understanding. For everyone else, a short account of why a model produces confident text without knowing whether it is true does the useful work, and it takes a paragraph rather than a module.

How do you know the baseline landed?

Look at output rather than at completion rates. Pick a sample of work from the month after rollout and ask two questions of each piece: was anything in it checked against a source, and did anything go into a prompt that the page says should not have. A completion rate tells you the module was opened and nothing else.

References

  1. 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that the workplace failure mode is a confident wrong answer on a task nobody could tell was outside the model's range.
  2. 2. AI Literacy - Questions & Answers European Commission (Shaping Europe's digital future / AI Office), 2026. digital-strategy.ec.europa.eu Supports the claims that the EU literacy duty requires no testing or certificate, and that the sufficient-level wording is no longer operative after the Digital Omnibus amendment.
  3. 3. The Rapid Adoption of Generative AI (NBER Working Paper 32966) National Bureau of Economic Research, 2025. nber.org Supports the baseline figures for how many workers report using generative AI at work, used to size who the page has to reach.
  4. 4. Timeline: implementation of the EU AI Act European Commission - AI Act Service Desk, 2026. ai-act-service-desk.ec.europa.eu Supports the claim that the Annex III high-risk requirements apply from 2 December 2027, moved from the widely quoted 2 August 2026.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.