Teams

How to Build an AI Upskilling Program That Changes the Work

An AI upskilling program that changes the work has three layers: access to the tools, supervised practice on the team's own real work, and a verification step where someone reads one finished work product per person before and after the program. Course completion is not the measure. The measure is whether a reviewer can point at a change in the artifact: criteria written before anything was generated, a claim checked against a source, a draft rejected with a stated reason.

The takeA completion percentage is a receipt for time spent, and it is the number most programs report because it is the only one they collect. The harder version: if nobody can hold up two deliverables from the same person and describe what is different about the second, the program did not land, whatever the dashboard says. Budget the reading before the training. The reading is the part that survives a skeptical question from finance, and it is the part every catalog leaves out.

Where Olive fits

Open a role and see what the work shows

The six dimensions Olive reports on describe what capable AI work looks like in a job: framing before generating, sourcing the claim that matters, holding a delegation boundary, structuring the work, rejecting an output, and verifying against something outside the conversation. A human reviewer writes each finding from what happened in the session.

Rank your shortlist

What goes in each of the three layers?

Access, practice, verification, in that order and at very different costs. Access is licences, an approved-use note and a list of what stays out of the tool. Practice is supervised sessions on the team's own live deliverables rather than sandbox exercises. Verification is one person, senior in the occupation, reading a finished work product per participant and writing down what changed.

Access is a procurement task and it is nearly always already done badly: people have a licence and no statement of what they are allowed to put in it. Fix that first, in one page, because every hour of training on top of an unclear permission gets spent on the permission.

Practice is where the design decisions live. Four sessions of ninety minutes, each one on a deliverable that was going to be produced anyway, with a facilitator who stops the room at the moment the assistant produces something plausible and asks how anyone would know. The work is real, the deadline is real, and the facilitator's only job is to make the checking visible.

Verification is the layer that gets cut, and cutting it turns the other two into an expense with no evidence behind it. It costs roughly two hours per participant across the whole program: one to read the earlier artifact, one to read the later one and write three sentences. Protect it in the budget line before the course fees, because course fees are the part vendors will discount and reader time is the part nobody will donate.

Why does course completion tell you nothing?

Because it records attendance at a curriculum, and the thing that goes wrong with AI at work is not absence of instruction. It is a confident answer accepted on a task that looked like the ones these tools handle well. In the Boston Consulting Group field experiment, on one task deliberately placed outside the model's capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer than the control group 1.

The two AI conditions did not fail equally: the group given a prompt-engineering overview did worse, down 24 points, than the group given no coaching at all, down 13 points 1. One task, one sample, a 2023 model, so nothing about it is a law. What it does show is that instruction which raises confidence without raising checking can move the wrong way, which is exactly the failure a completion rate cannot see.

The gains are real and they are uneven, which is the second reason a single headline number is useless to a program designer. Pooling three company-run randomized trials across 4,867 developers, an AI coding assistant raised completed tasks by 26.08%, with a standard error of 10.3% and the largest gains among less experienced developers 2. In a trial with 640 Kenyan small-business owners given a GPT-4 business assistant over WhatsApp, there was no detectable average effect on revenue or profit: high performers at baseline gained just over 15% while low performers did about 8% worse 3.

Different work, different populations, different outcomes. The pattern underneath is that the benefit tracks whether the user can judge the output, and judging an output is something a program can teach and a reader can check afterwards.

Who reads the work, and how long does it take?

Someone senior in the occupation, and about two hours per participant across the program. The reader needs to know which claim in this particular kind of document is expensive to get wrong, which is knowledge that does not transfer from a training function. Give them the before artifact, the after artifact, and three questions to answer in writing.

The three questions are the whole rubric.

  • Is there a statement of what was wanted before anything was generated? Constraints, audience, and what a wrong answer would have looked like.
  • Did a claim that carries the decision get checked against something outside the conversation? A source, a system of record, a colleague.
  • Was anything rejected, with a reason? An assistant that produces nothing worth discarding is being used as a transcription service.

Those are the same behaviours worth reading in a baseline exercise, which is why the two fit together: run the baseline assessment first, then set the program against what it showed. If the same reader does both, the comparison costs almost nothing extra.

One thing to decide before the first reading. The reader's notes are evidence about the program, not about the person, and saying so in advance changes what people submit. Where the notes become a performance record, participants produce the artifact they think will read well, and the program starts measuring presentation. Keep the individual notes with the individual and report the pattern upward.

Should everyone get the same program?

No. Access should be universal, practice should be occupational, and verification should be sized to how expensive a wrong answer is in that job. A marketing team and a claims team need the same permission and the same three questions, and almost nothing else in common, because the load-bearing claim in a campaign brief and the load-bearing claim in a coverage decision fail in different ways and at different costs.

That is also the answer to the sequencing question everyone asks. Start where the wrong answer is cheapest, so the practice is genuinely low-stakes, and expand toward the functions where it is expensive once the reading habit exists. The reverse order looks more ambitious and produces a program that has to be defended before it has any evidence behind it.

Two things to resist. The first is running the program because a mandate arrived rather than because a function needs it; if that is the situation, what HR actually has to build in ninety days is a narrower list than the mandate implies. The second is treating training as a substitute for hiring, or hiring as a substitute for training, without pricing either: the hire-or-train question has a different answer per function and it is worth answering per function.

Set the reporting line before the first session. Aggregate self-reported time savings from a nationally representative US survey ran to 5.4% of work hours among people who use these tools at work, which works out to 1.4% across all workers 4. That is the number a program will beat easily on a survey and hard on an artifact, and reporting the survey version to an executive who has read the research is a fast way to lose the budget. Report what the reader found instead, and keep training and verification as separate line items so the second one survives the first round of cuts.

See the benchmarks

Common questions

Do we still need the vendor courses if the practice sessions are the real program?

Yes, as the access layer, and free options cover it. A free vendor curriculum is a reasonable way to get everyone past the vocabulary and the interface so the supervised sessions are not spent on basics. What it cannot do is teach the occupational judgment about which claim matters in your work, because the course author does not know your work. Use the catalog for the floor and the supervised sessions for everything above it.

How do we report progress to an executive who wants a number?

Give a rate rather than a count, and attach an artifact. Something like: of thirty participants, twenty-two produced a later deliverable where the reader could point to a source check that was absent from the earlier one. That sentence carries a denominator, a threshold someone else could apply, and a document behind it. A completion percentage carries none of those, which is why it never survives the second question.

What if the team is already using AI heavily without any program?

Then the baseline reading matters more, not less. Heavy unsupervised use produces habits that feel productive and leave no trace of checking, and a program arriving on top of that is competing with something people already believe works. Read a current artifact from each person before announcing anything. The gap between what the deliverable shows and what the team reports about itself is the most persuasive material available for why the program exists.

How long should the whole program run?

About six to eight weeks for the practice layer, with the before artifact collected in week zero and the after artifact in the last week. Shorter and there is not enough real work in the window for a habit to attach to. Much longer and the model versions, the tool access and half the participants have changed underneath, so the comparison stops being about the program. Repeat annually rather than extending.

Does an AI policy have to be in place first?

A one-page permission statement does, and a full policy does not. People need to know what they may put into a tool, which tools are approved and who to ask when a case is unclear. That is enough to start practising. A complete governance document with retention rules, vendor terms and incident handling is worth writing, and waiting for it is the most common reason a program slips a quarter without anything being learned.

References

  1. 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that the characteristic AI failure is a confident wrong answer outside the model's capability, and that a prompt-engineering overview did not protect against it.
  2. 2. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers MIT Department of Economics (working paper; later Management Science), 2025. economics.mit.edu Supports the claim that measured AI gains are real, wide-ranged and concentrated among less experienced staff.
  3. 3. The Uneven Impact of Generative AI on Entrepreneurial Performance eScholarship, University of California (Berkeley Haas / Harvard Business School), 2024. escholarship.org Supports the claim that the same assistant can help strong performers and hurt weaker ones, depending on how the advice is judged and used.
  4. 4. The Rapid Adoption of Generative AI (NBER Working Paper 32966) National Bureau of Economic Research, 2025. nber.org Supports the caution that self-reported time savings are small in aggregate and are a weak thing to report as program impact.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.