Assessment design

Is AI Training Enough, or Do You Have to Check the Work?

Finishing AI training does not show someone can do the work. A completion rate records attendance, and a survey afterwards fixes nothing: people misjudge their own throughput badly. What transfers depends on the role. Where AI use is procedural, throughput settles it inside a quarter and a separate check adds little. Where the job's AI use is judgment (what to check, what to refuse, what to keep), run a short work sample on your own material four to eight weeks out, assistant on, scored on what the person actually did.

The takeThe uncomfortable question is which working habits a course can actually change. Framing a problem and demanding evidence respond to instruction. Refusal, I'd expect, does not: throwing out a fluent, plausible answer is a disposition. A curriculum hands out prompt patterns; the instinct to distrust one arrives some other way. Until someone measures refusal before and after a course and shows it moving, read the training line as buying access and the hiring bar as buying judgment. Training is still worth buying. The argument is about which line item does which job.

Where Olive fits

Open a role and see what the work shows

If you build the after-training check yourself, the expensive parts are the answer key and the evidence trail rather than the task. Olive is employer-purchased and built for hiring, and the shape it ships is the one this article argues for: twelve authored cases per occupation, and six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session rather than to a number.

Rank your shortlist

Does finishing AI training mean someone can do the work?

No. Completion is an attendance record, and self-report is not much better. In a randomized trial, sixteen experienced developers working real issues in their own repositories took 19% longer with AI tools allowed, and still believed afterwards that the tools had made them 20% faster 2. If practitioners misread their own throughput by that margin, a post-course survey is not evidence of anything.

Completion does tell you three real things, none of which is capability: a person had access to the material, they spent the time, and they could answer questions about the material immediately afterwards. Recall decays. Judgment was never measured in the first place.

The metrics most rollouts report make the gap harder to see rather than easier:

  • Seat time and completion rate. An attendance figure wearing a percentage sign.
  • Satisfaction scores. They track how a session felt, and the sessions that feel best are frequently the ones that asked least of the room.
  • License usage. Volume of use is not a virtue. Someone who ran forty prompts and shipped every answer untouched used AI more than a colleague who ran four and threw out three.
  • Self-reported confidence. The developers above are the clearest available warning, and they were measuring their own work in codebases they knew well 2.

Keep all four on the dashboard, by all means. None of them answers the question being asked in the budget meeting, which is whether the work got better. Whether to train the capability or hire it is a separate decision with its own arithmetic; this article assumes the training already happened and the invoice is already paid.

Why does transfer differ so much between roles?

Because some jobs use AI as procedure and some use it as judgment. A support agent working from a suggested reply is doing procedure, and the gain arrives fast: access to a conversational assistant raised issues resolved per hour by 14% on average, and by 34% among novice and low-skilled agents, with minimal effect on the most experienced 1. Judgment-heavy work moves the other way.

The METR participants were maintainers working on repositories they knew intimately, where most of the value sits in knowing what not to change. AI access made them slower while feeling faster 2. One curriculum, delivered to both populations, would produce a 34% throughput gain in one and a 19% slowdown in the other, and a single completion number would show the same green bar for both.

Where AI sits in the jobWhat training can carryWhat has to be checked afterwards
Drafting from a known patternAccess, prompt shape, house styleLittle; throughput shows it within a quarter
Summarizing a documentWhere the tool fits the workflowWhether the summary was ever compared against the source
Interpreting a numberVery littleWhether the figure was re-derived before it shipped
Taking a position a client acts onEscalation rules and disclosureWhat was refused from the model, and on what grounds

So the transfer rate is a property of the role, not of the training vendor, and no case study from a different company answers it for yours. Before deciding whether a check is worth running, write down what AI is actually doing in each job you trained: drafting, summarizing, interpreting, or committing the firm to a position. What the AI does in the role is the input to that decision; a course catalogue is not.

What does an after-training check have to observe?

Five acts, all of them visible while the work happens: how the problem was framed before anything was generated, what evidence was demanded for the claim the answer rests on, what the person kept rather than handed over, what they refused from the model and why, and what they tested against something outside the conversation. A check that observes none of these is grading output, which is the part that got cheap.

A survey of 319 knowledge workers describing 936 real tasks found the effort does not disappear when a model arrives. It moves: from gathering information toward verifying it, from problem-solving toward integrating a partial response, and from execution toward stewarding a task the model is doing part of 3. Those three destinations are what an after-training check is for.

  • Framing before generating. The first move names what is being decided and what would make an answer wrong. The failure shape is asking for the deliverable in the first message and letting the framing arrive with it.
  • Evidence for the claim that matters. Not a general request for citations, but a source demanded for one load-bearing claim, and opened. The failure shape is confident assertions flowing into the work as facts.
  • The delegation boundary. Something of consequence was kept and done by hand, deliberately. The failure shape is a record in which the person did nothing themselves.
  • A refusal with a reason. A direction was rejected on substance, and the reason is stated. Accepting everything and tidying the edges is the failure shape, and it is the most common one.
  • A test against the world. A figure recomputed, a script run, a page opened, a customer called. Describing a check is not running one.

The last act is field-specific and cannot be borrowed from a generic rubric: an analyst ties a number back to the filing, an engineer runs the code, a paralegal pulls the case. The same survey found that higher confidence in the tool predicted less critical thinking, while higher confidence in one's own ability predicted more 3, so the most enthusiastic graduate of your program is not automatically the one doing the checking. A rubric for AI-assisted work has to score the act rather than the polish, which is the same reason a finished-looking take-home tells you so little.

Build the check as a work sample, not a quiz

A quiz measures recall of what the course said. A work sample measures the thing you paid for. Give one task drawn from the team's own queue, with the assistant available, forty-five to sixty minutes, and a rubric written before anyone sits down. Work samples earn their standing precisely because the person performs the work instead of describing it 4.

Five decisions make or break the exercise:

1. Use real material with a wrong answer in it. The task should contain something a model will assert fluently and incorrectly: a figure that does not tie out, a source that says less than it appears to, a requirement that contradicts another. If nothing can go wrong, nothing can be observed. 2. Leave the assistant switched on. A check that bans AI measures a job nobody is doing any more. The question is not whether the person can work without it. 3. Write the rubric first, in observable terms. "Demanded a source for the revenue claim and opened it" is scorable. "Shows good judgment" is a memory of how the session felt. 4. Capture the working, not just the deliverable. Two people can hand in the same document having done entirely different work. The intermediate artifacts and the exchange with the assistant are the evidence; the deliverable alone is not. 5. Have a second person score a sample of them. Agreement between two scorers on the same session is the cheapest reliability evidence you will ever collect, and disagreement usually means the rubric is describing a feeling.

If you would rather adopt an existing frame than write one, the four Ds rubric is the most widely used starting point and maps onto the five acts above without much translation. Either way, run it once as a diagnostic before it decides anything. Piloting the assessment before it becomes a gate is what separates a measurement from a grievance.

What changes if the check decides anything?

It becomes a selection procedure. The moment a result affects who is promoted, reassigned, or moved onto an AI-heavy team, it sits under the same expectations as a hiring test: administered consistently, related to the job, and explainable if the pattern of who passes turns out to be uneven 5. That is an argument for writing the rubric down, not an argument for skipping the check.

Four things follow, and none of them is expensive:

  • Same task, same time, same rubric, same accommodations. Consistency is most of defensibility, and it costs nothing at the point of design.
  • Keep the evidence. The session, the rubric, and the scorer's notes. A result you cannot reconstruct in a year is a result you cannot defend in a year.
  • Separate diagnosis from consequence. A first run that names which of the five acts is missing is a training input. The same instrument attached to a promotion decision is a test, and it should be announced as one before anyone takes it.
  • Ask any vendor for the same evidence you would have to produce. Validity for the job, adverse-impact analysis, and what the tool actually observes. The questions worth putting to a vendor are the ones you will be asked yourself.

The close is simple. Training buys exposure; the check is what tells you whether exposure became capability, and in which roles it did not. Run the check where the job's AI use is judgment, skip it where the throughput number already answers the question, and write the standard down before the first person sits it rather than after the first argument about a result. If the bar itself is the open question, setting a defensible AI proficiency bar is the piece to settle first.

See what gets scored

Common questions

How long after the training should the check run?

Long enough for real work to have happened, which is usually four to eight weeks. A check run on the last day of the course measures recall, and recall is the one thing the course guarantees. Waiting until the person has taken the new habits into live material is what turns the exercise into evidence about the job. Waiting much past a quarter costs you the attribution: by then, other things have changed too, and a result no longer tells you whether the training did anything.

Isn't a certification exam enough?

A certification exam tells you someone can recognize a correct answer among four options. The behavior that matters is what a person does when no options are listed and the model has already produced something plausible. Multiple-choice items cannot observe a refusal, a source being opened, or a figure being re-derived, because those are acts rather than answers. Certifications are useful as a floor for tool access and vocabulary. They are not a substitute for one task done on real material with the assistant switched on.

What if someone fails the check?

Treat the first run as a diagnostic and say so beforehand. The value of a five-act rubric is that a weak result names which act is missing (framing, evidence, delegation, refusal, or verification), and each one has a different remedy. A person who never demands a source needs a different intervention from one who never keeps any work themselves. If the result will instead affect an assignment or a promotion, announce that before the exercise rather than after it, and apply the same standard to everyone in scope.

Do you have to check everyone who was trained?

No, unless the result carries a consequence. To find out whether the program worked, a sample per role is enough, and stratifying it by function tells you more than testing everyone in one team. The moment a result affects who advances, sampling stops being defensible for that population: a step that decides who moves forward has to reach everyone in scope on the same terms. Decide which of the two you are running before you schedule the first session.

Does this apply to non-technical teams?

Yes, and the five acts (framing, evidence, delegation, refusal, verification) do not change. A recruiter checks a drafted requirement against the actual role and cuts what cannot be justified. A communications lead opens the statistic the model put in the second paragraph. A finance manager re-derives the number before it reaches the board pack. The object being checked belongs to the function, but the act (testing a claim against something outside the conversation, and changing the work when it fails) is identical and just as observable.

How does Olive relate to a post-training check?

Olive is employer-purchased and built for hiring rather than for internal L&D, so it answers the same question one step earlier: whether a person already works this way. The session is a 40-to-60-minute assignment authored for their occupation with an AI assistant available, and a human reviewer writes six findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each attached to a timestamped moment. Outcomes are demonstrated, partly demonstrated or not demonstrated, never a number, and the candidate receives the identical report.

References

  1. 1. Generative AI at Work (NBER Working Paper 31161) Brynjolfsson, Li and Raymond, National Bureau of Economic Research, 2023. nber.org 5,179 customer support agents: access to a conversational assistant raised issues resolved per hour by 14% on average, 34% for novice and low-skilled workers, with minimal effect on the most experienced.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers across 246 randomized real issues took 19% longer with AI tools allowed, and believed afterwards they had been 20% faster.
  3. 3. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks and Wilson (Microsoft Research and Carnegie Mellon, CHI 2025), 2025. microsoft.com 319 knowledge workers described 936 real tasks: effort shifted toward information verification, response integration and task stewardship, and higher confidence in the tool predicted less critical thinking.
  4. 4. Work Samples and Simulations U.S. Office of Personnel Management, 2024. opm.gov Work samples require the person to perform tasks that mirror the job rather than describe how they would do them.
  5. 5. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Any step used to decide who advances is a selection procedure: it must be job-related, applied consistently, and reviewed for adverse impact.

5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.