Teams
Prompt Engineering Is Not the Skill Your Team Is Missing
Skip the prompt engineering course and buy supervised practice on the team's own work instead. Three habits belong in it: stating the constraint and what a wrong answer would look like before generating, demanding a source for the one claim that carries the decision, and testing an output against something outside the conversation. Those survive a model release. A prompt library is stale on the next one.
The takeThe course is not a fraud. It is answering a question the team has outgrown. Asking for prompt training is usually a request for permission and a request for protected time, dressed as a request for syntax, and the syntax is the only part of it anyone can sell by the seat. Grant the permission, protect the time, and put a senior reader at the end of it. What comes back is a set of habits that hold when the model changes underneath them.
Where Olive fits
Open a role and see what the work shows
The six dimensions Olive reports on are the ones that outlast a model release: how the problem was framed, where the evidence came from, what stayed out of the assistant's hands, how the work was structured, what was rejected, and what was verified. A human reviewer writes each finding from what happened in the session.
Rank your shortlistWhat part of prompting still matters?
The specification, not the phrasing. Saying what you want, what constraints apply, and what a wrong answer would look like is work no model release removes, because the model cannot know any of it. Choosing the wording that coaxes a better response out of a particular version is work each release absorbs a little more of.
The AI Fluency Framework splits it the same way. Written by Rick Dakan and Joseph Feller and turned into courses with Anthropic, it names four competencies: Delegation, Description, Discernment and Diligence 1. Description is the one a prompt course teaches. The other three are about deciding what to hand over, judging what came back, and taking responsibility for the result.
Under Diligence the framework lists Deployment Diligence, defined as taking responsibility for verifying and vouching for AI-assisted outputs, including fact-checking, testing for accuracy and validating claims 2. That item is the one summaries most often drop, and the one an employer cares about most. The framework is a taxonomy with no levels and no scoring anchors attached, so it will not grade anyone. It is still a better map of the territory than a syllabus of prompt patterns.
A team that phrases requests beautifully and vouches for whatever comes back has the wrong half of the skill.
Why does a syntax course go stale?
Because its content is indexed to a model version and its examples are indexed to a product. Temperature settings, role preambles, few-shot scaffolds and the particular incantations that worked around a weakness are all answers to how one system behaved at one time. Specification and verification are answers to what the work requires, which does not change when the vendor ships.
A sharper reason to be careful comes from the Boston Consulting Group field experiment. On a task deliberately placed outside the model's capability, consultants using GPT-4 did worse than the control group. The two AI conditions did not fail equally: the group given a prompt-engineering overview was down 24 points against control, and the group given no coaching at all was down 13 3.
One task, one sample, a 2023 model, and the paper's own framing is that people could not tell which side of the capability line they were on. It is not evidence that prompt instruction makes people worse in general. It is a reason to stop assuming that instruction which raises fluency also raises caution, because on this task it moved the other way.
The same pattern shows up in how people read their own performance. Sixteen experienced developers in a randomized trial forecast a 24% speedup from AI tooling, believed afterwards that they had got a 20% speedup, and were measured 19% slower 4. A small, specific setting, and still the cleanest available demonstration that comfort with a tool and effectiveness with it are separate quantities.
What should a half day of practice contain?
Three exercises on live deliverables, run in the order the work happens. Each one is thirty to forty minutes, on something the person was going to produce anyway, with a facilitator whose only job is to stop the room at the right moment. The output is a real deliverable plus a short written trail, not a certificate.
The three exercises, in order.
1. Write the brief before the prompt. Ten minutes, no tool open. The constraint, the audience, the format, and one sentence describing what a wrong answer would look like. Then work as normal. Compare the finished piece against the sentence. 2. Name the load-bearing claim and go and check it. Every deliverable rests on something. Find it, leave the chat window, confirm it against a source or a system of record, and write down what was confirmed. 3. Reject something and say why. Take the first draft the assistant produced and mark what does not survive. A session where nothing is discarded means the assistant is being used as a transcription service.
Then four supervised sessions over the following weeks, on real work, with the same three moves expected each time. That is the whole curriculum. It has no vendor, no slide deck and no expiry date, and it produces a stack of artifacts a manager can read.
If the team wants a definition of the target before they start, what AI fluency means and how to test for it covers the ground without the vocabulary a course would sell them.
How do you tell whether it worked?
Read two deliverables from the same person, one from before and one from after, and ask whether a reader can point at the difference. Not a completion rate, not a satisfaction score, not a self-rating. Three questions decide it: is there a stated brief written before generation, was a load-bearing claim checked outside the conversation, and was something rejected with a reason attached.
That reading takes a senior practitioner about two hours per person across a whole program, and it is the only part of this that cannot be bought. It is also the same reading that anchors a wider upskilling program, so where one already exists this practice belongs inside it. Budget it explicitly, because it is the line that disappears when the quarter tightens and the courses do not.
Two failure modes worth watching for. The first is speed becoming the proxy: work coming back faster is easy to measure and says nothing about whether the checking happened, and measuring speed against judgment is a choice worth making deliberately. The second is polish becoming the proxy: AI-assisted drafts read well by construction, so a reader grading on prose quality is grading the tool. What good AI use actually looks like describes behaviour, which is what the three questions ask for.
One more thing to settle before the first session, because it decides how seriously anyone takes the rest. Say who owns the result when an AI-assisted deliverable is wrong. If the answer is unclear, the verification habit has no consequence attached and it will not stick, and who is accountable for an AI mistake is worth answering in writing while nothing has gone wrong yet.
Common questions
Is prompt engineering still a real skill?
Parts of it are, and the parts split cleanly. Describing a task precisely, giving the model the context it cannot infer, and structuring a long piece of work into steps are professional habits that transfer across tools and versions. Memorised phrasings, jailbreak-adjacent workarounds and settings tuned to one model are tooling details with a short shelf life. Teach the first group as part of how the work is done. Do not build a curriculum on the second.
Our engineers want a course and our marketers want one too. Same course?
Same three habits, different cases. The framing exercise, the source check and the rejection work identically in both functions, but the load-bearing claim in a deployment plan and the load-bearing claim in a campaign brief fail in different ways and cost different amounts. Run the sessions per function with that function's own live deliverables, and keep the facilitator's questions constant across them. That way the results are comparable without the content being generic.
What if the team has already done a prompt course and wants more?
Move them to supervised practice rather than a second course. A team that has done the syntax material has the vocabulary and usually not the habit, because a course cannot supply the moment where a real deadline meets a plausible wrong answer. Run the framing, source-check and rejection exercises on their live work. If they resist, the useful question is what they expected the first course to change and whether anything did.
How much should half a day of practice cost?
Almost nothing in cash and a real amount in salary. There is no licence to buy, no vendor and no exam fee. The cost is thirty people times half a day, plus a facilitator, plus roughly two hours of senior reading per participant across the program. Cost it at loaded salary and compare it against a seat price times thirty. The comparison is usually not close, and the practice version leaves artifacts behind.
Do we need a specialist facilitator?
No. The facilitator needs occupational credibility, not AI credentials, because the job is to ask how anyone would know a claim is right at the moment it appears. A senior practitioner from the same function does this better than an external trainer, who cannot tell which claim in that deliverable carries the decision. Give them the framing, source-check and rejection exercises and the three questions used to read the results. That is the entire brief.
References
- 1. Framework for AI Fluency ringling.libguides.com Supports the claim that description is one of four competencies, alongside delegation, discernment and diligence.
- 2. Framework for AI Fluency ringling.libguides.com Supports the definition of Deployment Diligence as verifying and vouching for AI-assisted output, and the caveat that the framework is a taxonomy with no scoring anchors.
- 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that on one out-of-frontier task the group given a prompt-engineering overview did worse than the group given none.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) arxiv.org Supports the claim that comfort with an AI tool and measured effectiveness with it are separate quantities.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.