Teams
The Tasks Worth Keeping Out of the Model
Deciding which tasks to keep away from AI comes down to three tests, run task by task: is a plausible wrong answer expensive, would the error surface late or never, and is the person doing it still building judgment they will need later. Any two of three, and a model may assist but must not produce: the human writes the draft and uses the model to attack it. Write the resulting list down, because an unwritten exemption is indistinguishable from someone being slow.
The takeMost of these lists get written once, in a meeting, by people naming the tasks they personally hold sacred. That produces a list nobody can apply to a task that was not discussed in the room. Tests travel where a list of examples does not: a new task arrives, three questions get asked, and the answer comes out the same whoever asks them. Write the tests, derive the list from them, and let the list stay short. A long one is a sign the tests were skipped.
Where Olive fits
Open a role and see what the work shows
Delegation boundary is one of the six things an Olive reviewer writes a finding on: whether a person kept the judgment that should not have been handed over, shown by a timestamped moment in the session. That is the same question these three tests ask, applied to one person's work rather than to a task list.
Rank your shortlistWhy is the standard list a security list?
Because the first people forced to write one were legal and information security, and their question was where the data goes rather than whether the answer is right. That list is worth having. It answers a containment question: no client records in a consumer tool, no personal data in a prompt, nothing regulated leaving the tenant. It says nothing at all about the tasks where a model returns something plausible, wrong and expensive.
The two lists overlap only by accident. A competitor summary built from public pages breaks no data rule and can still put a false claim in front of a board. A contract clause read back in confident paraphrase leaks nothing and can be wrong in the one direction that matters. A market assumption at the top of a plan is safe by every security test on the page and carries every number below it.
So the containment list stays where it is, owned by whoever owns it, and a second list gets written next to it. This one is about work quality, and the question it answers is narrower: on which recurring tasks is a fluent wrong answer likely to survive review.
Test each task on cost, discovery time and skill
Three questions, asked task by task rather than tool by tool. Is a plausible wrong answer expensive here. Would the error surface late, or never. Is the person doing this still building judgment they will need later. Any two of the three, and the model may assist but must not produce the thing that ships.
Cost of a plausible error. Not the worst imaginable case: the ordinary one. A wrong number in a board deck is expensive precisely because it is believable enough to be acted on. A wrong number in an exploratory scratch file costs nothing, because nothing depends on it yet.
Time to discovery. Code with a real test suite fails in minutes. A policy interpretation fails in a year, in someone else's hands, with no note attached saying where it came from. The reason this test matters is that people cannot reliably tell which case they are in. In the field experiment with Boston Consulting Group consultants, on one task chosen to sit outside what the model could do, 84.5% of the control group reached the correct answer against 60% and 70% in the two AI conditions 1. The task looked like the others. That is one task in one 2023 experiment, and the frontier has moved since, but the failure mode has not: the tasks that go wrong look like the tasks that go right.
Skill still being built. Whether the tool helps at all appears to depend on the user's ability to judge what comes back. Among 640 Kenyan small-business owners given a business assistant over chat, high performers gained just over 15% while low performers did about 8% worse, a gap the authors trace to which advice owners chose to act on rather than to differences in the advice itself 2. Different setting, different outcome measure, and no detectable average effect at all. The direction is the useful part: protecting the tasks that build discernment is a capability decision, not a nostalgia one.
What does 'assist but not produce' look like on Monday?
The human writes the first version and then points the model at it as an adversary. Same tool, same hour, opposite order. Generate-then-edit produces a document whose spine came from the model and whose corrections are cosmetic. Draft-then-attack produces a document whose spine came from a person and whose weak points have been hunted deliberately.
Four moves make it concrete. Write the frame: what the answer has to contain and what would make it wrong. Write the draft, badly if necessary. Ask the model for the three strongest objections a competent critic would raise. Ask it what evidence that critic would demand, then go and get one piece of it.
The inversion has a real cost, which is why the list has to stay short. Short self-contained writing is exactly where the measured gains concentrate: in a pre-registered experiment, 444 college-educated professionals doing occupation-specific writing finished 37% faster with a chatbot than a control group averaging 27 minutes, and graders scored the output 0.45 standard deviations higher, with the largest gains going to the weakest writers 3. Those were one-off tasks done online for pay, with no colleagues and no consequences, so the numbers do not transfer wholesale. They do establish that the drafting hour is not free to give away.
Which is the trade being made deliberately. On the two or three tasks where a late, plausible error is expensive, the slower order is worth its cost, and everywhere else it is not. If a candidate is being assessed on the same distinction, that is the argument for hiring on verification rather than production.
Write the list down and put a date on it
One page, per team. Task on the left, which of the three tests it failed on the right, and a date at the top. An exemption that lives in a manager's head is indistinguishable from someone being slow, and the person who ends up defending it is whoever is most junior in the room. Writing it down moves the argument from a person to a rule.
Keep the reason attached to each entry, because the reason is what lets someone challenge it. "Rate filings stay human" invites nothing. "Rate filings: a wrong figure is acted on immediately and surfaces at audit" can be argued with, and should be, the first time someone builds a check that closes the gap.
Then set the review date, because the tests are stable and the answers are not. A task fails the discovery-time test until the day someone writes a validation that catches the error in an hour, and then it does not. Quarterly is enough. Anything more often turns into a standing meeting about tooling.
Pair the page with the register of who checks what, since the two answer adjacent halves of the same question: this one decides which tasks a model may not produce, and that one decides who is accountable when AI-assisted work turns out wrong. A team that has written the first and not the second has a rule with nobody attached to it.
Common questions
Where does data protection fit if these tests are about quality?
It stays exactly where it is, as a separate list with a separate owner. Nothing here relaxes a rule about client data, regulated records or personal information; those rules answer where information goes, and they are usually already written. The quality list answers a different question that no security review will ever raise: whether a fluent wrong answer on this task would survive review and get acted on.
Should the list be per team or company-wide?
Per team, with a short company-wide floor if one is genuinely needed. The three tests depend on facts only the team knows: how expensive an ordinary error is, how long it takes to surface, and which tasks are currently teaching someone their trade. A central list written by people who do not do the work tends to name the tasks that sound serious rather than the ones that fail quietly.
What if someone thinks a listed task is fine to automate?
Have them argue with the reason rather than the entry, which is why the reason is written next to it. Most challenges land on the discovery-time test, and most of those are won by building a check: a validation, a reconciliation, a second reader. When a check makes an error surface within hours instead of quarters, the task genuinely changes category and the list should change with it.
Does this mean juniors should not use AI on those tasks?
No. It means the reasoning stays with them while the typing does not. A junior can run the whole task with a model open and still do the thinking, provided they state the hypothesis before generating, read the primary source themselves, and can say what the model got wrong. Banning the tool teaches avoidance; narrating the work teaches judgment, and only one of those is still useful next year.
How often should the list be revisited?
Quarterly, and immediately after any incident where a plausible wrong answer reached a customer or a decision. The tests do not change. What changes is which tasks pass them, as validation improves and as model behaviour shifts under the same prompt. A list with a stale date on it gets ignored quietly, which is worse than no list at all because it looks like a decision was made.
References
- 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that people cannot tell which tasks sit outside a model's capability, so the failure arrives as a confident wrong answer.
- 2. The Uneven Impact of Generative AI on Entrepreneurial Performance escholarship.org Supports the claim that the same assistant helps or hurts depending on the user's ability to judge the advice.
- 3. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) economics.mit.edu Supports the claim that drafting is where measured gains concentrate, which is the cost of inverting the order on a listed task.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.