Teams
Where Juniors Get Their Reps When AI Takes the Grunt Work
Juniors still get their reps when AI takes the grunt work, but only if somebody designs them in. Pick two or three recurring tasks that historically built judgment: reproducing a defect, re-deriving a number, reading the primary source. Make them narrated rather than banned, so the junior does the reasoning and the model does the typing. Put the reps on a calendar, give one senior the teaching as a named responsibility with protected time, then test whether the mental model formed, by handing over an output that is confidently wrong.
The takeEvery plan for this fails at the same line item. The teaching gets assigned to a senior who already has a full week, with no hours removed from anywhere else and no name written against the task, and that is how these programmes get quietly dropped rather than argued down. Whose week the reps come out of is the only genuinely hard decision here; everything else in the design follows from it.
Where Olive fits
Open a role and see what the work shows
An Olive assignment is built so the assistant will do the whole thing if nobody stops it, which makes the person's own reasoning the part a reviewer can see. Each of the six findings names the moment in the session it rests on, so a manager can talk about a specific decision rather than an impression.
Rank your shortlistWhich tasks built judgment, and which only took time?
Separate the two before deciding anything, because the grunt work was never uniform. Reproducing a defect, re-deriving a number from source data, and reading the primary document instead of the summary all built judgment. Reformatting, transcribing, chasing a template into shape and renaming files took time and built nothing. Only the first group is worth protecting, and it is usually three or four tasks rather than a category.
The test for the first group is whether the reasoning is the real output and the artifact is a byproduct. Reproducing a defect teaches how the system actually behaves, which no summary of the defect conveys. Re-deriving a number teaches where the number comes from and therefore which way it moves when an assumption is wrong. Reading the primary source teaches how far a competent summary can drift from what a document says while remaining defensible.
Stanford's Digital Economy Lab has put a number on the backdrop. Using administrative payroll records through June 2026, employment of workers aged 22 to 25 in AI-exposed occupations stood 19% below where it would have been had it kept pace with less-exposed peers, with no comparable gap for experienced workers, and the gap opened through reduced hiring rather than separations 1. The authors are explicit that these are descriptive indicators and not causal estimates of AI. Whatever is driving it, the junior already on the team is the scarcer asset, which changes the arithmetic on whether their training is worth an hour of somebody's week.
Make the reps narrated, not banned
Narrated means the junior does the reasoning out loud and the model does the typing. State the hypothesis before generating. Read the primary source first, then let the model summarise it and mark where the summary drifted. Banning the tool teaches avoidance, which stops being useful the moment they leave the building. Narration keeps most of the speed and keeps the thinking where it has to be.
Three rules make it operational. The junior writes the expected answer, or the shape of it, before anything is generated. The junior opens the source that carries the decision, personally, and can say what it said. The junior states one thing the model got wrong or one thing it left out, every time, and if there is genuinely nothing then that itself gets said out loud.
The reason not to reach for a ban is that juniors are exactly who the tool helps most. Pooling three company-run randomized trials across 4,867 developers, access to an AI coding assistant raised completed tasks by 26.08%, with a standard error of 10.3%, and less experienced developers both adopted it more and gained more 2. The standard error is wide enough that the honest reading is a real gain of uncertain size, the outcome is completed pull requests rather than defect rates, and the tool was code completion rather than an agent. Even discounted, the point stands: a ban takes the most from the people it is meant to protect.
The narrow exception is the small set of tasks a model should not produce at all, where the reasoning is not the only thing at stake. Everywhere else, narration beats prohibition.
Who pays for the slower path?
The team does, on purpose, in hours that have a name on them. One senior takes the teaching as a named responsibility, with time removed from their delivery load rather than added to their week, and it appears in their own objectives where it can be assessed. Unnamed and unfunded, the reps get skipped in the first busy fortnight and nobody notices for a quarter.
The calendar mechanics matter more than the curriculum. Put the reps in a recurring slot rather than promising them. Run one task at a time, four to six weeks each. A programme that covers everything gets announced once and never run. Cap the number of juniors a single teaching senior carries, because the second one costs far more than the first. When the slot collides with a deadline, move it rather than cancelling it, since a cancelled rep does not come back on its own.
Have the cost argument in the open, because it is a real cost and pretending otherwise is what gets these programmes quietly dropped. It is also bounded and visible, which the alternative is not: a team that can produce anything and cannot tell when it is wrong shows up as a series of unexplained incidents rather than as a line in a plan. That is the same problem arriving from the other direction when new grads never learned to work without AI.
How do you know a mental model formed?
Hand them an output that is confidently wrong and watch what happens. Take a real deliverable from their own area, seed one wrong number and one source that does not support what it is cited for, and ask for a decision based on it. Either something happens or nothing does. The answer arrives in under an hour, and nobody has to rate anybody.
A quiz will not do this. Researchers building an AI literacy assessment for a Navy robotics training programme reported that a scenario task simulating AI use on the job outperformed the tests they had adopted from prior research or written themselves, and argued that prevailing assessments emphasise foundational technical knowledge over practical knowledge such as interpreting model outputs 3. One programme, one occupational context, and a preprint reporting no sample size and no effect size, so read it as a design argument rather than a benchmark. The seeded-error exercise is that argument applied to one team.
A pass looks unremarkable. They stop, they name the claim they cannot support, they go and check it, and they come back with a corrected version and one line about what they checked it against. A fail also looks unremarkable, which is the problem: a fluent decision, built on the seeded error, delivered on time, with nothing in it that reads as careless.
Run it once per task rather than once per person. The mental model forms per domain, so someone who catches a seeded error in a defect report may miss one in a financial model six weeks later, and that is information about the next rep rather than about them. If this is hardening into a standing programme, the shape question is apprenticeship, internship or rotation.
Common questions
Is this just banning AI for juniors?
No, and the difference matters. A ban removes the tool; narration removes the shortcut while keeping the tool. A junior on a narrated rep still has the model open, still gets the draft written faster, and still has to state the hypothesis first, read the source personally, and say what the output got wrong. The protection covers the reasoning step on the two or three tasks selected for it, and nothing else.
How long until a junior can use the fast path?
Per task rather than per person, and the seeded-error exercise is the gate. When someone reliably catches a planted mistake in a given kind of work and can say why it was wrong, that task moves off the narrated list for them. Expect that to happen at different times for different tasks, which is why a blanket promotion from junior habits to senior ones tends to be premature in one area and late in another.
What if the senior does not want to teach?
Pick a different senior, and take the refusal seriously rather than negotiating. Teaching badly is worse than not teaching, and a reluctant mentor produces a junior who learns to stop asking. The named responsibility should be a real assignment with hours attached and a place in that person's own objectives, which also makes it a fair thing to decline rather than something absorbed silently.
Does this slow the team down enough to notice?
On the two or three tasks where the junior is required to reason before generating, yes, visibly, and that is the deliberate part. Across the whole workload it is small, because most work is not protected and the model is still in use throughout. The honest framing for a planning conversation is a known cost on a named set of tasks, traded against a capability that otherwise has no other way of being built.
Is hiring fewer juniors the better answer?
That is a separate decision, and it is worth making on its own evidence rather than as a side effect of not wanting to run reps. Entry-level hiring in AI-exposed work has contracted, so the market answer is already visible; what it does not tell you is what your own team looks like in three years without anyone who came up inside it. Decide the hiring question first, then design the training for the juniors you do hire.
References
- 1. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence digitaleconomy.stanford.edu Supports the claim that entry-level employment in AI-exposed occupations has fallen behind, stated as a descriptive hiring slowdown rather than a causal finding.
- 2. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers economics.mit.edu Supports the claim that less experienced staff adopt an AI assistant more and gain more from it, which is why a ban costs juniors most.
- 3. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arxiv.org Supports the claim that a realistic work task reveals applied AI skill better than a knowledge test, offered as a design argument rather than a benchmark.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.