Teams
New Grads Who Never Worked Without AI: Training Gap or Hiring Mistake?
If your new grad has never worked without AI, start from a training gap, not a hiring mistake. The missing fundamentals (opening the source, recomputing a figure, sizing a number before asking the assistant) are procedures, and procedures land inside a quarter. Two things don't. Shipping a confident wrong answer with no flicker of doubt is a calibration problem, and onboarding rarely installs doubt. A competency the role needed on day one that your loop never tested was a hiring call. Train the first. Screen for the other two.
The takeCalling this a training gap and stopping there flatters the employer. Nobody arrives with doubt already installed. It forms where being wrong is visible to the person who was wrong, and the entry-level seat was where that used to happen, mostly by accident. Payroll data already shows that seat thinning for the youngest cohort in AI-exposed work. I'd expect the firms closing it fastest to find the same deficit in their mid-level bench about five years out, and to call it a talent shortage. What you're annoyed about isn't a generation's flaw. It's the first bill for a training ground nobody agreed to replace.
Where Olive fits
Open a role and see what the work shows
The same six dimensions name what a training plan is aiming at: framing the problem before generating, demanding a source for the claim the recommendation rests on, keeping the judgment that shouldn't be handed over, building something between the brief and the answer, refusing output on stated grounds, and testing a claim against something outside the conversation. Olive reads those from a recorded occupational session and returns six separately evidenced findings, so a diagnosis rests on what someone did rather than on how they describe their own habits.
Rank your shortlistWhich failure are you actually looking at?
Two failures wear the same clothes. In one, the new grad knows the figure ought to be checked and has never been shown how. Nobody told them where the filing lives or what a reconciliation looks like. In the other, nothing about the confident answer suggested checking at all. The first is a training gap and costs you weeks. The second is a hiring signal, and it costs you the quarter.
One afternoon separates them. Hand over a real task from the work you do, an assistant that will complete all of it, source material with one wrong thing in it that only checking catches, and a deadline tight enough that checking everything is unavailable. Then read what happened. "I wasn't sure the segment number was right and didn't know how to settle it" is a gap: the doubt fired and the method was missing. A clean deliverable built on the bad figure, handed over without a flag, is the other thing.
The distinction matters more than it used to because the job changed shape underneath it. Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 first-hand uses of generative AI at work and found the thinking that remains shifting toward information verification, response integration and task stewardship, with higher confidence in the tool associated with less critical thinking, and higher confidence in one's own ability associated with more 1. The part your new grads skipped is now most of the entry-level job.
And the seat itself is being repriced. Employment for 22-to-25-year-olds in the most AI-exposed occupations sits about 19% below where it would have been had it kept pace with less-exposed peers, through reduced hiring rather than separations, with no comparable gap among experienced workers 2. If you intend to keep hiring juniors (and the choice between hiring the skill and training it is a real one), the diagnosis has to be right, because the two failures take completely different money to fix.
What can you train, and how fast?
Most of it, inside a quarter, provided the practice carries a cost. Opening the filing, re-adding a column, running the test against a case the fixture doesn't cover, finding the sample behind a statistic: those are procedures, and a capable graduate picks them up in weeks once somebody names them. What takes longer is the doubt that triggers a procedure, because that only forms where being wrong is visible to the person who was wrong.
It is worth understanding why the habit never got built. In an MIT Media Lab study, 54 people wrote essays across three sessions in three conditions: an LLM, a search engine, and no tools. The LLM group showed the weakest and least distributed brain connectivity of the three and struggled to quote their own work, and when a subset was moved off the tool in a fourth session, engagement stayed low 3. That is a small preprint on a lab essay task, not a workplace result, so treat it as a mechanism rather than a measurement: work that is produced without being constructed leaves less behind.
A training loop that fixes this is four rules, not a course:
- One check per deliverable, named. Pick the claim the recommendation rests on, settle it outside the conversation, and write in the deliverable what it showed. Peripheral facts don't count.
- Something between the brief and the answer. A criteria list, an outline, a back-of-envelope estimate, made first, then the draft read against it.
- One number a week by hand. Estimate before asking, then compare. The gap between the two is the whole lesson.
- Review the check, not the polish. The reviewer's first question is what would have made this wrong, and the second is what would have had to be true to change the recommendation.
Run that for a quarter and the procedures land. What stays stubborn is stance, which is why the split between training the habit and verifying it in hiring never fully collapses into training. A person who does not experience a fluent answer as a claim will do the four rules as compliance and stop the week nobody checks.
Which fundamentals still have to be in the person's head?
The ones a wrong answer travels through without meeting a second reader. That is a field-level judgment and it does not generalize: an engineer can delegate syntax and cannot delegate knowing what a passing test proves; an analyst can delegate formatting and cannot delegate the order of magnitude. Write the list for your own function before you grade anyone against it.
Six versions of the same line, drawn from where the errors actually escape:
- Financial analysis. Keep the order of magnitude and the reconciliation back to the filing. Delegate formatting, first-pass summarization, comparable screens.
- Software engineering. Keep knowing what the test proves and what it doesn't. Delegate boilerplate, API recall, and the first draft of the change.
- Legal operations. Keep reading the paragraph behind the pin cite. Delegate clause drafting and the summary of a long document you will still open.
- Marketing. Keep the sample behind the statistic. Delegate variants, headlines, and the structure of the brief.
- Supply chain. Keep normalizing three quotes onto one set of terms. Delegate the comparison table once the terms agree.
- Healthcare revenue cycle. Keep classifying the denial correctly. Delegate the appeal prose, which is persuasive either way.
Getting this wrong is expensive in both directions, and the two bills look nothing alike. Over-preserve and you teach juniors to redo work the tool does better, pay senior rates for typing, and lose the people who notice. Under-preserve and the wrong thing ships into a place where nobody reads it twice, and a client tells you. The discriminator is not how important the work feels; it is how far a wrong answer travels before somebody competent sees it. See how Olive measures this.
There is a cheap test for whether someone holds the fundamental rather than a memory of it. Ask what would make this answer wrong. Someone who holds it names a condition in a sentence: the segment definition changed, the fixture uses equal values, the survey was the vendor's own customers. Someone who doesn't restates the answer more carefully, which is also what good AI use looks like from the outside and isn't.
When is it a hiring mistake?
When the role needed the competency on day one and your loop never tested for it. The Office of Personnel Management draws that line: a work sample suits a role where the competency is expected on entry, not one where it will be trained after selection 4. So the question isn't whether a new grad arrives with the gap, which most now do, but whether your review loop can carry one for a quarter without a wrong answer escaping.
The second condition is calibration, and it is the one that doesn't train out on your schedule. In a controlled study of AI code assistants, participants with access to the assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure. Participants who were more skeptical of the tool and worked harder at their prompts produced fewer vulnerabilities 5. The failure was confidence, not knowledge.
Nobody can self-report their way out of that, including people far more experienced than your new grads. METR ran a randomized trial with 16 experienced open-source developers on 246 real issues in repositories they knew well. With AI tools allowed, the work took 19% longer, and afterwards the developers still believed the tools had sped them up by 20% 6. If experts inside their own codebase cannot feel a swing that size, an interview question about how someone uses AI collects a story, and your impression of them across a quarter is not much better evidence.
Three signals that it was a hiring call rather than a training one, readable by about day sixty: no unsolicited uncertainty ever reaches you, only answers; the same class of error survives being corrected once; and the deliverable changes when the person is questioned rather than when the evidence changes. One of those is a bad month. All three, after a real correction, is a person who does not experience output as a claim.
How do you stop hiring the same gap twice?
Test the act instead of the account of it. Leave the AI open, hand over a task from the field you're hiring for, plant one wrong thing in the source material that only checking catches, and give less time than a comfortable person would want. Then grade four things from the record: which claim was checked, what it was checked against, whether that happened before or after drafting, and what changed in the answer.
Pilot the task before it decides anything. Two people already doing the job take the same task; one should find the planted problem inside ten minutes, and neither should call it a trick. If both miss it, the problem is unfair rather than diagnostic. Keep the task inside one occupation (hand an engineer a market-sizing packet and you have measured reading comprehension) while the four-column rubric carries across all of them unedited. This is the same instrument as hiring for verification rather than production, pointed at a cohort you already employ.
Then fix the intake that produced the gap. An onboarding built around shipping volume in the first month teaches exactly the behavior you are complaining about, and a review culture that only ever comments on polish teaches it twice. If the first thirty days reward output and nothing rewards a flagged uncertainty, you will reproduce this cohort every year and keep calling it a hiring problem.
And hold the limits honestly. A work sample watches one person for one afternoon; it says nothing about whether they still check in month four with a real deadline on top. The comparison between a junior with AI and a senior without is settled by your review loop as much as by either person. No hiring process reaches month four. The loop does, which is why the loop is the cheaper thing to fix first.
Common questions
How long does it take to train a missing fundamental?
Weeks, when the practice has a cost attached. Procedures (opening the source, recomputing by hand, running a real test) are learnable inside a normal review cycle, and most graduates pick them up faster than managers expect. What takes a quarter or more is the doubt that triggers them, because that forms only where the person sees their own wrong answer land. Set the rules, review the check rather than the polish, and give it one quarter before deciding it was a hiring call.
Should you ban AI for a new grad's first ninety days?
No, as a blanket rule. A ban trains someone for a job that no longer exists, and they will use it anyway once the deadline gets real. Ban it for one step instead: the estimate before the model's number, the first read of the source document, the classification decision that routes everything downstream. That produces the same practice a full ban aims at, keeps the rest of the work current, and gives you something specific to review rather than a rule nobody can verify.
Can you tell the difference in an interview?
Rarely, from questions alone. Asking how someone checks AI output gets you an answer they have rehearsed, and the research is clear that people cannot accurately report their own AI-assisted performance. What does work in a conversation is anchoring it to something they actually made: ask what would have made the answer wrong, what they left unchecked and why, and what they would have done with another hour. Someone who checks names a condition. Someone who doesn't restates the conclusion.
What if the whole cohort has the same gap?
Then it is your intake and your review loop, not the people you picked. A pattern across a cohort is a process fact: the screen tested production, the onboarding rewarded volume, and nothing in the first month made a flagged uncertainty worth raising. Fix the loop before you fix the hiring bar, because a stricter bar applied to the same loop mostly buys you fewer candidates and the same behavior six months later, at a higher salary.
Is a new grad who never worked without AI worse than one who did?
Worse is the wrong frame. Some of them are better at the parts that now carry the work: stating criteria before generating, refusing an answer with a reason, structuring a task so an assistant can be checked. Others produce polished output they cannot defend. The distinction that predicts anything is whether the person treats a fluent answer as a claim, and that has nothing to do with when they graduated. Test for it directly rather than reading it off a birth year.
References
- 1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org Survey of 319 knowledge workers and 936 first-hand examples: the thinking that remains shifts toward information verification, response integration and task stewardship; higher confidence in the tool is associated with less critical thinking, higher confidence in one's own ability with more.
- 2. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence ✓ digitaleconomy.stanford.edu ADP payroll data: employment of 22-to-25-year-olds in AI-exposed occupations sits about 19% below its counterfactual, driven by reduced hiring rather than separations, with no comparable gap for experienced workers.
- 3. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task ✓ arxiv.org 54 participants across three sessions in LLM, search-engine and no-tool conditions: the LLM group showed the weakest, least distributed connectivity and struggled to quote their own work, and showed reduced alpha and beta connectivity when moved off the tool. A preprint on a lab essay task, cited here as mechanism rather than workplace effect.
- 4. Assessment and Selection: Work Samples and Simulations ✓ opm.gov Work samples require performing tasks that mirror the job rather than describing them, and suit roles where the competency is expected on entry rather than trained after selection. Undated guidance, page text verified 2026-08-24.
- 5. Do Users Write More Insecure Code with AI Assistants? ✓ arxiv.org Participants with access to an AI code assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure; participants who were more skeptical of the tool and iterated harder on prompts produced fewer vulnerabilities.
- 6. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues in familiar repositories took 19% longer with AI tools allowed, and still believed afterwards that the tools had sped them up by 20%.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.