Teams
Should You Care That New Hires Ask AI Before the Team?
When a new grad's first stop for questions is the model, worry about the routing, not the loyalty. Asking a model first is rational for anything that fails loudly when wrong: a syntax question, a first draft. It gets expensive where your team's own context is the only check, exactly the context a new hire lacks. Treating it as disloyalty moves the habit out of sight, and a drop in questions proves nothing. Name the questions that belong to a person, then read their first real artifact against your records.
The takeThe habit is a measurement of your team. A model can answer everything except what nobody here wrote down, so the questions that come back wrong draw a map of your undocumented ground: the March deploy, the carve-out, the claim compliance already refused. Nobody has priced how much working knowledge sits in three people's memory, and on the evidence so far nobody is trying. A manager who reads this as a loyalty problem onboards the next five hires the same way. Read as a gap in the written record, it is the rare onboarding complaint whose fix helps everyone already here.
Where Olive fits
Open a role and see what the work shows
The same six dimensions describe the routing this article is about: framing before generating, demanding a source for the claim that matters, keeping the judgment you should not hand over, and testing a claim against something outside the conversation. Olive reads those from a real session on occupational material rather than from a self-assessment, and a person writes all six findings.
Rank your shortlistWhich questions is it safe for a new hire to ask a model?
The ones where checking the answer costs less than getting it. A syntax question, a definition, a first draft of a regex, the shape of a standard filing: if the answer is wrong, the code fails, the search returns nothing, or a two-minute read catches it. Those questions were always cheap, and routing them to a model costs the team nothing it was actually getting.
The dangerous class is the mirror of that. Ask which questions only your team can settle and the list is short and specific: why the retry logic is deliberately wrong in one place, which client's contract has the carve-out, why the previous version of this report got pulled, what the reviewer rejects on sight. A model answers all of those fluently and from nothing.
What that split looks like in the work:
- Software engineering. "How do I write a migration in this ORM" is cheap: it runs or it does not. "Is it safe to backfill this column in place" is not, because the answer lives in traffic patterns and a bad deploy from March that nobody wrote down.
- Financial analysis. A model will define a normalization adjustment correctly. It will not know that this issuer changed segment reporting last year and the prior-year figure in the pack is on the old basis.
- Legal operations. Clause language is cheap. Whether your counsel accepts a mutual indemnity at that cap is precedent, and the precedent sits in a folder and in three people's memory.
- Marketing. Copy structure is cheap. Which claim compliance has already refused twice is not, and the model will write it back confidently.
- Healthcare revenue cycle. Denial-code definitions are cheap. Whether this payer pays on a corrected claim or requires a formal appeal is local, and the model does not know which payers you are contracted with.
So the rule is one line: route by whether the answer can be checked without a person. If the check is cheap, the model is a reasonable first stop and probably a faster one than you. If the only check is somebody's memory of what happened here, that question was never the model's to answer.
Is asking AI first a sign they're avoiding the team?
Rarely, and the data suggests a duller reason. It is a sign the cheapest channel changed. In a questionnaire of 1,000 Americans who had started a new role in the previous twelve months, 44% said they turned to AI first when they got stuck, ahead of asking a colleague at 25% and a manager at 20% 1. The study calls itself non-scientific and exploratory, and its data was collected in October 2025. Read it as a direction, not a measurement.
What changed is the price of the question, not the loyalty behind it. Asking a person has always cost something: an interruption, a moment of looking like you do not know, a wait until they are free. The model costs none of that and it is awake at eleven at night. That trade shows up well beyond your team. After ChatGPT launched, activity on Stack Overflow fell by 25% within six months, measured against Russian and Chinese counterparts where access was limited and against mathematics forums where the model performs poorly, and the authors read their estimate as a lower bound 4. The decline was larger for the most widely used programming languages, and post quality did not measurably drop.
People do still route to a person for the things a person is for. In Stack Overflow's 2025 survey of more than 49,000 developers across 177 countries, asked what would still send them to another human in a future with advanced AI, the top answer was "When I don't trust AI's answers" at 75.3%, ahead of ethical or security concerns at 61.7% and wanting to fully understand something at 61.3% 3. The human channel is getting reserved for judgment, which is the same rule you want your new hire using.
The loyalty read also costs you something concrete. In that same new-hire questionnaire, 66% said they had used an AI-generated idea or response without disclosing where it came from 1. A manager who treats asking the model as a small betrayal gets the same behavior with the receipts removed, and undisclosed use is use nobody can check. This is a different problem from never having learned to work without the tool, which is about missing fundamentals. This one is about a habit that is mostly rational and occasionally very expensive.
Why is this riskier in month one than in year five?
Because verification runs on knowledge a new hire does not have yet. A senior reading a confident wrong answer notices that it contradicts how the billing service actually retries, or that no client on that contract has ever been invoiced that way. Someone in week three reads the same paragraph and finds nothing to disagree with. The answer is plausible, and plausible is all they can currently assess.
The research points the same direction. In a survey of 319 knowledge workers who supplied 936 first-hand examples of using generative AI in real work tasks, higher confidence in the tool was associated with less critical thinking, while higher task-specific self-confidence was associated with more of it 2. A new hire is low on the second and, having watched the model be right about everything they have so far been able to verify, high on the first. That is the worst corner of that chart to be standing in, and it is where onboarding puts people by construction.
The same study names the mechanism without euphemism: 58 of the 319 participants reported barriers to inspecting AI output, including not possessing enough domain knowledge to judge whether an answer was correct 2. One month in, that describes nearly everything about your business.
And the failures that get through are not the obvious ones. In the same 49,000-developer survey, the most-cited frustration with AI tools was "AI solutions that are almost right, but not quite" at 66% 3. Almost-right survives a read by someone who does not yet know better, and it fails later, at the point where it meets your data, your customer or your reviewer. That is the shape behind candidates who demo brilliantly with AI and then struggle in month one, and behind juniors who can produce anything but cannot tell when it is wrong. See how Olive measures this.
What does good routing look like on an ordinary Tuesday?
It looks like a question that arrives already narrowed. Someone routing well does not ask you what a deferred revenue schedule is. They ask whether this contract's ramp counts as a modification, because the model gave two defensible answers and only your team's precedent settles it. The question carries what they tried, what it returned, and the one specific thing they could not check.
Four things worth watching, all of them visible in ordinary work rather than in a conversation about AI:
- The questions narrow. Early questions are broad because everything is unfamiliar. By week six the ones reaching you should be the ones no search resolves: precedent, exceptions, and why a thing is the way it is.
- They bring the model's answer with them. "The assistant says do X, but the runbook says Y" is the best single sign on this list. It means they checked, found a conflict, and brought the conflict to the person who can settle it.
- They tell you what they could not confirm. A caveat attached to their own work ("I could not confirm the restated 2024 figure") is worth more than a clean deliverable, and it is the habit that keeps working as they get more senior.
- The work survives contact with your records. Not whether it reads well. Whether the number ties to the system, whether the client name is the right legal entity, whether the query returns what the dashboard returns.
Two things look like tells and are not. A drop in total questions is ambiguous: it means either that they are routing well or that they have stopped asking anything at all, and only the shape of the remaining questions separates those. And a first draft that reads like a model wrote it says nothing about whether anybody checked it, because polish and checking are unrelated. What good AI use actually looks like is a set of acts, not a writing style.
Write the routing rule in their first week
Say out loud which classes of question belong to a person, and name the person. Two or three lines covers it: anything touching a customer commitment, anything where the team has a precedent the internet does not, and anything about to ship that nobody has read. Everything else is theirs to route however they like. A rule nobody stated is a rule nobody is breaking.
Then make asking cheap enough to compete with a text box. Name one person whose job it is to answer for the first month, and say it in front of the team so the new hire is not spending a favor each time. Put a standing fifteen minutes on the calendar rather than telling them the door is open: an open door still costs a decision to walk through, and a scheduled slot costs nothing.
Seed the check. Most local answers live in three or four places, usually the runbook, the deal folder, the ticket history and whoever was here in 2023, and a new hire who does not know those exist cannot consult them. One afternoon spent walking through where the local record lives converts a whole class of uncheckable questions into cheap ones, which is the only durable version of this fix.
Check at the boundary rather than over their shoulder. The first artifact that reaches a customer, a repository or a filing is the honest test, and it gets read against the local record instead of for polish: does the figure tie to the system, does the clause match the executed version, does the recommendation survive the one number nobody opened. Say in advance that this is the review, so it reads as a standard and not as suspicion.
And ask the routing question directly in the one-to-one. "What did you check yourself this week, and against what" has a good answer available, and the answers improve the moment people know it is coming. It also surfaces the honest version of the problem, which is sometimes that the last person they asked was short with them, and that is something you can fix this week. It settles the older question of whether training is enough or the work has to be checked afterwards too: a routing rule with no check at the boundary is training with nothing attached to it.
Common questions
Should you tell a new hire to ask people instead of AI?
No, and a blanket version of that instruction backfires. Tell them which questions belong to a person: anything touching a customer commitment, anything where the team has a precedent the internet does not, and anything about to ship unread. That is a rule they can apply on a Tuesday. "Ask people first" is not, because most of what they ask a model is genuinely too small to interrupt anyone with, and an instruction they cannot follow honestly just moves the same behavior out of sight.
How do you tell whether a new hire is checking AI answers?
Read one artifact against your own records rather than asking them. Does the figure tie to the system, does the clause match the executed version, does the query return what the dashboard returns. Checking leaves traces: a caveat about what could not be confirmed, a conflict between the assistant and the runbook brought to you unresolved, a number recomputed rather than repeated. A clean, fluent deliverable with no caveats anywhere is not evidence of checking, and it is not evidence against it either.
Is it a problem if a new hire stops asking questions entirely?
It is ambiguous, so look at the shape of what is left rather than the count. Someone routing well asks fewer but harder questions: precedent, exceptions, and the reasons behind decisions no search will surface. Someone who has stopped asking anything is usually not stuck less, they are checking less. The fastest way to tell them apart is to ask what they could not confirm this week. A person routing well has an answer ready; a person who stopped checking has nothing to name.
Should new hires disclose when an answer came from AI?
Ask for something more useful than a disclosure. A note saying what they verified and against what tells you whether the work is sound; a checkbox saying AI was involved tells you nothing and invites people to skip it. In one questionnaire of 1,000 recent starters, 66% said they had used an AI-generated idea or response without disclosing its origin 1, which is what a rule with no purpose behind it produces. Make the verification note part of the deliverable and it stops reading as a confession.
Does asking AI first slow down how fast someone learns the job?
For the cheap questions, no, and it probably speeds it up. The loss is narrower than it feels: the questions that used to travel to a colleague also carried context nobody would have thought to say out loud, and that context does not arrive any other way. So put it back deliberately rather than hoping a habit returns. Walk them through where the local record lives, pair them on the first real artifact, and make one person visibly responsible for answering during the first month.
What if the AI answer is right and the team's way is worse?
That happens, and it is worth more than the compliance you were hoping for. A new hire who brings a conflict between the assistant and the runbook has done the thing you want: they checked, they found the disagreement, and they raised it instead of quietly picking one. Settle it on the merits and say which won and why. If the answer is that the runbook is stale, the routing worked exactly as designed and the runbook was the defect.
References
- 1. New Hires Rely on AI in Their First 90 Days ✓ onlinedegrees.nku.edu Questionnaire of 1,000 Americans who had started a new role in the previous 12 months, data collected October 2025 and described by its authors as non-scientific and exploratory: 44% turned to AI first when stuck, against 25% who asked a colleague and 20% who asked a manager; 66% used an AI-generated idea or response without disclosing its origin.
- 2. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ microsoft.com Survey of 319 knowledge workers supplying 936 first-hand examples: higher confidence in the AI tool is associated with less critical thinking while higher task-specific self-confidence is associated with more; 58 of 319 participants reported barriers to inspecting AI responses, such as not possessing enough domain knowledge.
- 3. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co More than 49,000 responses from 177 countries: the top reason developers would still ask another person for help in a future with advanced AI is "When I don't trust AI's answers" at 75.3%, ahead of ethical or security concerns at 61.7% and wanting to fully understand something at 61.3%; the most-cited frustration with AI tools is "AI solutions that are almost right, but not quite" at 66%.
- 4. Large language models reduce public knowledge sharing on online Q&A platforms ✓ academic.oup.com Stack Overflow activity fell 25% within six months of ChatGPT's release, measured against Russian and Chinese counterparts with limited access and against mathematics forums where the model performs poorly; the authors treat the estimate as a lower bound, the decline was larger for the most widely used programming languages, and post quality showed no significant change.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.