Teams
The People Best With AI Leave When Their Judgment Stops Mattering
The people who are genuinely good with AI leave when their objections stop changing what ships, and a bigger comp band rarely fixes that. Run one check first: in the last quarter, did anything they flagged or refused change a shipped decision? If not, they are running a check nobody acts on. Keeping them means giving that check teeth: a named point where a deliverable goes back, and a manager who holds the line when the date is tight. That costs a schedule argument, not a raise.
The takeA retention bonus paid to somebody whose judgment gets overridden buys a few months and teaches the whole team what the check is worth. Everyone watches which flag stuck and which one lost to the date, and that is the real signal about whether the work matters. If the answer to whether their objection ever changed anything is no, money just makes it a better-paid version of the same job. Fix the override first, then look at the band if it is still wrong.
Where Olive fits
Open a role and see what the work shows
Turning down an output that does not hold up, and saying what was missing behind it, is one of the six dimensions Olive reports on. It comes back as a written finding with the moment in the session attached, which is the same kind of evidence a manager needs to back a send-back internally.
Rank your shortlistRun the override count before the comp analysis
Count, for the last quarter, how many times this person raised an objection to a deliverable and how many of those objections changed what shipped. Two numbers, an hour of work, and they diagnose the problem better than any survey. A ratio near zero is not a sign the person is wrong. It is a sign the process has nowhere to put them being right.
The numbers are already written down in places nobody reads back: review comments, pull request threads, document comments, the meeting notes where somebody said this figure has no source. Pull a quarter of them and sort into three piles. Objection acted on. Objection overruled with a reason recorded. Objection that nobody answered at all.
The third pile is the one that predicts an exit. An overrule with a reason is a decision somebody made, and a person can live with losing an argument. Silence is the message that the argument was never live.
Before calling the check expensive, price what it catches. In the Boston Consulting Group field experiment, on one task deliberately chosen to sit outside the model's capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer, 84.5% of the control group against 60% and 70% in the two AI conditions 1. One task, one 2023 model, and no basis for a general claim about AI making people worse. The useful part is the failure mode: nobody in the room could tell which side of the line the task was on. The person who can tell is the one whose objections are sitting in the third pile.
Why the compensation answer misses this
Because the market data everyone quotes describes postings, and the person sitting in front of you is not a posting. Lightcast reports that salaries advertised in postings mentioning AI skills run 28% higher than postings that do not, roughly $18,000 more a year, and that over half of the 2024 postings asking for AI skills sat outside IT and computer science 2.
Read that carefully before it becomes a band. It is a raw comparison of advertised salaries between two groups of postings, not a like-for-like wage premium: AI-mentioning postings skew senior, urban and toward higher-paying industries, and it measures what employers advertise rather than what anyone is paid. Lightcast sells skills data, which does not make the figure wrong and does make it interested.
The figure is still useful for one thing. It tells you a market exists, so a person who wants out can find a door. It says nothing about whether this particular person is leaving through it, and the answer to that is almost always sitting in the override count.
Match the market when the market is the problem. Somebody with three approaches a month and a named competitor in play is a comp conversation, and pretending otherwise wastes everyone's time. Somebody who has stopped raising objections is not, and a raise offered there reads as payment for silence.
Give the check a named point and an owner
Name the point in the workflow where a deliverable can be sent back, name who can send it back, and write down what happens when that collides with a date. Most teams have the first, almost none have the third, and the third is the one that decides whether the other two mean anything.
The whole rule fits in four lines:
- The gate. Where in the flow work stops for a check, named as a step, not as a culture.
- The owner. One person, not a committee. A committee cannot hold a line under deadline pressure and everybody knows it.
- The override. Who can overrule the owner, in writing, with the reason recorded at the time and not reconstructed later.
- The read-back. Once a quarter, somebody senior reads the overrides in one sitting. That is the only step that makes the other three real.
The argument that overrides a check is almost always about time, and it is usually made on a belief about how much time AI saved. That belief has been measured and it was wrong in direction, not just size: sixteen experienced developers forecast a 24% speedup, believed afterwards they had gained 20%, and were 19% slower across 246 real tasks 3. Sixteen developers on codebases they knew intimately is a narrow result. It is still enough to stop treating a felt speedup as a reason to skip the step.
Writing the rule down also settles the question that surfaces the first time something wrong ships, which is worth deciding before it does rather than after: who is accountable when an AI-assisted mistake reaches a customer.
What does a stay conversation ask about here?
Ask what they raised that nobody acted on, and ask it that specifically. A general stay interview gets general answers, and this failure mode hides inside the phrase feeling heard. Name a decision, ask what they thought at the time, ask what happened to that view, and then ask what would have had to be true for it to land.
Four questions carry the conversation:
1. In the last quarter, what did you push back on that shipped anyway? 2. When that happened, did anybody tell you why? 3. What is the thing you have stopped raising? 4. If you could send one category of work back without arguing for it, which category?
Question three is the one that matters, and a fast answer to it is a resignation with a notice period attached. Somebody who has already decided which objections are not worth making has finished evaluating the role.
The answers usually point somewhere concrete. A gate in the wrong place, a manager who folds in the last week, a review step that runs after the commitment rather than before it. Occasionally the answer is that the person has outgrown the scope, in which case the retention move is a bigger decision to own, not a better rule: how to tell whether an internal candidate can really move into an AI-heavy role. And if they do leave, the replacement arrives into whatever gate was left in place, which is the thing to settle before deciding whether to hire for AI skills, or train the team you already have.
Common questions
Is a retention bonus ever the right answer?
Yes, when the person is genuinely being pursued and the role itself is fine. The tell is what they talk about unprompted. Somebody describing a competing offer, a title, or a market rate is having a compensation conversation. Somebody describing a decision that went the wrong way is not, and a bonus there buys a quarter and confirms that arguing was pointless. Count first how many of their objections changed a shipped decision; it takes an hour and it tells you which conversation to have.
What if the person is right about the work but hard to work with?
Separate the two and address both. A correct objection delivered badly is still a correct objection, and overruling it to avoid a difficult conversation is how the catch gets lost. Handle the delivery as a coaching problem with specific examples, and keep the gate intact while you do. Teams that fix the tone by quietly routing work around the person end up with neither the tone fixed nor the check running.
How do I tell a real judgment complaint from someone who just says no a lot?
Look at what happened downstream. Objections from strong judgment name a specific claim, name what is missing behind it, and turn out to matter later, sometimes visibly. Reflexive blocking is general, arrives at the same volume regardless of stakes, and leaves nothing behind that can be checked. Read a quarter of the objections against what shipped and the two patterns separate quickly. The read-back is the same exercise a stay conversation needs, which is why it is worth doing on a schedule.
They already resigned. Is anything worth doing?
In the exit conversation, ask whether anything they raised in the last quarter changed a shipped decision, and use the answer on the team instead of on the person. A resignation is usually final by the time it is spoken, and a counter-offer at that point pays somebody to stay inside the conditions they are leaving. What the conversation can still produce is a list of specific moments where a check was ignored, which is the most honest audit of the process available and costs nothing. Fix the gate before hiring the replacement into the same one.
Does this apply on a team where AI is used lightly?
The mechanism is the same wherever verification is somebody's job and speed is somebody else's. What changes with heavier AI use is the volume: more output arrives per week, so more checks run, so overrides accumulate faster and the person doing the checking reaches the conclusion sooner. A team using AI lightly has the same failure available and more time before it lands.
References
- 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that the failure mode worth catching is a confident wrong answer on a task outside the model's capability, measured at 19 percentage points on one such task.
- 2. Beyond the Buzz: Developing the AI Skills Employers Actually Need lightcast.io Supports the claim that employers advertise a premium for AI skills, at 28% and roughly $18,000, and that the figure is a posting-level comparison rather than a controlled wage premium.
- 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) arxiv.org Supports the claim that a felt AI speedup is unreliable grounds for skipping a verification step, with a forecast 24% speedup, a believed 20% gain and a measured 19% slowdown.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.