Teams
Three Failures, Three Deadlines, and Only One Is the Hire's Fault
How long to give a new hire before deciding it is not working depends on which of three failures you suspect. A missing hard skill is readable inside four to six weeks and usually calls for a training decision. A judgment gap is invisible until somebody checks work that looked finished, so schedule that check in month two and decide in month three. A wrong role definition is not a hire failure at all: fix the role the week you notice, or the second hire fails the same way.
The takeHire slow, fire fast is advice with no content, and the other stock answer (document it, coach them, escalate) sets no clock at all. The expensive error is rarely waiting a month too long. It is ending the wrong relationship because nobody separated a training problem from a definition problem, and then writing the same requisition again. Sixty days of deliberate checking costs less than a second bad hire into the same badly drawn role, and it produces a file the first person deserves either way.
Where Olive fits
Open a role and see what the work shows
Olive puts the judgment check before the start date rather than in month two: a role-grounded assignment with an assistant that will do all of it unless the candidate stops it, and a human reviewer writing six findings with the moment in the session behind each. It is an input to a manager's decision and never the decision itself, and the candidate is granted the identical report.
Rank your shortlistHow long do you give a new hire before deciding?
Set it by the failure, not by the calendar: the three common ones have genuinely different discovery times. A hard-skill gap surfaces in four to six weeks with no effort from you. A judgment gap surfaces only when somebody deliberately checks work that looked finished. A role-definition failure surfaces as a person doing well at the job you wrote and badly at the job you needed.
Each carries its own clock:
- Skill gap: decide at week six, and the decision is about training. You can name the missing thing, which is what makes it a skill gap. The question is whether it is teachable inside the runway you have, by someone who has time to teach it.
- Judgment gap: create the evidence in month two, decide in month three. Nothing arrives on its own, so the deadline is whatever date you put on the check.
- Definition failure: decide as soon as you notice, and the decision is about the job. Rewriting the role is cheaper than replacing the person, and replacing the person without rewriting the role buys a repeat.
What all three share is a cost that is smaller than the numbers people quote at each other. A Center for American Progress review of 30 case studies published between 1992 and 2007 put the typical cost of replacing a worker at 21% of annual salary across the 27 cases that set executives and physicians aside, and 16% for positions paying under $30,000 a year, with individual estimates ranging from 5.8% to 213% 4. That is a median across incompatible cost models from a labour market two decades old, and it is not specifically the cost of replacing an early-tenure hire. Quote it as the defensible figure it is, and treat the widely circulated band of 90% to 200% as something to check before repeating.
Which of the three failures is this?
Name it by what you can and cannot point at. A skill gap is the one where you can name the missing thing. A judgment gap is the one where you cannot point at anything wrong and keep finding yourself re-reading their work. A definition problem is the one where the work is good and nobody wanted it. Write your best guess in one sentence before you do anything else, because it changes every subsequent step.
The skill gap is the easy case and it is the one managers over-weight, because it is legible. Somebody cannot write SQL, cannot read a contract, does not know the payer rules. It shows up early, it has a name, and the honest response is a training decision with a date on it. Firing on a nameable, teachable gap is a decision to pay the replacement cost instead of a fortnight of teaching.
The definition failure is the one that gets misattributed most often. The symptom is a hire performing the job description competently while the team's actual need sits somewhere else, usually because the requisition copied the departing person's title while the work moved on. Replacing the person leaves the requisition intact, which is why the second hire fails the same way. The fix is to rewrite the role and have the conversation about whether this person wants the rewritten one.
The judgment gap is the expensive middle case, and it has been getting harder to see. It does not produce visible errors early, because assisted output arrives looking finished, and a manager who is waiting for something obviously wrong is waiting for a signal that has been dampened. It is worth reading what this looks like from the other end: why candidates who demo brilliantly with AI fall apart in month one.
Run the check that makes a judgment gap visible
Hand over one piece of work in month two where the confident generic answer is wrong for your context, say up front that context is the hard part of this one, and see whether it gets caught. This is the only one of the three failures that will not announce itself, so you have to build the check, and building it takes an hour of your time and none of theirs.
What makes a check work:
- The wrong answer is plausible. A trick question tests attention. You want a task where the standard answer is genuinely defensible everywhere except here.
- The local fact is discoverable. It has to be in a document, a ticket history or a colleague's head. If nothing could have told them, you are testing luck.
- You know what right looks like. Decide the answer before you hand it over, in writing, so the grading is not retrospective.
- Ask how they got there. Who they checked with, what they read, where they stopped. A correct answer arrived at by accident is a different result from one arrived at by checking.
The underlying failure has been measured in a cleaner setting. In a field experiment, on one task deliberately chosen to sit outside AI capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer: 84.5% of the control group got it right, against 60% and 70% in the two AI conditions 1. One task, one sample, a 2023 model, and the group given a prompt-engineering overview did worse than the group given none. The transferable finding is not that assistance makes people worse. It is that capable professionals could not tell which side of the capability line the task sat on, which is precisely why the check has to be scheduled.
The interview version of the same question is a good source of tasks, since a loop that failed to settle a capability question usually knows what it wanted to see: why candidates who interviewed brilliantly struggle in their first quarter is the same gap discovered late.
Write three dated observations before you write the decision
A decision built on a month of accumulated impression is both worse and harder to explain than one built on three dated notes, and the notes take ten minutes each. Write what you handed over, what came back, what you checked and what you found. Do it the day it happens. Nothing about this accumulates on its own, and a manager at day 80 with a strong feeling and no record is the ordinary outcome.
The reliability problem is old and well documented. Reviewing why employers resist decision aids, Highhouse reports that the interrater reliability of the traditional unstructured interview is so low that even with a perfectly reliable and valid criterion, interview-based judgments could never account for more than 10% of the variance in job performance, and that a 1996 survey of 201 HR executives rated the unstructured interview more effective than any paper-and-pencil procedure 2. Both numbers are that review reporting other people's work, and the 10% is a ceiling implied by reliability rather than a measured validity. The pattern travels anyway: confidence in an unaided impression runs well ahead of what an unaided impression can carry, and a first-quarter judgment is exactly that.
What you write matters more than that you write. Text-mining the post-interview notes on 7,650 hired candidates at a large Chinese internet technology company, researchers found that the number of job-related capabilities named in the notes tracked later performance and promotions and moved inversely with turnover, at roughly a 2 percent rise in performance per standard deviation of the matching score 3. Small effect, one firm, correlational, range-restricted because only people who were hired could be observed, and drawn from interview notes, a different setting from a manager's performance file. The usable part is the direction: notes that name the capability the role requires carry signal that a rating does not, so write down which capability the work needed and what you saw of it.
Then say it to the person before the meeting where it counts. Every observation in the file should already have been a conversation, which is both the fair version and the defensible one. If the period ends without that having happened, the problem is the process, and it is worth reading what a probation period actually does and does not do before the next start date.
Common questions
Is 90 days the right deadline for deciding a new hire is not working?
It is a reasonable outer bound and a poor default. A nameable skill gap is readable by week six, so waiting to day 90 wastes six weeks of teaching time. A judgment gap will not be visible at day 90 either unless somebody deliberately checked, so the date arrives and the manager has an impression instead of evidence. A badly drawn role is visible the moment you compare the job description to the work in front of the team. Use 90 days as the date the decision must be made, and set the earlier dates that produce the evidence.
What if I cannot tell which of the three failures it is?
Assume it is the role definition until you have checked, because that is the one where acting on the wrong theory costs twice. Take the job description to the two people who work with the hire most and ask what they actually needed in the first quarter. If their answer matches the requisition, you have a person problem, and the next step is to hand over work in month two where the standard answer is wrong for your context. If it does not, you have a definition problem, and no amount of performance management on the individual will touch it.
Should a performance improvement plan run during the first 90 days?
Run one only when it is a genuine attempt to close a specific gap, and the specificity is the test. A plan naming a teachable skill, a teacher and a date is useful at any tenure. A plan written to paper a file before an exit is visible as one to everybody including the employee, and it wastes the weeks it occupies. In a first quarter the more honest instrument is usually the named gap with a check date, which is the same thing without the ceremony. Formal process questions belong with counsel and HR.
How do I check judgment without setting a trap?
Use real work and say what you are doing. Hand over an actual task where the standard answer is wrong for your context, tell the person that context is the hard part of this one, and make the local fact findable in a document or a colleague. That is not a trap, it is the job. A trap is a task nobody could have got right, graded in secret, and it produces a result you cannot act on and a hire who learns that the environment is adversarial.
Does a fast decision protect the team or damage it?
Both, depending on which failure it lands on. Ending a relationship quickly on a nameable skill gap teaches a team that gaps are fatal, and the visible consequence is that people stop admitting them. Ending one late on a judgment gap, after months of work nobody checked, is the more expensive version and it damages trust in a different direction. The speed is not the variable worth optimising. The evidence behind the decision is, and evidence takes the number of weeks it takes.
References
- 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that a judgment gap does not announce itself: 19 percentage points worse on a task outside AI capability, 84.5% control against 60% and 70%, with the prompt-coached group faring worse.
- 2. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the reliability ceiling on unaided judgment (no more than 10% of the variance in job performance) and the documented gap between practitioner confidence and what an unstructured impression can carry, from a 1996 survey of 201 HR executives.
- 3. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company frontiersin.org Supports the claim that notes naming job-relevant capabilities carry signal a rating does not: interviewer post-interview notes on 7,650 hired candidates, at roughly a 2 percent performance rise per standard deviation of the matching score.
- 4. There Are Significant Business Costs to Replacing Employees cdn.americanprogress.org Supports the replacement-cost figure used here: a 21% median of annual salary across the 27 case studies that set executives and physicians aside, 16% under $30,000, with a full range of 5.8% to 213% across all 30 case studies from 1992 to 2007.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.