Pipeline
Measuring Quality of Hire Without a Satisfaction Survey
Measure quality of hire with one or two role-specific outcomes, written down before the offer goes out and collected on a schedule the hiring manager does not control. A ramp milestone the team already tracks, a work-quality check a peer performs, and a retention window matched to the role's own cycle each carry more than a satisfaction rating does. Then compare cohorts hired under different processes, because the process is the only thing you can change next quarter.
The takeQuality of hire is the metric talent teams most want and most often produce as a number nobody acts on, and the reason is that the standard recipe was built to be easy to collect rather than hard to fake. A survey of the deciders plus a retention rate can be assembled in a week, and it will move with onboarding, manager turnover and the calendar. If your measure has never once contradicted a hiring manager, it is not measuring the hire.
Where Olive fits
Open a role and see what the work shows
Olive returns six evidenced findings on one candidate, each anchored to a timestamped moment in the session, which gives a cohort comparison something specific to point back at when an outcome disagrees with the loop. It is priced per attempt rather than per seat, with ten attempts a month free.
Rank your shortlistWhy Does the 90-Day Survey Read as Onboarding?
Because ninety days is long enough to see whether somebody was set up well and short enough to hide almost everything else. In most roles the first quarter is training, shadowing and small supervised work, so a rating taken then reports how the onboarding went and how much the manager enjoys working with the person. Both are worth knowing. Neither is the hire.
Who answers is the other problem. The hiring manager chose this person, often over other candidates they also liked, and is now being asked to grade that choice. Grading a choice you made yourself is hard in an ordinary way that has nothing to do with honesty. A randomized trial gives the cleanest illustration available: sixteen experienced developers forecast that AI tooling would cut their task time by 24%, still believed afterwards that it had saved them 20%, and were measured 19% slower 1. Sixteen people in one setting is a small study about a different question, and the magnitude belongs to that setting. Carry the rest: belief and measurement pointed in opposite directions, held by the people best placed to know.
A self-rating and a task-based measure of the same skill can come apart the same way. In a study of 288 teachers who took both a self-report and a knowledge-based test of AI literacy built on the same framework, correlations between the self-reported and objective factors ran from 0.07 to 0.24, with profiles in both directions 2. Teachers are not hiring managers and none of this is a hiring context, so treat it as a caution about the instrument itself.
Pick Two Outcomes Before the Offer Goes Out
Two is the working number: one that shows up early enough to be actionable, one that shows up late enough to be real. Write both into the role's scorecard before the offer, in the same document that named what you were hiring for, so the definition cannot be adjusted later to fit whoever you ended up hiring. Anything defined after the fact is a description, not a measure.
Three families are worth choosing from, and the right pair depends on the role's cycle:
- A ramp milestone the team already tracks. First unsupervised piece of work shipped, first case closed without escalation, first account carried alone. The team is usually recording this anyway, which is what makes it cheap and hard to bend, and how long a new hire should take to get up to speed covers which of the two ramp dates is worth planning against.
- A work-quality check performed by a peer. One sample of real output, read against the same rubric used for everyone at that level, by somebody who did not interview the candidate. It costs about twenty minutes and it is the only one of the three that looks at the work itself.
- A retention window matched to the cycle. Not ninety days for everyone. Match it to the role: a full sales cycle, a full close, a full teaching term. The right window is the point at which somebody who was going to struggle would have.
Avoid promotion speed and performance-review ratings as the primary measure. Both arrive too late to inform a process change and both carry more manager variance than hire variance. Keep them as background context. If the role is one where AI has changed what the first quarter even looks like, structuring the first ninety days of an AI-heavy hire sets the milestones you would then measure against.
Who Collects the Measure, and When?
Somebody with no stake in the answer, on dates fixed before the hire started. A recruiting coordinator, a people-ops analyst, or an automated pull from the system that already holds the milestone. The hiring manager can be a source of fact, naming what happened and when, and should not be the source of the judgment. That single split removes most of the circularity from the measure.
Fix the calendar at the same time as the definition. Collection on a fixed date beats collection when someone remembers, because remembering correlates with how the hire is going. Put the two dates in the same place the requisition lives. Treat a missed collection as missing data and do not backfill it from memory two months later.
Write down the exceptions in advance too. A hire who leaves because the role was cut, whose manager left in month two, or who transferred internally is not evidence about the selection process, and deciding that afterwards is how a cohort quietly becomes whatever you wanted it to be. Name the exclusion rules first, apply them mechanically, and report how many records each rule removed.
One more discipline: keep the interview evidence attached. When a measure eventually disagrees with the loop, the useful question is which stage said what, and that is answerable only if the scorecards are still readable next to the outcome. Whether assessment scores still predict performance once AI is in the workflow is the version of that question worth asking first.
Compare Cohorts, Not People
Group the hires by quarter, by process version and by role family, then read the groups. A quality-of-hire number attached to one individual invites a conversation nobody benefits from and cannot be acted on anyway. The same numbers read as cohorts answer the question you actually have, which is whether the change you made last spring did anything. Score the process, not the person.
Cohorts also make the noise visible. Twelve hires in a quarter will not distinguish a real effect from a run of luck, and looking at them one at a time hides that. Grouping forces the sample size into view and usually reveals that the honest answer is not yet, which is a better place to stand than a confident chart.
What you do with the answer matters more than the precision of it. The selection evidence points toward a process built from two or three different kinds of evidence: a composite of predictors reaches about .61, and taking cognitive ability out of that composite costs .05 3. Those are corrected correlations under assumptions about how predictors are combined, and they are not a hit rate. The composite assumes mechanical combination with sensible weights, which is not what four unstructured conversations stacked on top of each other produce. Read the direction: the fix for a weak signal is usually a different second kind of evidence.
So the cohort comparison feeds one decision, and it is a design decision. When it says a process version is not paying for itself, proving a hiring change actually worked is the method for testing the replacement, and pricing what the mistakes cost is what turns a difference between cohorts into a number a budget owner can use.
Common questions
How long until a quality-of-hire measure says anything useful?
Two or three cohorts, which for most teams is two to four quarters. The constraint is hires per cohort rather than elapsed time: forty hires in a quarter reaches a readable signal faster than eight hires over a year. Declare that up front, along with an interim measure you are willing to name as a proxy, or the pressure to report something will quietly turn the proxy into the headline.
Can the hiring manager be involved at all?
Yes, as a source of fact rather than of judgment. Asking a manager when a new hire first shipped unsupervised work, or whether a specific milestone was met, gets you an observation. Asking whether they are happy with the hire gets you a rating of their own decision. Keep the questions concrete, timestamped and answerable with a date or a yes, and route the interpretation to somebody who was not in the loop.
Is retention a good enough measure on its own?
It is the cheapest one and the most easily misread. Retention captures the worst outcomes and nothing above them, so a team where everyone stays and half the hires underperform scores perfectly. It also tracks the labor market as much as the hire. Use it as one of two measures, with the window matched to the role's cycle rather than a standard ninety days, and never as the only one.
What about quality of hire for internal moves?
Same method, different baseline. Internal hires arrive knowing the organization, so the ramp milestone fires earlier and tells you less. Weight the work-quality check more heavily and lengthen the retention window, since the risk with an internal move is usually a poor role fit surfacing at six to nine months rather than an early exit. Keep internal and external cohorts separate; blending them hides the effect of both.
Should candidates who were rejected be part of the measure?
They cannot be, directly, and that is the known blind spot in every quality-of-hire programme. You only observe outcomes for people you hired, so the measure is silent on strong candidates the process turned away. Some teams sample near-miss rejections and record what happened next where it is publicly visible. That is rough, and it is more than the zero information the standard measure provides.
Does this work for a team hiring fewer than ten people a year?
The definitions do; the cohort comparison does not. At that volume, write the outcomes down before each offer and read them one hire at a time as case evidence, without computing rates. The value is in specifying what good would look like before you know who you hired, which sharpens the scorecard immediately. Aggregate across years or across similar teams once there is enough to aggregate.
References
- 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) arxiv.org Supports the claim that belief and measurement about one's own work can point in opposite directions: a forecast 24% speedup, a believed 20% speedup afterwards, and a measured 19% slowdown.
- 2. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arxiv.org Supports the claim that a self-rating and a task-based measure of the same construct diverge: correlations of 0.07 to 0.24 across 288 teachers taking both instruments.
- 3. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors cambridge.org Supports the claim that combining two or three kinds of evidence beats hunting for one perfect assessment: a predictor composite of about .61, falling by .05 when cognitive ability is removed.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.