Teams
Put the Hiring Manager in the First Pass, Not Only the Shortlist
Both the recruiter and the hiring manager should read the applications, in different jobs. The recruiter owns the pass itself, the pace and the record, because consistency across hundreds of applications is impossible if the reader keeps changing. The hiring manager owns two bounded pieces: writing down the two or three facts that separate a real candidate before screening starts, and reading a sample of the rejections afterwards. That costs a manager about an hour per requisition, and it is the cheapest hour in the process.
The takeThe handoff argument is a distraction. Teams spend meetings on service levels and ownership while the actual failure is that nobody with subject knowledge ever sees what was thrown away. A manager who rejects most of a shortlist is reporting that the criteria upstream are wrong, and the only way anyone finds out which ones is by reading rejects. Twenty rejected applications, read once, is a better use of a manager's hour than any number of calibration conversations about the passes.
Where Olive fits
Open a role and see what the work shows
Olive puts subject knowledge on evidence instead of documents: a role-grounded assignment done with an AI assistant, returned as six findings a hiring manager can argue with because each carries the timestamped excerpt it rests on. The candidate receives the identical report.
Rank your shortlistWhy the standard division of labour stopped working
The standard split stopped working because surface fit became cheap to produce, so the least discriminating filter is now the one narrowing the pool before anyone with subject knowledge sees it. It gives volume to the recruiter and the shortlist to the manager, on the premise that the first pass is pattern-matching a non-expert can do from a job description. That premise held while a document resembling the posting usually came from somebody who had done the work.
The over-narrowing predates all of this, which keeps the argument from being about AI. In the Harvard Business School and Accenture survey of 2,275 executives, 88% agreed that qualified high-skills candidates are vetted out of their own process because they do not match the exact criteria in the job description, rising to 94% for middle-skills roles 1. That figure is self-reported agreement about a tendency in their own process, weaker evidence than a measured exclusion rate, and it was fielded in early 2020 before generative tools were in play. Employers already knew the filter cut too hard. What changed is that the filter also stopped separating the people it let through.
So the shortlist rejection that starts this argument is a symptom with a location. When a manager turns down five of six shortlisted candidates, the recruiter did not read badly. The criteria the recruiter was handed described a document rather than a job, and nothing in the process was ever going to surface that, because the manager only ever saw the output of those criteria and never their casualties.
Moving the manager later, or asking them to read more of the shortlist, does not touch this. It doubles the reading at the stage that is already working.
What the hiring manager should do before screening starts
Write down the two or three facts that separate a real candidate for this role, in terms a non-expert can check without exercising judgement. Not a competency list and not a wish list: the specific things that, if absent, mean this person has not done the job. Twenty minutes with the person currently doing the work produces better criteria than any template, because they know which detail is impossible to fake.
The test for each line is whether two readers who have never done the job would agree on the answer. "Has run a month-end close" passes if the application says where and when. "Strong communicator" fails, because it is scored from prose and every reader means something different by it. A criterion that fails this test still belongs in the process, at a stage where somebody demonstrates it, which is the whole argument in writing screening criteria before the first application opens.
There is evidence for keeping this contribution at the criteria level. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed job tenures by just over 25%, and comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6% to 7% shorter job durations 2. Read the limits: not a randomised experiment, tenure as a match-quality proxy in a high-turnover setting, an average exception rate of 22% that made overrides ordinary, and firms that chose to buy the test. The finding is about averages across managers, and on average a manager overriding a structured signal case by case was not exercising superior private information. Nothing in it establishes that the test was right about any individual person.
Setting the rule is where a manager's judgement pays most, because the rule applies to everyone the same way. Overturning it one candidate at a time is the version the evidence questions.
Read the rejects, not only the passes
Rejects are where a broken criterion shows up, and passes never reveal it. A shortlist that all looks reasonable is consistent with criteria that are excellent and with criteria that are cutting your best applicants for a reason nobody wrote down, and there is no way to tell those apart by reading the survivors. Twenty rejects, once per requisition, distinguishes them in under an hour.
Run it as a blind sample. The recruiter pulls twenty rejected applications at random from the last closed requisition, strips the rejection reason, and the manager marks the ones they would have wanted to meet. Two or three marks is inside the noise you should expect. Five or more is a sign the criteria are selecting for something other than the job, and the fix belongs upstream in the criteria.
What the manager writes next to each mark is the part that pays. "I would have wanted this one because they ran the same migration this team is about to run" is a criterion you can add. "Something about this one" needs another round of questions before it becomes anything. This is also where a manager rejecting people for how a document sounds gets caught early: nobody has shown that guess to be reliable, and it is handled directly in how to stop hiring managers rejecting candidates for sounding like AI.
One scheduling note. Do this on a closed requisition. On a live pile the manager is deciding, the sample stops being a sample, and the exercise quietly turns into the manager screening the whole pool, which is the outcome this design exists to avoid.
Who owns what, and what it costs
The recruiter owns everything that has to be identical across hundreds of applications, and the manager owns everything that requires knowing the work. That line settles most of the disputes teams actually have, and it does not move when volume changes. What changes with volume is only how many people the recruiter's pass is allowed to pass forward.
- Recruiter: the pass itself, the pace, the record of what was decided and why, the reject sample pull, and every candidate communication. Consistency is the product here.
- Hiring manager: the criteria, written before the pool arrives and dated. The reject sample, once per requisition. The final read of whoever reaches the evidence stage.
- Both, together, once: an intake conversation that ends with the criteria written down, which is what separates an intake meeting that produces something from a meeting about the role.
The cost is roughly an hour per requisition: twenty minutes on criteria, thirty on the reject sample, ten on the write-up. Compare that with the alternative price. An hour saved at the screen reappears as interview hours across a panel, and that time is usually the largest real cost of a loop while staying invisible in the accounting. SHRM's cost-per-hire definition sums agency fees, advertising, job fairs, job boards, referral costs, travel, relocation, recruiter pay and talent acquisition systems, divided by the number of hires, and counts no hiring manager or interviewer time at all 4.
Do not solve this by adding readers instead. Even at the interview stage, where the scoring is deliberate, around 38% of scorecard pairs in one large benchmark set differ by at least a point, and nearly half of those one-point differences cross the yes-or-no threshold on a four-point scale 3. That is disagreement rather than error, from a dataset with no outcome measure attached, and it carries no demographic variable, so it says nothing about bias. What it does say is that ratings are unstable exactly where the decision gets made, which is a reason to expect a second reader working without shared written criteria to produce a second unstable rating. Agreement comes from what the readers were told to look for.
Common questions
Our hiring manager has no time. What is the minimum?
Twenty minutes of criteria writing before the requisition opens, and nothing else. If only one piece of this happens, make it that one, because everything downstream is applying criteria somebody wrote, and if the manager did not write them the recruiter guessed them from a job description. The reject sample is the second thing to add, and it can run quarterly across roles rather than per requisition if the manager is genuinely at capacity. Reading the shortlist more carefully is the item to drop, since that stage is already getting expert attention.
Does this mean the recruiter is just executing a rule?
No, and framing it that way loses the harder half of the job. Running a consistent pass across hundreds of applications, keeping a record that survives a question three months later, spotting the moment a criterion has stopped separating people, managing candidate communication and pacing a pipeline against real interview capacity are all judgement, and none of it is judgement about who is good at the job. That is the split. The recruiter owns whether the process is sound; the manager owns what the process is looking for.
How often should we re-check the criteria?
Once per requisition through the reject sample, and immediately whenever a manager turns down most of a shortlist. Criteria decay in a specific way: they are written for a job as it was, and roles now change faster than requisition templates do. A criterion nobody has re-read in a year is usually describing the last person who held the seat rather than the work. Date the criteria file so the age is visible, and treat the first shortlist rejection on a new requisition as a trigger to look at the criteria again.
Should the hiring manager see every application?
Doing that trades one failure mode for its opposite. A manager reading everything is slow, inconsistent between sessions, and produces a record nobody can reconstruct, which is exactly what the recruiter's pass exists to prevent. The manager should see two things: a random sample of rejects, which is small and diagnostic, and everyone who reaches the evidence stage. Between those two, the pass runs without them. Volume handling and subject knowledge are different jobs, and asking one person to do both at scale degrades both.
What if the recruiter and the manager disagree about a candidate?
Resolve it against the written criteria rather than by seniority. If the candidate meets every stated criterion and the manager still wants to cut them, there is an unwritten criterion in play, and the useful move is to name it and decide whether it should be added for everyone. If the candidate misses a criterion and the manager wants them anyway, the criterion may be too strict, and the same question applies. Either way the disagreement is information about the rule. Settling it case by case leaves the rule unchanged and the next disagreement identical.
References
- 1. Hidden Workers: Untapped Talent hbs.edu Supports the claim that employers already agreed their own criteria vet out qualified candidates: 88% for high-skills roles, 94% for middle-skills, from a survey of 2,275 executives fielded in early 2020.
- 2. Discretion in Hiring nber.org Supports keeping the manager's contribution at the criteria level, across 15 firms hiring low-skilled service workers: testing raised completed tenures by just over 25%, and a one standard deviation higher exception rate went with 6% to 7% shorter job durations, at an average exception rate of 22%.
- 3. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the claim that ratings are unstable at the decision boundary: around 38% of scorecard pairs differ by at least a point, with nearly half of those crossing the yes-or-no threshold on a 1 to 4 scale.
- 4. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) shrm.org Supports the claim that the standard cost-per-hire definition leaves out hiring manager and interviewer time: SHRM sums third-party agency fees, advertising, job fairs, job boards, referral costs, applicant and staff travel, relocation, recruiter pay and benefits, and talent acquisition system costs, divided by the number of hires.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.