Screening

The Manager Owns the Bar, the Recruiter Owns the Evidence

Qualified is two decisions, and they have separate owners: the hiring manager owns the bar, which is what the person has to be able to do. The recruiter owns the evidence, which is whether what sits in front of them meets that bar. Escalation to HR is not a third opinion about the candidate: it is a sign that the bar was never written down, so there is nothing for the evidence to be measured against.

The takeThe split is not a hierarchy, and it fails in both directions. A manager who owns the evidence as well as the bar is running a search on taste, and nobody in the loop can tell them they are wrong. A recruiter quietly owning the bar is guessing at a standard they will be blamed for missing. One sentence in the intake document settles it, and it costs nothing to write before a candidate is attached to the argument.

Where Olive fits

Open a role and see what the work shows

Olive puts something in front of that argument both people can read: a role-grounded assignment worked with an AI assistant, written up by a human reviewer as six findings, each carrying the timestamped excerpt it rests on. The report is an input to the decision, and the candidate is granted the identical copy.

Rank your shortlist

Who owns which half of the decision?

The manager owns the bar because they are accountable for the work getting done, and because they are the only person who knows what a wrong answer looks like in this material. The recruiter owns the evidence because they see every candidate while the manager sees the four who reached round two, and because judging evidence against a written standard is a skill the manager practises twice a year and the recruiter practises daily.

That second half is the one companies get wrong, usually by assuming the manager's judgment should also govern the reading. It should not, and there is evidence about what happens when it does. Across 15 firms hiring low-skilled service workers, Hoffman, Kahn and Li found that introducing a job test raised completed job tenures by just over 25%, and that comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6% to 7% shorter job durations 1.

Every limit on that result matters. It is not a randomised experiment. Tenure is the quality measure because turnover is what these employers pay for, so it is a match-quality proxy rather than a performance rating. The instrument is an online questionnaire scored into a green-yellow-red band, from before generative AI, and overrides were common in these firms. The finding is that the average override did worse. Nothing in it says the test was right about any particular person, and nobody should read it as an argument for letting an instrument decide.

Managers are good at saying what the job needs, which is what their own screen with a candidate should be buying. The place accuracy leaked in these firms was the step after that, where an impression got layered on top of a structured signal. That is the argument for keeping the bar and the reading in different hands, and for putting both in writing.

Why calibrating on resumes now agrees about nothing

Two people agreeing about resumes today are substantially agreeing about the same drafting assistants, which is consensus about nothing in particular. The standard exercise is recruiter and manager reading a stack of sample resumes together and settling what strong, average and weak look like, and it worked while a resume was a document a candidate wrote unaided.

Calibrate on something a model did not smooth. In descending order of how much they cost to set up: a work sample from a recent hire and a recent decline, both anonymised. Structured screen notes from the last ten candidates, with the ratings removed. Recordings or transcripts of two screens. What all three share is that the evidence is about what a person did in a specified situation. The reason every resume in the inbox now looks perfect is the same reason the exercise stopped working.

There is an older reason to be careful with a calibration session, which is that agreement is not accuracy. In a review of why employers resist selection decision aids, Highhouse notes that although it is commonly accepted that some interviewers are better than others, the research on variance in interviewer validity suggests the differences are due entirely to sampling error, and reproduces Sarbin's 1943 result in which high school rank plus a college aptitude test correlated .45 with academic achievement while the same two predictors plus counselors' intuitive judgment correlated .35 2.

That is an admissions study built on 1939 admissions data, so nothing in it is about hiring, and the summary about interviewer variance is Highhouse reporting somebody else's work. Read it as the modest claim it supports: your best interviewer is not an evidenced category, and layering a holistic judgment on top of structured evidence can lower accuracy. The output that counts from a calibration session is an edit to the written bar, and a session that ends in two people trusting each other's instincts has not produced one.

Run the five-applicant calibration in an hour

Write the bar first, then both parties independently pass or reject the same five recent applicants against it, then compare. An hour, once, at the start of a search. The bar has to exist before the exercise, or the hour produces a shared preference and calls it a criterion, and a shared preference is exactly what you already had.

The mechanics, because the details are what make it work:

  • Keep it to five. Five produces enough disagreement to be useful and few enough that the conversation stays specific.
  • Recent and real. Applicants from this requisition or the last one for the same role. Invented profiles calibrate people on fiction.
  • Independently, in writing. Pass or reject plus one sentence naming the requirement the decision turned on. Doing it together produces one opinion held by two people.
  • Every disagreement becomes a line in the bar. The exercise produces a criterion. It does not reopen anyone's application.

Disagreement at the boundary is common. In one ATS vendor's data, around 38% of scorecard pairs include at least one point difference between interviewers, and nearly half of those one-point gaps fall between 2 and 3, crossing the yes/no threshold on a 1 to 4 scale 3. Those figures count scorecard pairs, so they are not a rate of interviews or of candidates. The dataset carries no outcome measure, so nothing says which rater was right, and it carries no demographic variable either, so it supports nothing about bias. It supports one thing: ratings are unstable exactly where the decision gets made.

Unstable ratings are the case for a written bar. A pass or reject that names the requirement it turned on can be argued with. A number cannot, which is most of the difference between ranking candidates and passing them against a bar.

What a fight about one candidate is really telling you

A fight about one candidate is almost always the first time the bar has been stated precisely enough for anybody to disagree with it, which is why it is so rarely a disagreement about the candidate. That makes the argument useful and makes the verdict on that candidate the least important thing to come out of it. Get the edit to the bar written before the meeting ends, while both parties still remember what they were disagreeing about.

The recruiter's half is also structurally harder than it looks, and that is worth saying to a manager who thinks a screen should be catching more. Modelling a screening stage built only from methods that can run at volume, with no structured interview available, Berry and colleagues found the validity-maximising composite topped out at .51 against .61 when the structured interview was in the mix 4. That is a modelled ceiling under stated assumptions rather than an observed result, and it does not describe resume screening at all, which is not in the model. The structural point still lands: the methods that scale to the top of a funnel are collectively weaker than the one that does not scale, so a screen will always be passing people the loop later rejects.

One rejected shortlist is a normal week. The escalation worth worrying about is the same rejection three requisitions running with no edit to the written bar in between, which means the bar is living in one person's head and the recruiter is being measured against something they cannot read. That pattern is what actually belongs in front of an HR lead.

Take the last shortlist that got rejected and ask the manager one question: which written criterion did these four fail? If the answer is a description of somebody else, the criterion does not exist yet, and it gets written today. It goes into what an intake meeting is supposed to produce, and it gets tested by the screen that follows, which raises its own question about whether the resume screen still predicts anything.

See a sample report

Common questions

What if the recruiter thinks the manager's bar is unrealistic?

Say so with evidence rather than with an opinion, and say it early. The usable form is a market read plus the pipeline: here is what candidates matching this bar are paid, here is how many exist within the search area, here is what happened on the last two requisitions with a similar bar. That is the recruiter's expertise and the manager cannot supply it. The bar stays the manager's to set, but a bar set without those three facts is a decision made with one eye closed.

Should HR arbitrate when a recruiter and a hiring manager disagree?

Not about a candidate. HR's useful role is upstream and pattern-level: noticing that the same argument has happened on three requisitions, that the written bar has not changed once in that time, and that nothing in the process would reveal which requirement candidates are failing. Arbitrating a specific candidate makes HR a third opinion in a room that already has two too many, and it leaves the cause untouched.

How often should a bar be recalibrated?

At the start of each search, and again after the first five screens on a role nobody has hired for recently. The first five candidates are what reveal which requirement was imaginary, which one everybody was actually screening on without saying so, and which one the market cannot supply at the offered level. After that, recalibrate when something real changes: the level, the team around the role, or what the job now involves day to day.

Can the bar be set by the person currently doing the job?

They should write a large part of it, and they should not own it. The incumbent knows which step should never be handed off and what a wrong answer looks like in the material, and that knowledge is usually missing from the requirement list. What they cannot own is accountability for the hire or the trade-offs against level, budget and team shape. Fifteen minutes with the incumbent, then the manager decides what makes the list.

Does the split change if there is no recruiter at all?

Defining the bar and judging the evidence are still two jobs, so one person holds both, and the risk is that the reading quietly becomes the bar. The cheapest guard is sequence: write the criteria down before reading a single application, then read against them. It is a smaller version of the same discipline, and it is the one thing that keeps a founder-run search from being a search for whoever most resembles the last person they liked.

References

  1. 1. Discretion in Hiring National Bureau of Economic Research, Working Paper 21709 (November 2015, revised September 2017); published 2018 in the Quarterly Journal of Economics, 2017. nber.org Supports the claim that a manager overriding a structured signal is not on average exercising superior private information, at 6% to 7% shorter job durations per standard deviation of exception rate.
  2. 2. Stubborn Reliance on Intuition and Subjectivity in Employee Selection Industrial and Organizational Psychology, 1(3), 333-342, Table 1 (Scott Highhouse), 2008. edbatista.com Supports the claim that a better interviewer is not an evidenced category, and that adding intuitive judgment on top of two mechanical predictors lowered prediction from .45 to .35.
  3. 3. Recruiting Operations Benchmarks | 2026 Talent Trends Report Ashby, 2026. ashbyhq.com Supports the claim that interviewer ratings are unstable at the yes/no threshold, at around 38% of scorecard pairs differing by at least a point, with nearly half of those gaps falling between 2 and 3.
  4. 4. Insights from an Updated Personnel Selection Meta-analytic Matrix: Revisiting General Mental Ability Tests' Role in the Validity-Diversity Tradeoff Journal of Applied Psychology, 109(10), 1611-1634 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2024. filiplievens.squarespace.com Supports the claim that a screening stage built only from methods that scale is structurally weaker than one including a structured interview, .51 against .61 in the modelled composite.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.