Interviewing
The Manager's Screen Buys the Spec, Not a Second Opinion
A hiring manager's screen should cover the one thing the recruiter cannot: the hard part of the role, in the manager's own words, and what the candidate asks once they hear it. Nobody else holds that input: what the first ninety days actually contain, which decision this person will own alone, and what the last person in the seat found hardest. A second opinion on the candidate is the round's cheap output, and the case for layering one on top of structured evidence is thin.
The takeIf the manager cannot name the hard part, that is the finding, and the requisition is not ready to interview against. It is a more common outcome than anyone admits, and it is the reason loops drift: five interviewers each inventing a bar because nobody wrote one. Twenty minutes spent getting one hard problem out of the manager's head is worth more to the process than a fourth read on the same candidate, and it is reusable across every applicant after this one.
Where Olive fits
Open a role and see what the work shows
A screen can capture a candidate describing how they would check a confident claim; it cannot capture them checking one. Olive puts that in front of them as work, with an assistant that will overreach and a human reviewer who writes what actually happened in the session.
Rank your shortlistWhat should the hiring manager cover?
The hard part of the job, in enough detail that a stranger could recognise it. What the first ninety days contain, which decision this person owns without a second signature, and what the last person in the seat found hardest. Twenty minutes is enough for one of those told properly, and the output is written down before the call ends.
This is the scarce input at this point in the process, and it is scarce for a reason that is new. The job description used to carry it. Now the job description was drafted from the previous one with a model, reads like every other posting for the same title, and has been read back to you by every candidate who prepared. Whatever is genuinely specific about this role currently exists only in the manager's head.
So the round has a deliverable, and it is not a rating. One sentence naming the capability the hard part demands, handed to whoever runs the next round. That sentence is what the later rounds test against, and what a debrief argues over when two interviewers disagree.
A useful side effect: the same twenty minutes is where the material for a work sample comes from. A manager describing what actually goes wrong in the seat is describing an exercise, and finding out what AI actually does in the role is the same extraction pointed at a different part of the job.
Why isn't a second opinion worth the slot?
Because an extra holistic read layered on top of structured evidence can lower accuracy rather than raise it. Highhouse's review reproduces a 1943 study in which high school rank plus a college aptitude test correlated .45 with academic achievement, while the same two predictors plus counselors' intuitive judgment correlated .35 1. That is an admissions study from another era with no sample size on the page reproducing it, and the direction is the part worth carrying.
The same review reports that research on variance in interviewer validity suggests differences between interviewers are due entirely to sampling error, which leaves the best interviewer on a team an unevidenced category 1. That is Highhouse summarising work he does not reprint, so the honest reading is that the evidence fails to support stable interviewer differences. It does not prove none exist. It is still enough to stop building a round on the premise that one does.
Field evidence points the same way. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed job tenures by just over 25%, and comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6% to 7% shorter job durations 2. The limits travel with it: no randomisation, a high-turnover low-skill setting, tenure standing in for quality, and an average override rate of 22%, which makes overriding ordinary behaviour. The instrument was an online questionnaire scored into a green-yellow-red band, and none of it happened in an interview. The finding is about the average override, and it says nothing about whether the test was right on any individual.
None of that says managers should not meet candidates. It says the manager's judgment does more good aimed at the work than aimed at the person, and that a fourth impression arriving before the evidence is collected tends to become the anchor everyone else reads first.
Name the hard part, then watch what they ask
Spend the first ten minutes describing one real problem from the seat, including the part that goes badly, then stop and let the candidate react. A candidate can rehearse answers about themselves. They cannot rehearse a problem they are meeting for the first time, and the shape of what they ask about it is readable: scope, constraint, who decides, what has been tried.
Describe it honestly or the exercise collapses. A sanitised version of the problem produces sanitised questions, and the candidate learns nothing they can use to decide whether they want the job. Name the constraint that makes it hard, the thing that was tried and failed, and the person who has to agree. Managers worry this puts people off. The failure worth worrying about runs the other way: a candidate who accepts the sanitised description and meets the real one in month four.
What to listen for, in rough order of how much it tells you:
1. Questions that change the problem's shape. Asking what happens if the constraint is removed, or who owns the decision, means they are already working on it. 2. Questions about what has been tried. A candidate who assumes the obvious fixes were attempted is used to inheriting real systems. 3. Questions about success. Asking what good looks like in ninety days is fine and common, and it is a weaker signal because every preparation guide suggests it. 4. No questions. Worth one more prompt before you read it, since some people hold questions to the end out of politeness.
Run the same problem past every candidate for the requisition. It is the only way the reactions are comparable, and it turns a conversation into something a second person could read later. If your interviewers judge AI-assisted work differently from each other, that inconsistency shows up here first, and getting managers who do not use AI to judge AI-assisted work is the version of this problem one round further in.
Write the one line the rest of the loop tests
Write the line as a task somebody could set, in the words the manager used for the hard part. What the line is about does more work than the format of whatever tests it later, and the cleanest measurement of that sits in the job knowledge literature Sackett and colleagues drew on: all 164 studies averaged an observed .22, while the 59 using tests built for the job in question averaged .31 3.
Those are subsets of one meta-analysis rather than an experiment, both figures are raw correlations before any correction, and job knowledge tests assume people who already have the knowledge, so the comparison sits in experienced hiring. Reading it across to assessments in general is a stretch, and it is the one this section makes: content has to come from somewhere, and the manager's twenty minutes is where it comes from.
The difference shows up the moment an interviewer tries to build an exercise from the line. "Can hold a scope decision against a founder who wants it shipped" gives them something to build. "Strong communicator" and "analytical" give them a word to interpret. Write the version an exercise could come out of, because one is about to.
The order of the two screens matters as much as their content. The recruiter round tests a claim on the application and produces a record of what the candidate said, which is what a recruiter screen is supposed to decide. The manager round produces the specification everything after it tests against. Run them the other way round and the manager becomes the first opinion in the file, which is the anchor the evidence then has to argue with.
Before the next requisition opens, ask the manager one question and write the answer down verbatim: what will this person spend the hardest week of their first quarter on? If that answer takes more than a minute to arrive, the round after it was never going to have a bar.
Common questions
Should the hiring manager screen come before or after the recruiter screen?
After, in almost every case. The recruiter round tests a claim on the application cheaply and produces a written record; the manager round turns the role's hard part into a specification the rest of the loop tests against. Running the manager first spends the scarcer resource on candidates who have not been screened at all, and it puts the manager's impression at the top of the file, where every later interviewer reads it before forming their own. The exception is a very low-volume senior search, where there is no volume to screen.
What if the manager wants to run a technical screen instead?
A manager-run technical screen duplicates the round after it and costs the specification. It produces a second opinion about skills the next interviewer is about to assess directly, and it uses up the only slot in the process where the hard part of the role gets described out loud. If the loop genuinely needs an early technical filter, give it to someone else and keep the manager's twenty minutes for the work nobody else can do.
How do I stop the manager's read from anchoring the whole loop?
Change what the manager writes down. If the output of the round is a specification rather than a verdict, there is no rating for later interviewers to defer to. Two supports help: have interviewers submit their own evidence before seeing anyone else's notes, and keep the manager's line pointed at the role. The anchoring problem comes from circulating an impression early. Having one is unavoidable.
What does the candidate get out of this round?
An honest description of the hardest part of the job, which is the thing they cannot get from the posting and the thing that decides whether the job is right for them. That makes it the round most likely to lose a candidate, deliberately, and parting there is far cheaper for both sides than parting in month four. It is also the round most likely to win a strong candidate, because a manager who can describe a real problem clearly is the strongest evidence available that the job is well understood.
How long should the manager's screen be?
Twenty to thirty minutes, weighted toward the manager talking, which is the opposite of most first rounds. Ten minutes to describe the problem properly, ten for the candidate's reaction and questions, and a few minutes on what happens next. Longer than that and it starts absorbing the round after it. Shorter and the problem gets summarised into an abstraction, which produces abstract questions and no comparable signal across candidates.
References
- 1. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the Sarbin comparison in which adding an intuitive judgment lowered prediction, and the claim that interviewer-level differences are not evidenced.
- 2. Discretion in Hiring nber.org Supports the finding that managers overriding a structured signal more often were associated with shorter completed job tenures.
- 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range static1.squarespace.com Supports the job-knowledge comparison used to argue that the content of an assessment matters more than its format.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.