Screening
Seniority Shows Up in Decisions Under Ambiguity, Not in Titles
When every title says senior, a candidate's real level is the size of the call they were trusted to make when the answer was not available. Ask for three decisions: one they owned end to end, one they escalated and why, and one that went badly and what it cost, then follow up until the answers stop being general. Reporting line and team size are claims inside a document the candidate wrote; a decision history has details that either hold up or do not.
The takeTitle inflation is not a candidate problem, and treating it as one will cost you good people. Titles get handed out instead of raises, invented by companies with eleven employees, and translated badly across industries and countries. Nobody in your pipeline set their own. So the level has to be established in the room, in a conversation the candidate could not have scripted in advance, and the title's job is to tell you which conversation to have.
Where Olive fits
Open a role and see what the work shows
A title travels with a candidate and the record of a working session does not. Olive puts a person in front of a role-grounded assignment with an AI assistant that will overreach, and a human reviewer writes six findings about what they actually did, each one carrying the timestamp it came from.
Rank your shortlistWhat actually separates a senior candidate from a mid-level one?
Level shows in the biggest call they were trusted to make with no safety net, and in what it cost when that call was wrong. That is not on a resume, and years carry less than they look: prehire work experience correlates .06 with later job performance, and task-relevant experience does no better 1. That is a finding about performance, and level is a separate question. Two people with the same title and the same tenure can sit two levels apart.
The standard proxy list does not survive contact with the current market. Reporting line, team size and budget authority are all claims inside a document the candidate controls, and the tell that used to make an inflated claim visible, a sentence that read a little too large for the rest of the page, is gone now that drafting help is universal. The list also misses the thing it is trying to reach. A person can manage a team of twelve and never make a call that could not be reversed by their manager on the same day.
What does separate levels is decision scope under ambiguity. At mid-level, the question is usually available somewhere: an approach exists, a precedent exists, somebody senior will confirm it. At senior, the answer does not exist yet, the tradeoff has a real loser, and somebody has to own the consequence. That is why the most informative question in a screen is about a decision that went wrong. A candidate who owned real scope can price their own mistakes; a candidate who did not will describe a mistake somebody else made.
One stage earlier, the same reading problem arrives as a stack of applications that all look equally polished. What to actually screen on when every resume looks perfect is the top-of-funnel version of the question this article answers in the room.
Ask for three decisions, then keep asking
Three questions, in this order: a decision you owned end to end, a decision you escalated and why, and a decision that went badly and what it cost. Ask them of every candidate for the role, in the same words, and rate the answers on the same scale. The structure does more work here than the wording of any one question.
The numbers back that up. In the 2022 re-analysis of the selection literature, structured interviews estimate at .42 and unstructured interviews at .19, and before any correction the raw observed figures were .32 and .13 2. Structured there is a coding of the interview format: the same questions, rated on a common scale. The US Office of Personnel Management adds a third element to that definition, which is that the interviewers agree in advance what an acceptable answer looks like 3. Structure is a property of how the interview is run, not a script somebody buys.
The follow-ups carry the signal, so plan two per question and do not skip them when the first answer sounds good:
- Who disagreed with you, and what did they say? A real decision had an opponent. An invented one rarely does.
- What did you give up? Every senior call has a loser. A candidate who cannot name it was describing an outcome.
- What would have changed your mind at the time? This separates judgment from hindsight, which is the part that survives preparation.
- What did it cost, in a number or a date? Vagueness here is the most reliable tell in the whole set.
A thin answer to the third question sounds like this: the project shipped late and I learned a lot about scoping. A real one sounds like this: I chose to hold the release for the migration, it slipped eleven days, the sales team lost two renewals in the quarter, and I would make the same call again, because rolling back mid-migration would have cost the data. The second answer is checkable, priced and owned. That is the level.
Write the level definition before the screen, in behaviors
Write two or three sentences per level describing what a person at that level decides without asking, before anybody talks to a candidate. Levels written in years or in headcount cannot be applied to an answer. Levels written in decisions can. Without them, two interviewers hear the same story and file it at different levels, and the debrief turns into a negotiation between confident people.
The evidence on how those readings get combined is blunter than the evidence on how they get gathered. In a meta-analysis of selection and admissions studies, applicant data combined by a stated rule correlated .44 with job performance, against .28 when the same kinds of data were combined by expert judgment, and the authors classify a group consensus meeting as the second kind: what makes a method holistic is that the pieces get combined by judgment rather than by something applied the same way each time 4. That comparison rests on nine studies for job performance. Keep the debrief, which is how evidence gets surfaced and corrected, and agree the combination rule before it starts.
What gets written down after the interview matters too. Text-mining the post-interview notes on 7,650 candidates hired at one large Chinese technology company, researchers found that the number of job-related capabilities an interviewer named in their notes was positively related to later performance and promotions and negatively related to turnover, with a one standard deviation rise in the note-to-job-analysis match corresponding to roughly a 2 percent rise in performance 5. That is correlational, at one firm in one country, and only candidates who passed the interview could be observed. What it supports is notes that name the capabilities the role calls for. It says nothing about whether taking notes at all improves a decision.
The practical version fits on one page. For each level, name what the person decides alone, what they escalate, and what the largest reversible and irreversible calls look like. Then have every interviewer mark which of those they saw evidence for, in the candidate's own examples, before the debrief starts.
Verify the two claims that live outside the room
Scope and reporting line are the two claims a former manager can settle from memory, so put both of them on a reference call. Ask what the candidate owned without needing approval, and how large the biggest call was that they made alone. Use the same wording for every reference on every candidate, and keep notes you can compare side by side.
The validity figure usually quoted for reference checks, .26, comes from the 1998 table, and reference checking was one of eight procedures excluded from the 2022 re-analysis for insufficient information, which makes it at once the only published number and the least checkable one in the set 6. The US Office of Personnel Management's own guidance says reference checks predict better than years of education or job experience but less well than a cognitive ability test, that adding structure improves them, and that written requests produce low response rates and thin information; it budgets about 20 minutes per contact by phone with a minimum of three contacts 7.
So a reference call is a verification instrument, not an evaluation instrument. Scope and reporting line are things somebody else watched happen, which is why a call can settle them. The phrase they were great is reaching for something no call settles. Getting past that phrase is its own skill, worked through in how to reference-check for AI judgment when the reference just says they were great.
What is left after all this is a level you can defend in a sentence: this candidate has owned decisions of this size, with this consequence, confirmed by a person who watched them do it. That sentence beats a title, and it beats a number of years, and it is the one thing in the file that a model did not draft.
Common questions
Should I downgrade a candidate whose title looks inflated?
No, and doing it on the title alone punishes people for their employer's compensation habits. Titles get handed out instead of raises and mean different things at a startup, an agency and a bank. Treat the title as one weak piece of context, then ask for decisions and level the candidate on what they actually owned. The reverse error is the one nobody counts: a person carrying a modest title at a company that levels conservatively can be operating two levels above what the resume says, and a screen that starts from the title will never find out.
How do I ask about a decision that went wrong without putting the candidate on the defensive?
Say what you are looking for before you ask, in one sentence: I am trying to understand the size of the calls you owned, so a decision that did not work out is more useful to me than one that did. That takes about four seconds and it changes the answers you get. Then follow up on cost rather than on blame: what it cost, when you knew, what you did next. Candidates who owned real scope find this question easy. That ease is itself part of the signal.
Can a candidate prepare for decision-history questions with AI?
They can prepare a first answer, and they should. What preparation does not cover is the second and third question, because those depend on what the first answer contained. A prepared story survives who disagreed with you and what did you give up only if there was a real decision underneath it. Treat a polished opening answer as normal rather than as a warning sign, and put your attention on whether the detail holds when you push somewhere the preparation could not have anticipated.
What if the candidate cannot share details because of confidentiality?
Ask for the shape rather than the specifics. A candidate can describe the size of a decision, who had to be convinced, what the tradeoff was and what it cost without naming a client, a number or a product. One useful follow-up: ask them to describe the decision as they would to a peer at another company who cannot be told the details either. If the answer stays abstract even after that, ask about a different decision. Some work genuinely cannot be described at any level, and the candidate is not the person who set that rule.
Is a work sample better than an interview for judging level?
They measure different things and a level judgment usually needs both. A work sample shows current craft on a task you chose, which is strong evidence about capability and weak evidence about scope, because you set the boundaries rather than the candidate. Decision history shows what somebody was trusted with and what happened, which is the part a sample cannot reach. Where they disagree, that is information: a strong sample with no owned decisions behind it often describes a capable individual contributor with a senior title.
References
- 1. A meta-analysis of the criterion-related validity of prehire work experience digitalcommons.unf.edu Supports the .06 correlation between prehire work experience and later job performance, and the near-identical result for experience with relevant tasks, jobs or occupations.
- 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the .42 estimate for structured interviews against .19 for unstructured, and the uncorrected observed figures of .32 and .13.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the third element of the structured-interview definition used here, that the interviewers agree in advance on what an acceptable answer contains, alongside the same questions in the same order and a common rating scale.
- 4. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports the .44 against .28 comparison between mechanical and holistic combination of applicant data, and the classification of a group consensus meeting as a holistic combination method.
- 5. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company frontiersin.org Supports the finding that the number of job-related capabilities named in interviewer notes across 7,650 hired candidates related to later performance, promotions and turnover, at roughly a 2 percent performance rise per standard deviation of match.
- 6. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range static1.squarespace.com Supports the claim that reference checks were one of eight procedures excluded from the 2022 re-analysis for insufficient information, leaving the 1998 estimate of .26 as the only published and least checkable figure.
- 7. Assessment and Selection: Other Assessment Methods - Reference Checking opm.gov Supports the description of a structured phone reference check, its stated position relative to education, experience and cognitive ability tests, the low response rate of written requests, and the 20 minutes per contact with a minimum of three contacts.
7 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.