Screening
Put the Stage With the Best Evidence First, Not the Cheapest
Sort hiring stages by the evidence each produces per minute of reviewer time, which moves the resume out of first position. The first real gate should generate something a person can argue with: a short scoped task, or a fifteen-minute conversation about work the candidate has actually done. Read the resume afterwards, to decide what to ask, rather than beforehand, to decide who to drop. Keep a genuinely cheap disqualifier at the very top: work authorisation, a licence the role legally requires, a location that cannot flex.
The takeCheapest-filter-first was never a principle about evidence. It was a principle about recruiter minutes, and it held only while the cheapest document to read was also proof that a person had written it. Both halves broke at once. Bolting an automated screen onto position one keeps the order and runs the same sort faster, which is the one change that cannot help here, because throughput at that gate was never the thing that failed.
Where Olive fits
Open a role and see what the work shows
No screen reads a document and tells you who produced it, so Olive assesses the person instead: a 40-to-60-minute assignment grounded in one occupation, done alongside an AI assistant, returned as six findings each carrying the timestamped excerpt behind it. The candidate is given the same report the employer gets.
Rank your shortlistDoes the resume screen still belong first?
Not as a gate. As context, read after something arguable exists, it stays genuinely useful. The order everyone inherited put it first because it was the cheapest thing to read, and cheap meant cheap in recruiter minutes. That trade held while the document was also proof that a person had written it. Nothing about a resume carries that proof now.
What the screen reads is also the part with the least evidence behind it. A meta-analysis of 81 independent samples found prehire work experience, the amount, duration or type a candidate accumulated before joining, correlated .06 with later job performance, .11 with training performance and .00 with turnover, and experience with relevant tasks or occupations did no better at about .07 2. Those are corrected correlations, so weak correction does not explain them away. The finding is narrow in a way worth keeping: it covers the crude experience measures a screen reads at hire. Tenure once someone is in the job is a separate literature, and the finding says nothing at all about licensure or a legally required minimum.
Volume did the rest. Across more than 109 million applications and 247,000 jobs in one large applicant-tracking dataset, the average recruiter is processing 291 applications per hire against roughly 100 in early 2021, and the share of applications resulting in an interview fell from about 7-8% in 2021 to between 3.6% and 4.7% depending on role type 1. That is one vendor's customer base, skewed toward venture-backed technology employers, and the report itself names automated and fraudulent submissions as part of the rise without sizing them. The direction is the usable part: the first gate absorbs far more than it did, and a smaller share of what it absorbs reaches an interview.
Sort stages by the evidence they produce, not by cost per applicant
Sort order comes from one question: how much arguable evidence does this stage return per minute of reviewer time. Answer it stage by stage and the sequence rearranges itself. A fifteen-minute conversation about work the candidate actually did returns more per minute than an hour of resume reading, because the follow-up question is where the evidence sits, and a document cannot be asked one.
A sequence that survives that question usually looks like this:
1. Knockouts. Work authorisation, a licence the role legally requires, a location the role cannot flex on. Seconds per applicant, unambiguous, and defensible on their face. 2. The first real gate. A scoped task under an hour, or a short structured conversation about a piece of work the candidate names. Either produces an artifact or a transcript that a second person can disagree with. 3. The resume, read as context. Now it earns its keep, because it tells you which claims are worth probing in the rounds that follow. 4. The deeper rounds. Whatever the coverage of stages one to three left uncovered, and nothing else.
Skipping the screen entirely is a live option rather than a provocation, and the argument turns on whether your first gate is short enough to absorb the volume: dropping the resume screen and going straight to a work sample is where that trade gets priced properly. It is a large change to run, and the one that actually alters what the process knows.
What belongs at the top of a high-volume funnel?
A genuinely cheap disqualifier, and almost nothing else. Work authorisation, a required licence, a location constraint the role cannot flex on: each is a yes-or-no question that costs seconds to apply and can be explained to the person it removed. That is the job the resume screen was pretending to do. Keep the knockouts, promote the first real gate behind them, and stop asking one document to do both.
Whatever you put at the top of a funnel, that stage is structurally the weakest place in the process. Modelling a screening stage built only from methods that run at volume, with no structured interview available, Berry and colleagues found the validity-maximising composite topped out at .51 with an adverse impact ratio of .37, against .61 once the structured interview was in the mix, and pushing that screen-only battery past a .80 adverse impact ratio dropped it to .39 3. That is a modelled ceiling under stated assumptions with optimal weighting, not an observed result, and resume screening is not one of the methods in the model. The comparative point is what travels: the methods that scale are collectively weaker than the one that does not, so a screening stage can only be asked to remove the clearly ineligible cheaply and honestly.
Which is why replacing a keyword filter with a smarter filter changes very little. What replaces an ATS keyword filter once every resume matches turns on what position one is for, and whether knockout questions still filter anything is the narrower question underneath it.
A smarter filter also carries duties the keyword filter did not. Since July 2023 New York City has barred using an automated employment decision tool on a candidate unless a bias audit ran in the prior year, a summary of it is posted, and the candidate had ten business days' notice 5. That duty attaches where the tool substantially assists or replaces the decision, it binds employers using such a tool on New York City candidates, and other jurisdictions differ, so check yours with counsel.
What moves earlier when you promote a stage?
Its exposure moves with it. Whatever sits at position one produces the largest number of rejections, so promoting a stage promotes its adverse-impact risk and its explanation burden at the same time. A gate that removes most of a pool has to be able to say what each removal rested on, which is a heavier requirement than the resume screen was ever held to.
That is an argument for putting real structure at the new front. The structure is also less common than the vocabulary suggests: in SHRM's benchmarking of its member organisations, with data collected in 2021, 34% used structured interviews to assess executive candidates, 37% for middle management and 36% for individual contributors, while work sample interviews ran at 12%, 11% and 9% of organisations and simulation exercises at 6%, 6% and 7% 4. Those are self-reports against SHRM's own definition, from organisations large enough to have an HR function, and organisations routinely claim more structure than they run, so treat them as an upper bound. Even as an upper bound, a job-shaped exercise remains a minority practice.
Before promoting anything, find out what your current first gate is actually doing. How to tell whether a resume screen is throwing away the wrong people is the measurement that should precede a redesign, because a reordering that moves an unexamined gate to the front just concentrates the same errors.
On Monday, take one open req and count, per stage, how many candidates it removed and how many of those removals someone could defend by pointing at something specific: a missing licence, a task submitted, a claim that did not survive a follow-up question. A stage that removes a great deal and defends nothing is not first because it earned the position. It is first because nobody has moved it.
Common questions
Should we stop reading resumes altogether?
No. Stop reading them as a gate. A resume is a useful map of what to ask about: which employers, which projects, which claims are checkable in ten minutes. Read it after a candidate has cleared a stage that produced evidence, and read it with a question in mind, not a threshold. What it cannot do any more is tell you who wrote it or how much of the described work the candidate personally did, and those were the two jobs the screen was quietly relying on it for.
Does moving the work sample first mean assessing everyone?
Yes, everyone who clears the knockouts, which is why the exercise has to shrink. That is the real constraint, and it sets the design: a first-position exercise has to run in well under an hour, be scoreable in minutes, and be tolerable to send to a large pool. Teams that get this working usually shrink the exercise dramatically rather than assess fewer people. If yours cannot be shortened that far, keep it where it is and put a short structured conversation at the front instead.
What about high-volume, licensed or credential-gated roles?
Those keep a genuinely cheap disqualifier at the top, and it is not the resume. A licence number that can be checked against a public register, a certification that is a legal condition of the work, or work authorisation are all binary, fast and explainable. Apply them as knockouts, then run the same ordering principle on everything below. The distinction that matters is between a filter that removes people who cannot legally do the job and a filter that guesses at quality.
Is an AI screening layer a valid replacement for the resume screen?
No. It occupies the same position and inherits the same problem: reading a document faster does not recover evidence the document no longer carries. Automating the read also brings duties the human read did not, including New York City's bias-audit, posting and ten-business-day notice rules, in force since July 2023 5. Ask what evidence the layer produces that a human reader could not, and get that in writing. Rules vary by state and country, so check the specifics with counsel.
How much does reordering actually change who gets hired?
Enough to be worth measuring, and the measurement is the point, not the promise. Reordering changes which candidates are removed before anyone looks at their work, so the honest test is a stage-by-stage count of removals before and after, split the way your existing reporting allows. Run the new order on one req rather than the whole function, and keep the old sequence running elsewhere so the comparison means something.
References
- 1. Recruiter Productivity | 2026 Talent Trends Report ashbyhq.com Supports the applications-per-hire rise and the fall in the share of applications resulting in an interview, quoted with the dataset's technology skew.
- 2. A meta-analysis of the criterion-related validity of prehire work experience digitalcommons.unf.edu Supports the claim that the experience measures a resume screen reads carry almost no signal about later job performance, training performance or turnover.
- 3. Insights from an Updated Personnel Selection Meta-analytic Matrix: Revisiting General Mental Ability Tests' Role in the Validity-Diversity Tradeoff filiplievens.squarespace.com Supports the modelled ceiling on a screening stage built only from methods that run at volume, and the comparison with a composite that includes a structured interview.
- 4. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) shrm.org Supports the adoption figures for structured interviews, work sample interviews and simulation exercises, cited as self-reported upper bounds.
- 5. Automated Employment Decision Tools: Frequently Asked Questions nyc.gov Supports the New York City bias-audit, posting and ten-business-day notice duties named in the FAQ on AI screening layers, with the substantially-assists scope condition stated.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.