Screening

How Do You Know If Your Resume Screen Rejects the Wrong People?

Your hires can't tell you whether a resume screen rejects the wrong people: its errors are the people it rejected, and no outcome data on them will ever exist. Three checks work. Advance a random sample from just under the bar and follow it to an outcome the screen cannot have caused. Re-run old screen calls against work you saw later. Audit what the pass pool has converged on. Only the first sees anyone you rejected, and only at the margin, so name the limit where you can't run it.

The takeAn unfalsifiable screen survives because its failures never send anyone a bill. Nearly every other stage in the loop eventually produces a complaint: a bad hire, a blown onsite, an offer declined. The screen produces silence, and silence gets read as accuracy. That asymmetry is why an 80% rejection rate gets quoted as an efficiency win rather than as the size of the blind spot. Nobody has measured what a year of that costs a company, and the people who could measure it are the ones with the least reason to look.

Where Olive fits

Open a role and see what the work shows

The comparison check needs an observed-work signal on people your screen already judged, and ten attempts a month are free, so a sample drawn from under the bar can be assessed beside your current round rather than in place of it. Olive returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification), each written by a human reviewer against a timestamped excerpt, as an input to your decision rather than a filter, with the candidate granted the same report.

Rank your shortlist

Why can't a resume screen be validated on the people you kept?

Because the screen's errors were removed before anyone could observe them. Performance data exists only for people you advanced, interviewed and hired; the 80% you rejected generate no outcome of any kind. Researchers call this the selective labels problem, and a study of resume screening at a Fortune 500 professional-services employer states the constraint flatly: the authors "do not observe hiring outcomes for applicants who are not interviewed" 1.

So "the screen works" nearly always means nobody who got through has caused a problem yet. That is a statement about the top of a range you truncated yourself. It is equally compatible with an excellent screen and with one that has been discarding the strongest applicants in the stack for two years, because both produce the same clean record of people who did fine.

The standard vendor checklist stops short for the same reason. Precision, recall and false-negative rate all need labels, meaning a known right answer per applicant, and the labels are exactly what the screen destroyed. A tool evaluated against your historical advance-or-reject decisions is being measured on agreement with the judgment under test, not on accuracy about the job.

The material the screen reads is thinner than the confidence placed in it. Job experience in years correlates with job performance at .07, corrected for measurement error in the criterion, and years of education at .10, corrected for that plus range restriction: roughly 0.5% and 1% of the variance in job performance. That makes them "two of the least valid predictors of job performance (among commonly used screening criteria)" 2. It is the ground most resume bars are built on, and the same drift sits behind whether GPA still predicts entry-level performance.

Reviewer agreement does not rescue this. Two people applying the same rule tells you the rule got applied twice, and a screen every reviewer agrees on can still be wrong about the same applicants every time. Resumes make even that consistency hard to reach: applicants choose what to put in, so "there is no structured format across applicants" and no guarantee two files were ever read on the same fields 2.

Let a sample through under the bar, then track it

Take the band immediately below your bar, draw a random sample from it, and put those applicants through the same loop everyone else gets. It is the only move that manufactures the data the screen destroyed, and at 80% rejection the volume is small: one in twenty of the near-miss band adds a handful of interviews a month, not a second funnel.

Random is the load-bearing word. An "interesting rejections" pile hand-picked by the same reviewer reproduces the judgment under test and will confirm it. Draw mechanically: every nth application below the line, over a fixed window. Write the rule down before you run it, and blind the downstream stages if you can, because an interviewer who knows a candidate came from below the bar will find the reason, and the exercise then measures the interviewer.

The evidence that this pays is specific, and it is worth stating precisely. Li, Raymond and Bergman simulated screening policies against the historical applications of a Fortune 500 professional-services employer. A model built to explore, meaning one that advances candidates it is uncertain about, would more than double the share of selected applicants who are Black or Hispanic, from 10% to 23%, while the two standard supervised models would cut that share to approximately 2% and 5% 1. Those are counterfactuals computed on the firm's own records, not a program anyone ran.

The authors' own falsification test is the part worth borrowing. If the extra applicants the exploration model selected were truly weaker, it would learn that and stop selecting them; across their test sample it keeps selecting them even as the exploration bonuses fall 1. On the quality side, their decomposition puts average hiring rates among model-selected applicants at 25% and 30%, against the observed 10% among recruiter decisions, on the authors' stated assumption that there is no selection on unobservables 1. That is an estimate carrying a named condition, not a lift you can promise a CFO. What it establishes is that the people a screen passes over are not a settled question.

Measure at a point your own screen cannot contaminate: an offer accepted, a scored work sample, a manager rating at six months. Whether the panel liked them is your own judgment again, so it does not count. Keep the record too. Sampling on your screen's own margin is not selecting on a protected characteristic, but it changes selection rates, and selection rates are what the Uniform Guidelines (US federal, adopted 1978) ask you to document 3. Write the rule, the window and the counts before the first invitation, and put the design in front of employment counsel.

Compare old screen calls against work you later saw

Anywhere a candidate produced observable work after the screen, whether a work sample, a paid trial or a structured exercise, you hold a second reading of the same person the resume never gave you. Pull six months of those, hide the result, have two reviewers re-run the screen decision from the resume alone, then unblind. The disagreements are the finding.

Count both directions and keep them apart. A resume the screen would have rejected, attached to strong observed work, is a false negative you can finally see. A resume that sailed through attached to weak work is the cheaper error, but it is the one that names the criterion doing no work, usually the employer name or the years-of-experience band.

The criterion has to be worth trusting, which rules out most of what sits in an applicant tracking system. The Uniform Guidelines say that whatever criteria are used "should represent important or critical work behavior(s) or work outcomes" 5. A hiring manager's overall impression at ninety days is neither, and it quietly inherits the halo of the screen decision that produced the candidate. A scored exercise everyone took under the same conditions qualifies. So does retention, slowly.

The legal standard names the same evidence the audit needs. Evidence from a criterion-related validity study "should consist of empirical data demonstrating that the selection procedure is predictive of or significantly correlated with important elements of job performance" 4, and the sample subjects "should insofar as feasible be representative of the candidates normally available in the relevant labor market" 5. A validity study run on your hires is not that sample. This check narrows the gap; it does not close it.

The ceiling here is low. This only sees people who reached the work stage, so the applicants your screen cut hardest stay invisible, and it runs on data you already have, which is its entire appeal. It also raises the question of whether to drop the resume screen for a work sample, and it inherits the harder one of whether assessment scores still predict performance once AI is in the workflow.

Audit what your pass pool has quietly converged on

Describe the last hundred people who passed without using the criteria you believe you apply. Which employers, which schools, which title ladders, which resume format, which phrasings. A screen drifts toward whatever reviewers found easy to read, and after a year the pass pool is a profile nobody wrote down and nobody validated. This check needs no new data, just an afternoon.

Tabulate it plainly: prior employer, degree field, years-of-experience band, whether the file echoed your job posting's own vocabulary. Then put the pass pool beside the applicant pool on the same columns. Any attribute that separates the two by a wide margin and appears nowhere in your written bar is the screen's real rule, and it is the one to defend in writing or delete.

Run selection rates by group at the same time, because a resume screen is a selection procedure with obligations of its own. Under 29 CFR 1607.4(D) a selection rate below four-fifths of the highest group's rate "will generally be regarded by the Federal enforcement agencies as evidence of adverse impact" 3. The trap is the bottom line: where the total process shows no adverse impact, the agencies say they "will not expect a user to evaluate the individual components" 3. That is prosecutorial discretion, not a finding that your screen is sound, and a clean overall number can sit on top of a screen doing damage that a later stage compensates for. The Guidelines are US federal and were adopted in 1978, jurisdictions add their own rules on top, so treat this as public legal fact and take the design to counsel. The mechanics on a purchased tool are an adverse impact audit.

Convergence can be legitimate. Unexplained convergence is the thing to chase down: if everyone who passes came from four employers, either those employers genuinely produce the work or the screen learned their formatting, and the first of those is a claim the comparison check can test.

Which role families does your screen rot fastest for?

Look for roles where a resume used to evidence the work and no longer does. A screen is a bet that a document predicts an occupation, and that bet decays at very different speeds. It holds where a licence or a proctored exam makes the claim checkable outside the file, and it goes first where the first year has become judging generated output. Run the checks per role family; a company-wide verdict averages processes with different half-lives.

  • Accounting, audit, nursing. A licence, a board result or a cleared exam section is verifiable outside the document, and the curriculum overlaps the first year's knowledge. The screen keeps its signal here longest.
  • Software engineering. A repository link used to stand for authorship. Side projects now read uniformly finished, and "three years of React" never described the part of the job that is deciding what not to merge.
  • Marketing. A portfolio piece is cheap to produce and expensive to attribute. What separates candidates is what got checked before it shipped, which the artifact does not show.
  • Financial analysis. "Built DCF models" was a low-information line before any of this. With the deliverable now generatable, the screen falls back on tenure at a named employer, which is job experience in years, among the least valid predictors in common use 2.
  • Legal operations, revenue cycle. Named systems, named payers, named jurisdictions still carry information, because a claim that specific can be wrong.

One bar cannot be right across that spread. Where the screen still evidences something, tighten the written criteria and leave it alone. Where it does not, the honest options are to move the evidence earlier in the loop or to stop asking the document to carry it, which is the same argument behind what replaces ATS keyword screening when every resume matches. The revised validity estimates put the predictors specific to individual jobs at the top of the list, with structured interviews at .42 and an 80% credibility interval running .18 to .66, wide enough to read as a caution against trusting any single number too far 6. See how Olive measures this.

Write the verdict down per family, with a date and the evidence behind it. A screen nobody has checked in three years is not a standard. It is an inheritance, and the only thing it reliably predicts is who was easy to read in the year it was written.

See a sample report

Common questions

How many people below the bar do you need to advance?

Enough to see a difference you would act on, which is usually more than feels comfortable. Start with a fixed fraction of the near-miss band, say one in twenty, over a window long enough for outcomes to accumulate, and hold the rule steady instead of adjusting it when early results look bad. Small samples answer one question: does anyone below the bar succeed at a rate close to those above it. If the answer is yes even once, the bar is doing something other than what you think it does.

Is advancing rejected candidates a legal risk?

Sampling on your screen's own margin is not selecting on a protected characteristic, and a documented validation effort is a better position than an unexamined bar. It does change selection rates, and selection rates are what the Uniform Guidelines ask you to record. Write the rule, the window and the counts before the first invitation goes out, keep them with the rest of your applicant-flow records, and take the design to employment counsel first. The Guidelines are US federal and were adopted in 1978; state and city rules sit on top of them.

Can't you just check whether reviewers agree with each other?

Agreement is reliability, not validity. It tells you two people applied the same rule; it says nothing about whether the rule predicts the work, and a screen every reviewer agrees on can be wrong about the same applicants every time. Reliability is still worth measuring, because an unreliable screen cannot be a valid one, but it is a floor and not evidence. Only outcomes speak to validity, which means advancing a sample from below the bar or comparing old decisions against work you observed later.

What outcome should the screen be measured against?

Something the screen cannot have caused. A panel's impression of a candidate it interviewed inherits the screen's judgment, and a manager's overall rating at ninety days inherits it more quietly. The Uniform Guidelines ask for criterion measures representing important or critical work behaviors or outcomes, which in practice means a scored exercise taken under the same conditions, a concrete production or error measure, or retention over a real period. Pick the criterion before you run the check, and write down why it counts as performance.

Does buying an AI resume screener fix this?

No, and it can hide it. A model trained on your past advance-or-reject decisions learns the judgment under test, and precision and recall computed against those decisions measure agreement rather than accuracy. Ask any vendor which outcome their validity evidence uses, whether the study sample included applicants the incumbent process rejected, and what the model does with candidates it is uncertain about. A screener built only to repeat what worked before cannot learn anything about people it never advanced.

What if you genuinely cannot advance anyone below the bar?

Run the two checks that need no extra interviews. Re-run old screen decisions against work you observed later in the loop and count the disagreements in both directions, then audit what the pass pool has converged on against the applicant pool. Neither sees the deepest rejections, and both are worth more than the current evidence, which is none. Label the limit in the write-up: a partial check honestly described is a finding, and a validity claim built only on your hires is not one.

References

  1. 1. Hiring as Exploration (NBER Working Paper 27736) Danielle Li, Lindsey R. Raymond and Peter Bergman, National Bureau of Economic Research, 2020. nber.org Names the selective labels problem in resume screening: "we do not observe hiring outcomes for applicants who are not interviewed". In counterfactual policy simulations on the firm's historical applications, implementing an exploration (UCB) model "would more than double the share of selected applicants who are Black or Hispanic, from 10% to 23%", while static and updating supervised models "would both dramatically decrease Black and Hispanic representation, to approximately 2% and 5%, respectively". The paper's own falsification argument: "if the additional minority applicants selected by the UCB algorithm were truly weaker, the model would update and learn to select fewer such applicants over time. Instead... the UCB model continues to select more minority applicants... even as exploration bonuses fall." A decomposition puts "average hiring rates among those selected by the UCB and updating SL models" at "25% and 30%, respectively, compared with the observed 10% among observed recruiter decisions", and states "This approach assumes that there is no selection on unobservables." Data is "professional services recruiting within a Fortune 500 firm".
  2. 2. Resumes vs. application forms: Why the stubborn reliance on resumes? Risavy, Robie, Fisher and Rasheed, Frontiers in Psychology, 2022. pmc.ncbi.nlm.nih.gov Job experience in years validity .07 (corrected for criterion unreliability; about 0.5% of variance in job performance) and years of education .10 (corrected for criterion unreliability and range restriction; about 1% of variance), called "two of the least valid predictors of job performance (among commonly used screening criteria)". Also that applicants choose what to include, so "there is no structured format across applicants".
  3. 3. 29 CFR 1607.4 - Information on impact (Uniform Guidelines on Employee Selection Procedures) U.S. Equal Employment Opportunity Commission, via Cornell Legal Information Institute, 1978. law.cornell.edu The four-fifths rule: a selection rate below four-fifths of the highest group's rate "will generally be regarded by the Federal enforcement agencies as evidence of adverse impact". Also the bottom-line provision, under which agencies "will not expect a user to evaluate the individual components" where the total selection process shows no adverse impact.
  4. 4. 29 CFR 1607.5 - General standards for validity studies (Uniform Guidelines on Employee Selection Procedures) U.S. Equal Employment Opportunity Commission, via Cornell Legal Information Institute, 1978. law.cornell.edu Criterion-related validity evidence "should consist of empirical data demonstrating that the selection procedure is predictive of or significantly correlated with important elements of job performance".
  5. 5. 29 CFR 1607.14 - Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) U.S. Equal Employment Opportunity Commission, via Cornell Legal Information Institute, 1978. law.cornell.edu Criterion measures must represent "important or critical work behavior(s) or work outcomes", and the study sample "should insofar as feasible be representative of the candidates normally available in the relevant labor market".
  6. 6. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology (Cambridge University Press), 2023. doi.org "The predictors at the top of our list in terms of criterion-related validity are those specific to individual jobs"; structured interviews carry a mean validity of .42 with an 80% credibility interval of .18 to .66.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.