Screening

False Positives, False Negatives, and Which One You Pay For

Nearly every hiring process is tuned to avoid the bad hire rather than the missed one, rarely because anyone decided it. Write down which of the two errors yours is set against, then check that against the role's economics: a mistake you can recover from inside a quarter should be tuned differently from one you cannot. Then buy visibility into the half you never see, by sampling near-miss rejections or by replacing one impression-based stage with one that produces evidence, which shrinks both error rates instead of trading them.

The takeThe bar is not too high or too low. It is set by whoever last got embarrassed, which is the honest description of how most hiring standards are calibrated and the reason they only ever move one way. Nobody has been called into a room about a candidate they turned down. Until a process produces some record of its own rejections, the argument about where the bar belongs runs between one person's memory and another's, and the memory with a bad hire in it wins every time.

Where Olive fits

Open a role and see what the work shows

Olive is bought by the employer and reads the work rather than the paperwork: a 50-to-70-minute session grounded in one occupation, done with an AI assistant on the candidate's own clock, and written up by a human reviewer as six findings with the moment behind each one. The candidate is granted that same report, free, on every tier.

Rank your shortlist

Which Error Is Your Process Tuned Against?

Look at what happens to an ambiguous candidate. If a mixed signal at any stage resolves to a rejection, the process is tuned against the bad hire, whatever the hiring philosophy on the careers page says. That is a defensible setting for some roles and an expensive one for others, and almost no team has ever stated which it chose, because the setting arrived through a hundred small decisions rather than one.

The vocabulary is worth having, because it makes the trade sayable. A false positive is somebody who was hired and should not have been. A false negative is somebody who was rejected and would have done the job well. Any threshold that reduces one raises the other, holding the quality of your evidence fixed. That last clause is the escape route, and the whole final section is about it.

The economics differ by role, and this is where the decision actually gets made. A mistake recoverable inside a quarter, on a role with a short ramp, a defined probation, a large cohort and work that is visible early, tolerates a looser bar, because the process has a cheap second chance. A mistake that is not recoverable, on a small team where the ramp runs nine months and the person holds a relationship nobody else can hold, does not. Write the role into one of those two categories before arguing about the bar, and half the argument resolves itself.

Why the Missed Hire Never Shows Up in a Review

Because there is nothing to review. A rejected candidate leaves a row in an applicant tracking system and no further data, ever. They do not come back in eighteen months with a promotion and a performance rating to prove the decision wrong, so the error is not merely uncounted, it is structurally uncountable by any process that only observes the people it hired.

That asymmetry compounds. A bad hire generates a manager who remembers it, a post-mortem that names a stage, and a change to that stage. Nothing on the other side generates anything, so every feedback loop in the system pushes the bar in one direction, and each push feels like rigor rather than drift. Ten years of that and a process can be rejecting on evidence nobody would defend if it were written on a slide.

Knowing how much signal the underlying methods carry bounds both errors. In the 2022 re-analysis of the selection literature, structured interviews came out top ranked at .42, ahead of job knowledge tests at .40, work samples at .33, cognitive ability tests at .31 and unstructured interviews at .19 1. Those are corrected correlations with supervisor ratings rather than accuracy rates, and the .42 carries a wide credibility interval. Read them as a ceiling: the best evidenced single method in the field leaves a great deal of both mistakes on the table, so a process claiming to have eliminated one of them has simply stopped counting.

Whether a resume screen that rejects most applicants is throwing away the wrong people is the same question at the stage where the volume, and therefore the count of missed hires, is largest. What the mistakes cost in money is a separate calculation worth building from your own payroll.

Set the Error Rate Deliberately, Because You Already Did

A threshold is a choice even when nobody made it on purpose. Every stage with a cutoff has one, and moving it moves both error rates in opposite directions. The automated end of hiring shows this most plainly, because there the threshold is a literal number in a settings panel, and the published research on where those numbers land is unusually direct.

RAID, an evaluation of twelve text detectors across eight domains, found that a tool's false-positive rate is mostly a function of how it is tuned rather than a property of the tool: at naive default thresholds, one open-source detector flagged human writing at a 100% rate while closed-source tools stayed below 1.7%. Even after every detector was calibrated to a 5% overall false-positive rate, individual domains stayed far off, with 20.4% of human reviews, 33.4% of human recipes and 13% of Wikipedia text still flagged 2. None of the eight domains is a resume: they run from news and Wikipedia to books, reviews and recipes, so no number there is a measured rate on hiring documents. The transferable part is that the rate is a dial, and the buyer of a tool almost never sees where it is set.

The arithmetic of a small rate at volume is the second half. Vanderbilt University switched off Turnitin's AI detector after publishing the multiplication: a claimed 1% false-positive rate against 75,000 papers submitted in a year works out to roughly 750 papers wrongly labelled 3. Whatever your equivalent numbers are, run that multiplication before adopting any stage with a stated error rate.

And the errors are not distributed randomly across people. Seven detectors run over 91 essays written by non-native English speakers produced an average false-positive rate of 61.3% 4. One cohort, one document type, 2023 tool versions, and the figure is an average across seven tools rather than any single one. The durable finding is the direction of the error, which is that it lands hardest on a group defined by something other than ability. Which of a detector, a structured interview and a work sample actually defends a decision is the practical version of that choice.

Test the Half You Cannot See

Two moves are available, and the second is much better than the first. The cheap one is to sample your own rejections: take twenty near-misses a quarter, the ones a panel split on, and record where they landed. It is rough, incomplete and biased toward candidates whose careers are publicly visible, and it is still more than the zero information a normal process produces about its own false negatives.

The better move is to stop trading the two errors and improve the evidence instead. Both rates fall together when a stage produces something specific rather than an impression, because the trade between them is fixed only while the quality of the evidence is fixed. In practice that means one stage where the candidate does a piece of the work and somebody records what happened, replacing one stage where three people formed a view and averaged it.

A short checklist for the stage you are least sure about. Does it produce an artifact somebody else could re-read? Could two reviewers reach the same conclusion from it? Does it reject anybody on something the role never requires? Would you be comfortable telling a candidate exactly why they did not pass? A stage failing all four is generating both errors at once and is the cheapest thing in the loop to replace.

The common failure at this stage is a screen on how a document sounds, which is an impression wearing the costume of evidence. Stopping hiring managers rejecting candidates for sounding like AI is worth doing before touching the bar at all, and what a screen still reads for when every resume looks polished is where the replacement usually starts.

See a sample report

Common questions

Is raising the bar ever the right answer?

Yes, when the role's economics say a mistake is not recoverable and the evidence behind the decision is strong enough to bear the extra weight. Raising a bar on thin evidence just rejects more people at random, since a threshold applied to noise produces noise. Improve the evidence first, then decide where the bar goes, and write down which of the two errors you have chosen to accept more of.

How would I even know I have a false negative problem?

Three symptoms are readable from inside. Panels that split often and resolve to reject. Rejection reasons that name style, polish or fit rather than a job requirement. And a funnel where the stage rejecting the most people is the one with the least written evidence behind it. None of these proves a missed hire, and together they describe a process that would produce them at a rate nobody is measuring.

Does a higher volume of applications change the trade?

It changes the arithmetic, not the trade. With more applicants per opening, a given rejection rate discards more good candidates in absolute terms while the rate itself looks unchanged, and any stage with a fixed error rate produces proportionally more wrong outcomes. High volume is also the condition under which teams add cheap automated screens, which is exactly when the multiplication of a small error rate by a large pile matters most.

Can a scoring model be tuned to balance the two errors?

A threshold can be moved, which is not the same as balancing anything, and any automated decision stage brings its own legal and notice obligations that vary by jurisdiction. Check with counsel before relying on one. The tuning question is also not answerable without labels on the outcomes you care about, and the missing labels are precisely the rejected candidates, so a model trained on hires learns from one half of the evidence.

Who should own this decision?

Whoever owns the role's economics, which is usually the hiring manager, with people-ops owning the record that the decision was made and on what basis. It is a business decision rather than a recruiting preference, since the answer depends on ramp time, team size, cohort volume and how recoverable a mistake is. Put it in the intake conversation, in one sentence, before the first candidate is sourced.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the ceiling claim that every selection method leaves both errors on the table: structured interviews .42, job knowledge .40, work samples .33, cognitive ability .31, unstructured interviews .19.
  2. 2. RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 2024. aclanthology.org Supports the claim that a false-positive rate is a tuning choice rather than a property: 100% at a naive threshold against under 1.7% for closed-source tools, and 20.4%, 33.4% and 13% by domain after calibration to 5%.
  3. 3. Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector Vanderbilt University (Brightspace / Center for Teaching), 2023. vanderbilt.edu Supports the base-rate arithmetic that a small stated error rate at volume produces a large count of wrong outcomes: a claimed 1% rate against 75,000 papers is roughly 750 wrongly labelled.
  4. 4. GPT detectors are biased against non-native English writers Patterns (Cell Press), via PubMed Central, 2023. pmc.ncbi.nlm.nih.gov Supports the claim that selection errors fall unevenly across groups: an average 61.3% false-positive rate across seven detectors on 91 essays by non-native English speakers.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.