Screening
An AI Screener Buys Throughput, Not Accuracy
Buy an AI screening tool if the bottleneck is genuinely reading time, and buy it to summarise and group applications rather than to cut them. Put that restriction in the configuration itself, because a setting survives a busy month and a training slide does not. Ask for validity evidence against job outcomes, which is a different thing from agreement with human raters: matching a biased baseline is not accuracy. And ask the vendor what would make this tool wrong. A bias audit certificate is not an answer to that question.
The takeVendor comparisons argue accuracy and time-to-hire. The published measurements are bleaker than the deck: embedding-based screeners select at different rates depending on the name on the resume, and larger models make errors that correlate with each other instead of cancelling out. What you are actually buying is reading capacity, which is a real thing to want once one opening draws hundreds of applications rather than dozens. Price it as capacity, put it where being wrong is cheap, and stop calling it judgment.
Where Olive fits
Open a role and see what the work shows
Reading capacity is what a screening tool genuinely adds, and it says nothing about how a candidate works with AI. Olive puts that in front of them as work: an assignment done with an AI assistant on the candidate's own clock, returned as six findings a human reviewer wrote, each with the timestamped excerpt behind it.
Rank your shortlistShould you buy one?
Only if reading time is the constraint, and only with the tool's authority written down before it is switched on. If the real problem is that the posting attracts the wrong people, or that nobody agrees what a strong candidate looks like, a screener will process the wrong applications faster and hand back a confident ordering of them.
Three conditions make the purchase defensible. Volume that a person genuinely cannot read, meaning hundreds per opening rather than dozens. A stated bar the tool is asked to apply, written by your team and not inferred from past hires. And a configuration where the output summarises and groups instead of excluding, with the exclusion decision left to a person who saw the application.
The third condition is the one that gets negotiated away. It is also the one that changes what you have to explain later, the candidate experience and the failure mode all at once, because a tool that only summarises cannot silently remove anybody. If the business case only works when the tool cuts, the honest description of the purchase is a reduction in review coverage, and that is a decision to take deliberately, never one to acquire as a side effect.
Before any of this, be sure the screen you already run is worth defending. How to tell whether your resume screen is throwing away the wrong people is the prior question, and adding automation to an unmeasured screen just makes it faster.
What the measurement evidence actually shows
Worse than the sales material and more interesting than a blanket refusal. A resume audit study using text embedding models across nine occupations, with more than 500 resumes and 500 job descriptions, found the models significantly favoured White-associated names in 85.1% of cases and female-associated names in only 11.1%, with Black male names disadvantaged in up to 100% of the tested comparisons 1. That is selection rates differing by name alone.
The second finding is less discussed and matters more for a buyer. An evaluation across more than 350 models, including a resume-screening task, found substantial correlation in model errors: on one leaderboard dataset models agreed 60% of the time when both were wrong 2. Larger and more accurate models had highly correlated errors even across different architectures and providers.
The second result breaks the standard mitigation. Buying a second opinion from a different vendor does not give you an independent second opinion. If the market converges on similar models, a candidate the tools get wrong is wrong everywhere at once, which is a different risk from a single tool being imperfect. No comparable measurement exists for human reviewers, so keep the claim narrow: a second subscription buys more reading, and nothing shows it buys a check on the first.
Two honest limits on both studies. They test research configurations rather than any particular commercial product, so neither one is evidence about the vendor in front of you. And neither says these tools read nothing useful. What they undercut is the specific claim on the slide: that the tool decides better than the process it replaces.
Ask what would make this tool wrong
That single question separates vendors who have measured their product from vendors who have described it. A team that has tested its own tool can answer immediately, usually with a population or a role type where performance degrades. A team that cannot will reach for a certificate, and a certificate is a different kind of object from an answer.
What to ask for instead, in order of how much it tells you:
1. Validity evidence against job outcomes. Did people the tool rated highly perform better, and in which roles, measured how? This is the hard one and the one worth paying for. 2. Agreement with human raters, clearly labelled as such. It is useful and it is not accuracy. If the humans were inconsistent or biased, matching them is a defect described as a feature. 3. Selection rates by group on your data, not on theirs. Which is where the audit question actually lives. 4. What the tool does with an application it cannot parse. The silent failure mode nobody demos.
On certificates specifically. A posted New York City bias audit computes a selection rate and an impact ratio by sex, by race and ethnicity, and by intersectional category, and nothing else: not accuracy, not job-relatedness, not whether any individual was treated fairly 3. An auditor may exclude any category under 2% of the audit data from the impact ratio calculation, which most often drops the smallest groups in the pool. And an employer that has never used the tool may rely on an audit built on other employers' data or on synthetic data 3. So a first-time buyer can hold a genuine audit describing nobody who ever applied to it.
The four-fifths rule gets quoted the same way and means less than the slide implies. The agencies that wrote it said in their own interpretive guidance that the 80% figure is a rule of thumb, not intended as a legal definition, and that it speaks only to adverse impact rather than resolving whether discrimination occurred 4. Passing it is not a clean bill of health. For the wider vendor conversation, the questions that verify a screening vendor's claims covers what to send before a demo.
Measure the throughput you actually bought
Take a baseline before the tool arrives, because after it arrives nobody can reconstruct one. Count recruiter hours spent reading per opening, elapsed days from application to first human response, and the share of applications a person actually opened. Those three numbers are what the purchase is supposed to move, and they are the only claims you will be able to defend internally at renewal.
Be realistic about where the days go. In SHRM's benchmarking data the median time-to-fill for nonexecutive roles was 44 days, of which the median organisation spent about 5 days screening applicants, 7 days interviewing and 4 days deciding and extending an offer 5. Screening is real work and it is a small slice of elapsed time, so a tool that halves it moves the headline number by two or three days. If the business case promised more, the case was about something else.
Then check the restriction held. Two months in, pull a sample of applications the tool grouped at the bottom and have a person read them cold. You are looking for one thing: whether anybody was effectively excluded without a human opening the file. If they were, the configuration drifted or the workflow routed around it, and that is worth knowing before a candidate asks.
One last thing that belongs in the plan rather than in a crisis. Someone will ask why they were screened out, and the answer has to be something a person can say out loud. What you owe a candidate who asks why the AI screened them out is the version of that conversation worth rehearsing before the first one arrives.
Common questions
Is an AI screener more accurate than a recruiter reading resumes?
No vendor has shown that against hiring outcomes, and the published research points the other way on the parts that have been measured. Text embedding models used for resume screening selected at materially different rates depending on the name on the document, and large evaluations find model errors correlate rather than cancel. A recruiter reading resumes is also unreliable, which is the real argument for changing the stage rather than automating it. Neither comparison supports the accuracy claim on a vendor slide.
What is the difference between a bias audit and validity evidence?
A bias audit asks whether groups are selected at different rates. Validity evidence asks whether the thing being measured relates to doing the job. A tool can pass an audit and predict nothing, and it can predict well and still produce impact. Buyers routinely accept the first as though it answered the second. Ask for both, and ask which population each was computed on, because an audit run on another employer's applicants tells you very little about yours.
Can we use the tool to rank and let a recruiter make the final call?
Yes, though the split changes less than it sounds. A reviewer who opens an ordered list and works down it is treating the output as the heaviest single input. New York City's 2023 rule under Local Law 144 counts exactly that as substantially assisting a decision: a simplified output weighted more than any other criterion in the set 3. Whether that reaches your own process is a question for counsel. If the ordering exists, treat it as a decision the tool made and review it as one, or configure the output as groups rather than a sequence.
Does buying two tools reduce the risk?
Less than expected. Model errors correlate across providers and architectures, and correlate more among the larger and more accurate models, so a second tool is not an independent second opinion. Candidates that one gets wrong the other often gets wrong too. If redundancy is the goal, a person reading the file is at least not running the same model, which is the argument for spending the saved capacity on human reading further down the funnel rather than on a second subscription.
What should we do if we already bought one?
Three things, none of them expensive. Change the configuration so the tool groups and summarises rather than excludes, and confirm in the admin settings rather than in the training material. Pull a sample of what it put at the bottom and have someone read it cold. And ask the vendor, in writing, for selection rates by group computed on your own applicants. If they cannot produce that, you have learned something useful about the renewal conversation.
Will an AI screener reduce our time-to-hire?
A little, and less than the deck suggests, because screening is a small share of elapsed time. In SHRM benchmarking the median organisation spent about five days screening inside a median time-to-fill of forty-four days for nonexecutive roles. The rest of the clock runs on things a screening tool never touches. If time-to-hire is the goal, measure where your own days go first, because the answer is usually not the reading.
References
- 1. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval arxiv.org Supports the finding that Massive Text Embedding models used for resume screening, tested over nine occupations with more than 500 resumes and 500 job descriptions, significantly favoured White-associated names in 85.1% of cases and female-associated names in 11.1%, with Black male names disadvantaged in up to 100% of cases.
- 2. Correlated Errors in Large Language Models arxiv.org Supports the evaluation across over 350 models including a resume-screening task, the finding that models agreed 60% of the time when both erred on one leaderboard dataset, and that larger, more accurate models had highly correlated errors.
- 3. Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY 5-300 et seq.) rules.cityofnewyork.us Supports what a posted bias audit computes and omits, the exclusion of categories under 2% of the audit data from the impact ratio calculation, the allowance for a first-time buyer to rely on other employers' historical data or synthetic data, and the rule's definition of substantially assisting a decision as a simplified output weighted more than any other criterion in the set.
- 4. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures (Q.11, Q.19) eeoc.gov Supports the point that the 80% four-fifths figure is a rule of thumb rather than a legal definition, and speaks only to adverse impact rather than resolving unlawful discrimination.
- 5. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) shrm.org Supports the median time-to-fill for nonexecutive roles and the median days spent screening, interviewing and deciding inside it.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.