Policy

You Cannot Compare the Model to Recruiters You Never Measured

No one knows whether an AI resume screen is less biased than the human recruiters it replaced, because the recruiters were never measured. No employer wrote down why a screener set a resume aside, so there is no baseline for a model to beat or fall short of. What evidence exists is narrower than either side admits: a hiring test that beat managers' overrides on tenure, a Black-name penalty across 108 large employers, and one randomized trial of an AI interviewer at a single firm.

The takeThe comparison is the wrong question, and it survives because both sides benefit from it staying unanswerable. A vendor gets to claim an improvement over a baseline nobody can produce; a critic gets to claim a harm against the same missing number. The question that can be answered is which of the two leaves something behind. A disparity you can trace to a written reason can be argued with. A disparity in a column of scores can only be counted, and counting is not the same as knowing what happened.

Where Olive fits

Open a role and see what the work shows

Olive has not performed a bias audit, and olive.is says so, because attempt volume is too low for a four-fifths ratio to mean anything yet. What it does hold is the material such a review would need: six findings per session, each written by a person and anchored to a timestamped excerpt, granted to the candidate in the same form the employer reads.

Rank your shortlist

Why can't the comparison be settled?

Because the human side was never measured. A hiring team that replaced twelve screeners with a model has, at most, a record of who was rejected and no record of why, so there is nothing to compute a before-and-after against. The comparison also assumes a single system can stand in for a room of people, and that the room's errors were partly cancelling each other, which is a strong assumption nobody states out loud.

Where the human side has been measured, the ceiling is low. A review of why employers resist selection decision aids reports that the interrater reliability of the traditional unstructured interview is so low that, even with a perfectly reliable and valid criterion, interview-based judgments could never account for more than 10 percent of the variance in job performance 3. The same review reports a 1996 survey of 201 HR executives who rated the unstructured interview more effective than any paper-and-pencil assessment procedure. Belief and measurement pointed in opposite directions, and the belief won for decades.

That is a finding about interviews rather than about resume screening, and the gap matters: no comparable measurement of recruiters reading resumes exists at all. Which is the point. The half of the comparison a vendor is being asked to beat has never been quantified in any employer's own building, so an improvement claim and a harm claim are equally unfalsifiable. A vendor's own bias audit does not fill that gap: under New York City's 2023 final rules, an employer that has never used a tool may rely on an audit computed on other employers' historical data, or on synthetic test data where too little real data exists 7. Anyone who answers this question with a number is answering a different question, usually about their own product.

What does the evidence on each side actually say?

Less than either side claims, and none of it is a head-to-head comparison. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed job tenures by just over 25 percent, and comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6 to 7 percent shorter job durations 1. Human discretion over a structured signal was not, on average, superior private information.

Hold that result at the strength it was measured. There was no randomization; it compares managers at the same site who override at different rates. Tenure is the quality measure, chosen because turnover is what this employer pays for; no supervisor performance rating was collected. Overrides averaged 22 percent, so they were routine in these firms. The instrument was an online questionnaire scored by a proprietary algorithm into a green-yellow-red band, rather than an interview or a model reading resumes. It says the average override was worse, and says nothing about whether the test was right about any individual person.

What has been measured on the machine side is version-locked: the findings in model-driven resume screening are attached to specific model releases and specific prompts, and a release resets them. The human side has a steadier measurement, and it is not flattering. In the 83,000-application audit at 108 large US employers, distinctively Black names reduced the chance of employer contact by 2.1 percentage points, an effect equal to 9 percent of the Black mean contact rate, while the average gender contact gap was zero with a between-company standard deviation of 2.7 percentage points roughly symmetric about zero: some firms consistently favored men, others consistently favored women, and the two cancelled only across companies 2.

That last cancellation is where the argument against the machine usually starts, and it is weaker than it looks. The offsetting happened across 108 separate employers. Inside any one of them the gap held its direction, which is what a single consistent rule looks like from the applicant's side, and a model applied across a whole pipeline has that property at higher volume. So scale raises the cost of being wrong. It does not show that models are more biased than people, and no employer should read an aggregate of other companies as evidence that its own screeners cancel out.

Which of the two can you inspect afterwards?

The one that leaves a written reason behind. A disparity found in a column of scores can be counted and nothing else, because there is no account of how any individual score was reached; a disparity found in a set of written reasons can be read, disagreed with, and traced to the criterion that produced it. The difference comes from what a process stores, and a person reading resumes without notes stores nothing either.

The research supports the narrow version of that claim. Text-mining post-interview notes on 7,650 candidates hired at a large Chinese internet technology company, researchers found the number of job-related capabilities an interviewer named in the notes was positively related to later job performance and promotions and negatively related to turnover, with a one standard deviation rise in the match between notes and job analysis corresponding to roughly a 2 percent rise in performance 4. It is correlational, at one firm in one country, and range-restricted by construction, since only candidates who passed could be observed. What it shows is modest and worth having: what a reviewer writes down carries information that a rating on its own does not.

The one large randomized result on machines in this loop points the same way. In a field experiment where 70,000 applicants were randomly assigned to a human or an AI voice-agent interview, with human recruiters evaluating the transcripts and making every hiring decision in both arms, applicants interviewed by the agent were 12 percent more likely to receive an offer, with higher starts and retention and no drop in the productivity of hires 5. The mechanism the authors identify is consistency in gathering evidence. The costs are in the paper: 5 percent of applicants ended the interview rather than speak to an AI, and the agent hit technical difficulties in 7 percent of cases. One firm, one country, one job family, and a working paper that has not been peer reviewed. Reviewers who cannot judge the work in front of them are a separate failure with its own fix, which is where getting hiring managers who don't use AI to judge AI-assisted work starts.

Start recording what each screening decision rested on

Start with the record you wish you had two years ago. From today, store the criterion each screening decision was made against and the specific thing in the application that met or missed it, for every rejection as well as every advance. Six months of that is a baseline, and it is the only version of this comparison you will ever be able to run on your own pool.

1. Write the criteria before the applications arrive. A reason recorded after a decision is a rationalization, and it will read as one later. 2. Record rejections in the same detail as advances. The rejected side is where a disparity lives and where nobody keeps notes, which is exactly why the question is unanswerable now. 3. Keep it at the individual level. A rate is an aggregate of things that already happened; only the individual record can be re-read when someone asks what happened to a particular person. 4. Decide in advance what you will do if you find something. California's amended FEHA employment regulations, effective October 1, 2025, make evidence, or the lack of evidence, of anti-bias testing relevant to a discrimination claim and to any defence, including the quality, recency and scope of the effort and the response to the results, and they extend records retention from two years to four with automated-decision system data expressly included 68. They impose no duty to test and set no standard for adequate testing; what changes is what a court may weigh. Testing and then ignoring the result is the worst of the available positions. That is California as of 2025, and the register here is check with counsel; none of this is advice.

Six months in, you will not have settled whether models are less biased than people. You will have something better for the decision in front of you: your own numbers, at your own stages, with an account of every decision behind them. The debate will still be running. The record is an asset, and it is the one thing a candidate who challenges an outcome can be shown, which is the standard an AI-skills assessment has to hold up to as well.

Read the evidence

Common questions

Doesn't research show algorithms outperform human recruiters?

One good study shows something narrower. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed tenures by just over 25 percent, and managers who overrode the test's recommendation more often got hires who stayed 6 to 7 percent less time. That is evidence about discretion over a structured signal in a high-turnover setting, measured by tenure rather than performance, and it is not randomized. It does not say a model reading resumes beats a recruiter reading resumes, which nobody has measured.

Can a vendor's bias audit serve as your baseline?

No, and the reason is structural. A published audit describes selection rates on the sample it was computed on, and New York City's rules let an employer that has never used a tool rely on an audit built on other employers' historical data or on synthetic test data. Your baseline would have to describe your own recruiters, on your own applicants, in your own roles. No audit of a vendor's product contains that, and none claims to. It answers a compliance question, not the comparison question.

Is a model's bias worse than a human's because it applies consistently?

That is the strongest form of the concern, and it rests on reasoning about how errors combine. Nobody has measured it head to head. The offsetting people have in mind was measured across 108 separate companies; inside any one of them the gender contact gap held its own direction, with a between-company standard deviation of 2.7 percentage points and an average of zero. One rule applied across one pipeline offsets against nothing. Scale raises the cost of being wrong, which is not the same claim as models being more biased than people.

What is the minimum record worth keeping on a screening decision?

Four fields, and they fit in an ATS note. The criterion applied, the specific thing in the application that met or missed it, who decided, and when. That is enough to re-read a decision months later, enough to see whether one criterion is doing all the rejecting, and enough for a person to disagree with a call rather than simply observe a rate. Nothing about it requires a vendor, a tool purchase or demographic data you do not already hold.

Does keeping records create legal exposure if a problem turns up?

That is a question for employment counsel and the answer varies by jurisdiction. What is publicly stated in California's amended regulations is that evidence, or the lack of evidence, of anti-bias testing is relevant to a discrimination claim and to any available defence, including the response to the results. Read carefully, the exposure runs the other way from the intuition: the position that reads worst is finding a disparity and doing nothing, and the position with no evidence at all is not a safe harbour.

References

  1. 1. Discretion in Hiring National Bureau of Economic Research Working Paper 21709 (revised September 2017); published 2018 in the Quarterly Journal of Economics, 2017. nber.org Supports the just-over-25-percent gain in completed job tenures from introducing a job test across 15 firms, and the 6 to 7 percent shorter durations associated with a one standard deviation higher rate of hiring against the test, along with the 22 percent average exception rate.
  2. 2. Systemic Discrimination Among Large U.S. Employers NBER Working Paper 29053 (revised May 2022); published in the Quarterly Journal of Economics, 137(4), 1963-2036, 2022. nber.org Supports the 2.1 percentage point Black-name contact penalty equal to 9 percent of the Black mean contact rate across 83,000 applications at 108 large employers, and the zero average gender contact gap alongside a 2.7 percentage point between-company standard deviation roughly symmetric about zero.
  3. 3. Stubborn Reliance on Intuition and Subjectivity in Employee Selection Scott Highhouse, Industrial and Organizational Psychology, 1(3), 333-342, 2008. edbatista.com Supports the reliability ceiling under which unstructured interview judgments could never account for more than 10 percent of the variance in job performance, and the survey of 201 HR executives rating it more effective than any paper-and-pencil procedure.
  4. 4. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company Liu, Chang, Jiang, Ma and Zhou, Frontiers in Psychology, Volume 11, 2021. frontiersin.org Supports the 7,650 hired candidates, the relationship between the number of job-related capabilities named in interviewer notes and later performance, promotions and turnover, and the roughly 2 percent performance rise per standard deviation of the matching score.
  5. 5. Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews Brian Jabarian and Luca Henkel, arXiv:2607.28222, 2026. arxiv.org Supports the 70,000 randomized applicants, the 12 percent higher offer rate with human recruiters deciding in both arms, and the two costs the paper reports: 5 percent of applicants ended the interview and the agent hit technical difficulties in 7 percent of cases.
  6. 6. Final Unmodified Text of Proposed Employment Regulations Regarding Automated-Decision Systems (Attachment B), 2 CCR sections 11009, 11013 California Civil Rights Department, Civil Rights Council, 2025. calcivilrights.ca.gov Supports the relevance of anti-bias testing evidence and the response to its results in a FEHA claim or defence, and the move from two years to four years of records retention with automated-decision system data included.
  7. 7. Notice of Adoption of Final Rule: Use of Automated Employment Decisionmaking Tools (6 RCNY 5-300 et seq.) NYC Department of Consumer and Worker Protection, 2023. rules.cityofnewyork.us Supports the rule that an employer that has never used an AEDT may rely on a bias audit computed on other employers' historical data, or on synthetic test data where too little real data exists.
  8. 8. Rulemaking Actions - Civil Rights Council California Civil Rights Department, Civil Rights Council, 2025. calcivilrights.ca.gov The Council's own record of the automated-decision-system employment regulations: approved by OAL and filed with the Secretary of State, effective October 1, 2025.

8 sources, numbered by first appearance. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.