Screening

Is It Legal to Reject a Resume on an AI-Detector Result?

No federal statute bans rejecting a resume over an AI-detector flag. But the moment the flag decides who advances, it is a selection procedure under the Uniform Guidelines, and one that screens out a protected group at a higher rate is discriminatory unless validated against the job. No vendor sells that validation, and the measured false positives land hardest on non-native English writers. In New York City, screening with a detector also triggers Local Law 144's bias audit and 10 business days' notice before use.

The takeAuthorship stopped being a useful question before most screening teams noticed. Assume most resumes reaching you now passed through something, so a tool that answers who typed the sentence is answering a question that no longer separates applicants, and it charges an error rate aimed at people who learned English second. I suspect the detector market runs on the comfort of a number, and until someone publishes a validation study against job performance, that comfort is the whole product. A device that cannot tell you what a person can do is not a screen. It is a mood ring with a legal department.

Where Olive fits

Open a role and see what the work shows

A flag on a document says nothing about who did the thinking, so Olive assesses the person instead: a 40-to-60-minute occupational assignment done with an AI assistant, returned as six separately evidenced findings a human reviewer writes by hand. The candidate is granted the same report the employer gets.

Rank your shortlist

What law actually applies when a detector flag decides a rejection?

No federal statute names AI detectors, and none has to. The moment a detector result decides who advances, it is a selection procedure under the Uniform Guidelines on Employee Selection Procedures, defined there as "any measure, combination of measures, or procedure used as a basis for any employment decision" 1. From that point the ordinary rules apply.

Section 1607.3(A) is blunt: a selection procedure that has an adverse impact on the opportunities of any race, sex or ethnic group is discriminatory unless it has been validated 1. The EEOC's guidance on tests and selection procedures states the standard the employer has to meet, job-related and consistent with business necessity, which it explains as "necessary to the safe and efficient performance of the job" 2. It also names the second hurdle: even a job-related procedure loses if a challenger shows a less discriminatory alternative was available 2.

That last clause is the one to sit with. If a cheaper procedure answers the same question and flags fewer people wrongly, the detector is the harder thing to defend, not the safer one.

New York City reaches the practice by name. Local Law 144 covers a computer-based tool that uses machine learning, statistical modeling, data analytics or artificial intelligence, helps an employer make an employment decision, and substantially assists or replaces discretionary decision-making 5. A classifier that sorts a resume into one bucket or another is doing exactly that, and DCWP is explicit that an employment decision includes screening rather than only the final hire 5. Where the law applies, a bias audit has to be completed before use, its summary posted publicly, and NYC candidates given notice 10 business days ahead; enforcement began 5 July 2023 5.

New York City is not the only place writing rules for automated screening, so check the ones your postings reach. And read what a rejection reason has to survive before a candidate asks you for one.

How do detector false positives land, and on whom?

Not randomly. Seven widely used detectors were run over 91 human-written TOEFL essays and misclassified them as AI-generated at an average false-positive rate of 61.22 percent, with 97.80 percent flagged by at least one detector, while essays by US eighth-graders were classified with near-perfect accuracy 3. National origin is a protected class, so that error pattern is the adverse-impact case writing itself.

The follow-up experiment is the tell. When the same TOEFL essays were rewritten to enrich the word choices, the average false-positive rate fell from 61.22 percent to 11.77 percent 3. Simplifying the word choices in native-speaker essays pushed misclassification the other way 3. The tool is reading register, and register tracks first language, education and profession.

Then do the arithmetic on volume. Vanderbilt did it before disabling Turnitin's AI detector: against a claimed 1 percent false-positive rate and 75,000 submissions in a year, roughly 750 pieces of student work would have been wrongly flagged, and the university concluded the software was not an effective tool to use 6. One percent sounds like rounding until it meets a funnel. Whether detectors work at all is a separate question from whether their errors are survivable.

Title VII disparate impact requires no intent. A procedure that screens out non-native English writers at a materially higher rate is exposure whether or not anyone meant it, and the four-fifths rule is where the enforcement agencies start: a selection rate for any race, sex or ethnic group below 80 percent of the highest group's rate is generally regarded as evidence of adverse impact, and smaller gaps can still count where they are significant in statistical and practical terms 1.

Does one detector threshold work across job families?

No. Baseline AI use at work differs sharply by occupation: highest in computer and mathematical, management, and business and finance jobs at 43, 42 and 41 percent, with every occupation group somewhere between 15 and 50 percent 4. Writing was the work task people most often used it for, at 39.5 percent 4. One threshold therefore carries a different false-positive burden into every family you apply it to.

Adverse impact is assessed per job, against the total selection process, and the agencies keep the discretion to reach an individual component of it where appropriate 1. Running one detector across ten requisitions is not one exposure, then. It is ten, each with its own applicant pool and its own baseline rate of assisted drafting.

Register moves the same way. A compliance analyst writing to a controls template and a marketing coordinator writing to a brand voice do not produce prose with the same statistical texture before either of them opens an assistant. The same score means a different thing in each case, and a single cut line quietly sets a different standard for every job family it touches.

If a detector stays in the process, the minimum is per-family impact data. Records that disclose the impact a selection procedure has on employment opportunities by race, sex and ethnic group are already what the Guidelines ask users to maintain, with sampling permitted where applicant volume is large 1. The adverse-impact arithmetic on a screening step is not hard to run. Not having run it is the finding.

What to run instead of the detector

Screen the claim, not the authorship. Pick two or three job-related criteria, write them down before you open the pile, and judge each submission against them. Where the real question is whether the person can do the thing, put a short piece of the work in front of them with AI allowed and the working visible. Authorship stops mattering once the work is watched.

  • Write the criteria first. A rejection reason that existed before you saw the candidate is the one that survives being questioned.
  • Ask about a decision in the work, not about the prose. Which figure did they check, what did they discard, what did the source turn out not to support. A model produces the paragraph; it does not produce the reason the paragraph says what it says.
  • Move the weight to a work sample. Skipping the resume screen in favour of a short work sample converts an authorship question into an observation.
  • If the detector stays, demote it. A flag routes a file to a human read and never rejects on its own. Record what the reviewer read and the job-related reason they gave, because "substantially assists" is the phrase that decides whether Local Law 144 applies 5, and a rubber stamp assists substantially.
  • Tell candidates what you use. Notice under Local Law 144 runs 10 business days before use, not after a rejection 5.

None of this makes authorship answerable, and it is worth being straight that nothing does. Whether an AI-written resume should be disqualifying at all is a policy call to make explicitly. The alternative is making it implicitly, at a threshold a vendor chose.

See a sample report

Common questions

Is there a federal law that bans rejecting a candidate over an AI-detector result?

No. No federal statute names AI detectors. The exposure comes from the general rule for any screening device: under the Uniform Guidelines, a procedure used as a basis for an employment decision that produces adverse impact is discriminatory unless it has been validated 1, and the EEOC standard is job-related and consistent with business necessity 2. A detector output has no validation study behind it, so the live question is not whether the practice is banned but whether you could defend it.

Does an AI detector count as an automated employment decision tool in New York City?

If it substantially assists the screening decision, yes. Local Law 144 covers a computer-based tool that uses machine learning, statistical modeling, data analytics or artificial intelligence, helps make an employment decision, and substantially assists or replaces discretionary decision-making 5. DCWP states that an employment decision includes screening, not only the final hire 5. Where it applies, a bias audit must be completed before use, a summary posted publicly, and NYC candidates notified 10 business days in advance 5.

Does having a human review the flag remove the risk?

Only if the human actually decides. The New York City test turns on whether the tool substantially assists the decision 5, and a reviewer who forwards every flagged file to a rejection is not deciding anything. Under the Uniform Guidelines the position is the same: adverse impact attaches to the procedure as used, not to the reporting line above it 1. Keep a record of what the reviewer read and the job-related reason they gave.

Should you ask candidates whether they used AI to write their resume?

Ask, but do not treat the answer as a test. A yes tells you nothing about who did the thinking, and a no is unverifiable. The useful version of the question comes later and is about the work: which claim did you check, what did you throw away, what did the source not support. A disclosure question sets expectations. It is not evidence, and it will not stand in for one.

What records do you have to keep if a detector is part of screening?

Records that disclose the impact your selection procedures have on employment opportunities by race, sex and ethnic group. The Guidelines put that on the user and permit a sample where applicant volume is large 1. In New York City, add the bias audit and its public summary 5. Practically: keep the threshold you used, the date you last changed it, the rejection reasons written against job-related criteria, and what any human reviewer concluded.

What is the four-fifths rule and does it apply to a detector threshold?

A selection rate for any race, sex or ethnic group below four-fifths, or 80 percent, of the rate for the highest group is generally regarded by the federal enforcement agencies as evidence of adverse impact 1. It applies to any measure used as a basis for an employment decision, which a detector threshold is 1. Smaller differences can still count where they are significant in both statistical and practical terms 1.

References

  1. 1. Uniform Guidelines on Employee Selection Procedures (1978), 29 CFR Part 1607 Equal Employment Opportunity Commission, via govinfo, 2023. govinfo.gov Sec. 1607.16(Q) defines a selection procedure as any measure used as a basis for an employment decision; 1607.3(A) makes an adverse-impact procedure discriminatory unless validated; 1607.4(A) requires impact records; 1607.4(C) assesses the total selection process per job and reserves agency discretion over individual components; 1607.4(D) states the four-fifths rule; 1607.9(A) excludes promotional literature as validity evidence.
  2. 2. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov States the job-related and consistent with business necessity standard, explains it as necessary to safe and efficient job performance, and names the less discriminatory alternative test.
  3. 3. GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu and Zou (arXiv:2304.02819), 2023. arxiv.org Seven detectors over 91 human-written TOEFL essays: 61.22 percent average false-positive rate, 97.80 percent flagged by at least one, near-perfect accuracy on US eighth-grade essays, and 61.22 down to 11.77 percent after word-choice enrichment.
  4. 4. The Rapid Adoption of Generative AI (NBER Working Paper 32966) Bick, Blandin and Deming, National Bureau of Economic Research, 2024. nber.org Generative AI use at work by occupation group: 43, 42 and 41 percent for computer and mathematical, management, and business and finance; every group between 15 and 50 percent; writing the top-ranked work task at 39.5 percent.
  5. 5. Automated Employment Decision Tools: Frequently Asked Questions NYC Department of Consumer and Worker Protection, 2023. nyc.gov Defines an AEDT under Local Law 144 of 2021, states that an employment decision includes screening, and sets the bias audit before use, the public summary, the 10 business day notice and the 5 July 2023 enforcement date.
  6. 6. Guidance on AI detection and why we're disabling Turnitin's AI detector Vanderbilt University Center for Teaching, 2023. vanderbilt.edu Turnitin's claimed 1 percent false-positive rate applied to 75,000 submissions works out to roughly 750 wrongly flagged papers, and the institution disabled the tool.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.