Screening
Do AI Detectors Work on Resumes and Cover Letters?
AI detectors do not work well enough on resumes and cover letters to justify rejecting anyone. A comprehensive public test ran 14 tools against human, AI, translated, edited and paraphrased documents: the best scored 76%, and the field averaged 30% on edited text and 15% on paraphrased. Seven detectors flagged 61% of essays by non-native English speakers as AI-generated. One self-editing pass drops detection to 13%, so the pile you flag is the pile that wasn't trying to hide. Assume the model helped, and test the work instead.
The takeA detector is not really bought to find AI. Most likely it gets bought because the screen stopped separating people, and because a number with a vendor's name on it is easier to defend in a meeting than a reading is. What makes it stick is that it can never be caught being wrong. The cost of a wrong cut never lands on a dashboard, so nobody inside the company has ever had to argue with the tool. Most teams running a detector have, I would guess, no idea what it costs them, and no way left to find out.
Where Olive fits
Open a role and see what the work shows
A document cannot show you who did the thinking, so Olive assesses the person instead: a 40-to-60-minute occupational assignment done with an AI assistant, returned as six findings, each anchored to the moment it happened. The candidate is granted the identical report, free.
Rank your shortlistDo AI detectors actually work on resumes and cover letters?
Not reliably enough to reject on. One of the most comprehensive public tests put 14 detection tools against human, AI-written, machine-translated, human-edited and paraphrased documents. The best tool reached 76% accuracy overall, no tool cleared 80%, and across the field the tools averaged 30% on documents a human had edited and 15% on paraphrased ones 1.
The failure mode matters more than the headline number. The same test found the tools biased toward calling text human-written 1, which sounds like the safe direction until you work out what it implies. The AI-drafted applications you are worried about mostly sail through. The ones that trip the alarm are the odd documents: short, plain, structurally conventional. A resume is all three by construction.
Cover letters are not the safer case, whatever the detector vendors say about longer narrative text giving a model more to work with. More text buys a more confident number, not a more correct one, and confidence and accuracy come apart exactly where you need them not to. A candidate who drafts with a model and then rewrites the result in their own words, which is the normal case now, lands in the 30% band above.
Same arithmetic as applications that all look excellent: the screen stopped separating people, and adding a detector to it does not put the separation back.
What does a false positive actually cost you?
About ten people per four hundred applications, and you never learn which ten. The same 14-tool test measured how often a tool's output would get an innocent person accused: on plain human-written English that false accusation rate averaged 2.4% 1. Half the tools produced none at all. GPT Zero was the outlier: half of everything it flagged would have been a false accusation 1.
So the tool you pick matters more than the category does, and the headline accuracy number will not tell you which you have. On plain human text the field scored 94% correct 1, but the missing 6% is documents a tool failed to call human with confidence (unclear and partly-wrong verdicts, not rejections). What turns one of those into a cut is a person reading a score, and nothing in your funnel records that reading.
The errors are not scattered randomly across your pipeline either. Stanford researchers ran seven detectors over 91 TOEFL essays written by non-native English speakers and 88 US eighth-grade essays. The detectors were near-perfect on the eighth-graders and misclassified the TOEFL essays at an average false positive rate of 61.22%. All seven agreed on 19.78% of them, and 97.80% were flagged by at least one 2.
So the cost is not a rounding error spread evenly. It is concentrated on a group, and a screening step that disproportionately excludes people on a protected basis has to be job-related and consistent with business necessity if anyone asks about it 3. "The tool returned a 92" is not that answer. You would be defending a threshold you did not set, cannot inspect, and cannot explain to the person you rejected, which is the legal exposure a detector actually creates, separate from whether it works.
A false rejection is silent, and silence is what makes it expensive. A bad hire announces itself in a quarter. A good candidate you cut in the first ten seconds never appears in any number you look at.
Why does trained, formulaic writing get flagged?
Because detectors key on predictability, and a well-trained analyst writes predictably on purpose. The Stanford study showed the mechanism directly: take native-speaker essays that detectors called human, simplify the word choices, and the average misclassification rate jumps from 5.19% to 56.65% 2. Nothing about the authorship changed. Only the range of vocabulary did.
House style is that rewrite, applied over years. A banking analyst is trained out of adjectives, into a fixed memo skeleton, and toward the same few dozen verbs everyone in the group uses. A consultant writes to a deck template with a required action title. A paralegal writes to a form, an underwriter to a schedule. Each of those is deliberate low variance, and low variance is the property a perplexity-based detector reads as machine.
Which puts your flagged pile in a predictable order: candidates who learned English as a second language, candidates trained in a formulaic professional register, and candidates who are simply plain writers. That is not a random six percent. It is a list of the people you were most likely to want.
That finding has a second half, and it is worse for the tool. When the researchers applied a single self-editing prompt to AI-generated essays, detection dropped from 100% to 13% 2. Anyone actively trying to get past a detector does it in one move, for free, in about four seconds. The population you catch is the population that was not trying.
Stop scoring the document. Screen for what it proves.
A resume is now a jointly written document. That changes what it is evidence of, not whether it is evidence: it still shows what a candidate chose to claim, what they put first, and how specific they were willing to be. Screen for a described decision (a problem framed, an output refused, a number checked), because a specific claim is checkable in the next conversation and a generic one is not.
Two rules make that cheap to run, and both of them are things you write down before the pile arrives rather than after it:
- State the rule in one sentence and apply it to every application. Whatever you decide to look for, a rule you can state is a rule you can apply the same way on Friday afternoon as on Monday morning, and a rule you can hand to counsel if anyone asks how the process worked 3.
- Move the weight off the artifact. If the cover letter no longer does the screening work it used to, cut it or replace it with two role-specific questions in a short answer box, and put the weight on something you generated yourself. Some teams go further and are skipping the resume screen for a work sample outright.
Both get easier once you have settled the question underneath them, which is what AI actually does in this role day to day. A screen for AI skill that nobody has defined against the job is a keyword rule wearing a better name.
What to do when an application is obviously AI-written
Assume it was, and change what happens next rather than who gets cut. Being right about the authorship buys nothing. The question was never who typed the sentences, it was whether this person can do the work and can tell when the model is wrong. Put the same task in front of everyone at the next stage, let them use AI openly, and read what they do with it.
If you want to test the document anyway, test it the way a reference check works: ask about it. Pick one concrete claim in the cover letter and ask the candidate to walk through how they got there: what the first number was, what they threw out, what they checked. Someone who did the work carries detail the document had no room for. Someone who did not has a summary. Four minutes, and it produces something no detector returns: an answer you heard yourself.
The second failure is quieter than the false rejection and lasts longer. A team that starts acting on detector scores stops asking what it wanted to know, which was never "did a human write this." The two questions are unrelated, and only one of them says anything about the job.
Olive is not a detector. It is an assessment an employer opens for a role: the candidate works a real occupational task with an AI assistant available, and a human reviewer writes six findings about what they actually did, each carrying the moment it rests on. See how Olive measures this
Common questions
Can an AI detector tell if a cover letter was written by ChatGPT?
Not with enough accuracy to act on. In a 14-tool test published in the International Journal for Educational Integrity, the best tool reached 76% overall and the field averaged 30% on text a human had edited and 15% on paraphrased text. A cover letter drafted by a model and then revised by the candidate sits exactly in that band. The number a detector shows you is a confidence, not an accuracy.
Is it legal to reject a candidate based on an AI detector?
The AI use is not the legal problem; the instrument you used to infer it is. A screening device is a selection procedure, and a neutral procedure that disproportionately excludes people on a protected basis has to be job-related and consistent with business necessity. Detectors misfire hardest on non-native English writers, which is national-origin exposure you would have to defend with validation evidence a vendor threshold cannot give you. Ask counsel before building a rule on it.
What false positive rate would be low enough to screen on?
Decide the number before you see a vendor's, by working out what one wrong rejection costs. Then ask for the false positive rate on short, formulaic, human-edited text written by non-native English speakers, because that is what your application pile contains. Published rates come from clean essay corpora and do not transfer. In practice no measured rate has been low enough, which is the answer most teams end up at anyway.
Should you ask candidates to disclose how they used AI?
Yes, and ask everyone in the same words. A one-line field (what you used it for, what you changed) costs a candidate ten seconds and tells you more than any detector, because it produces something specific enough to ask a follow-up about. Treat it as material for the next conversation rather than as a filter. A disclosure used to reject is a detector with extra steps and the same defensibility problem.
What replaces the cover letter if a model writes all of them?
Two or three role-specific questions in a short answer box, asked of every applicant, read against something you wrote down first. You lose the prose and keep the part that was doing the work: what a candidate chooses to say about a concrete situation. Teams that go further drop the letter entirely and move the weight to a work sample, so the evidence is something the process generated rather than something the applicant submitted.
References
- 1. Testing of detection tools for AI-generated text ✓ edintegrity.biomedcentral.com 14 tools tested; best 76% accuracy, none above 80%; 94% on human-written, 30% on human-edited and 15% on paraphrased text; false accusation ratio on human-written text averages 2.4%, zero for half the tools, GPT Zero the outlier; bias toward classifying text as human-written.
- 2. GPT detectors are biased against non-native English writers ✓ doi.org 61.22% average false positive rate on TOEFL essays; 5.19% to 56.65% when word choice is simplified; detection of AI text falls from 100% to 13% after one self-edit prompt.
- 3. Employment Tests and Selection Procedures ✓ eeoc.gov A neutral selection procedure with disproportionate exclusionary effect must be job-related and consistent with business necessity.
3 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.