Screening
Do AI Detectors Work on Interview Transcripts or Written Responses?
AI detectors do not work well enough on interview transcripts or candidate answers to act on a flag. Seven detectors flagged 61% of essays by non-native English speakers as AI-generated while classifying native writers near-perfectly [1], and OpenAI withdrew its classifier after measuring it at 26% correct on AI text with 9% false positives on human writing [2]. A transcript adds a second uneven error before the detector reads a word, and a flag is a selection procedure you cannot explain or reproduce. Evaluate the work, not who typed it.
The takeDetectors sell because authenticating a document is cheap and reading the work is expensive, and the bill for that shortcut lands on whoever writes most plainly. Usually that is the person writing the way the job taught them to. Anyone determined to beat one can, with a single rewrite prompt, so the flag settles on the candidates who were not hiding anything. I suspect most detector purchases are really a wish to keep the old screen alive one more year. What a flag tells you is that your round has no question a rehearsed answer cannot survive. That is a fact about the round.
Where Olive fits
Open a role and see what the work shows
Authenticating the words is a dead end, so Olive assesses the work instead: a 40-to-60-minute occupational assignment done with an AI assistant, returned as six findings, each anchored to the moment in the session it rests on. The candidate is granted the same report the employer reads.
Rank your shortlistDo AI detectors work on interview transcripts?
Not well enough to act on. The published error rates come from text long enough to give a classifier something to work with; an interview answer or a short written response usually isn't. OpenAI's own classifier was rated unreliable below 1,000 characters (roughly a 150-word answer), and the company withdrew it in July 2023 for low accuracy 2.
The numbers behind that withdrawal are the ones to hold onto. On its own evaluation set, the classifier correctly identified 26% of AI-written English text and labeled human writing as AI 9% of the time, and OpenAI's published guidance was that it should not be used as a primary decision-making tool 2. That is the organization with the most data on how these models write, publishing its own ceiling.
Independent testing lands in the same place. A review of fourteen detection tools (twelve free ones plus Turnitin and PlagiarismCheck) concluded that the available tools are neither accurate nor reliable, that they skew toward calling text human-written, and that ordinary obfuscation makes them significantly worse 3.
The 94%-and-up accuracy figures you see quoted come from vendors testing themselves on their own corpora. Ask any of them for three numbers: the false-positive rate on answers under 200 words, on non-native English, and on text a candidate generated and then edited. Those three decide whether a flag means anything, and they are the three nobody publishes. If the live round is what actually worries you, that is a different problem. See what to do about AI help during a live interview.
Why do detector errors cluster by writing register?
Because the thing detectors measure is predictability, not authorship. Text with low perplexity (common words in an expected order) reads as machine-written to a classifier. Plain, formulaic prose is exactly what a competent engineering postmortem, a consulting deck summary or a claims file note is supposed to be, so the error rate lands unevenly across the roles you hire for.
The clearest published evidence of that skew concerns writers rather than genres. Seven detectors were run over 91 human-written TOEFL essays: 61.3% were misclassified as AI-generated on average, all seven agreed on 19.8% of them, and 97.8% were flagged by at least one tool. The same detectors classified US eighth-grade essays by native speakers almost perfectly 1.
Nothing in that mechanism is specific to language proficiency. It penalizes any writer whose register is constrained by house style, by a template, by a compliance requirement, by training that told them to write short declarative sentences and cut the adjectives. The candidate who writes the way your incident reports are supposed to read is the candidate most likely to be flagged.
The same study showed the failure runs the other way too. Asking a model to enrich the word choices in those essays dropped the average false-positive rate from 61.3% to 11.6% 1. One prompt moved the result by roughly fifty points, which tells you what the result is measuring. A candidate who wants to clear a detector can; the one who gets caught is the one who wasn't trying to. When a round starts producing suspiciously uniform answers, the polish itself is the thing worth reading, not the provenance.
Does the transcript itself add error?
Yes, before any detector reads it. A transcript is a machine's guess at what was said, and the guess is worse for some speakers than others: across five commercial speech recognition systems, average word error rate was 0.35 for Black speakers against 0.19 for white speakers 4. Running a text classifier over that output stacks one uneven error on another.
Consider what a transcription error does to a perplexity score. Every substitution, dropped word and mangled proper noun makes the text less predictable, which pushes a transcript toward "human." Every cleanup the service performs (inserted punctuation, dropped filler words, normalized speaker labels) pushes it back toward "AI." Neither movement has anything to do with the candidate.
Written responses carry a smaller version of the same problem. A candidate typing into your form on a phone, a candidate who drafted in a second language and tidied it up, and a candidate using the grammar checker that ships with their browser all produce text with different statistical fingerprints and identical underlying work. The instrument is reading tooling, not people.
Can you act on a detector flag?
Treat it as a selection procedure, because that is what it is. Any measure used to make an employment decision falls under the EEOC's guidance on tests and selection procedures, and a procedure that screens out a protected group has to be shown job-related and consistent with business necessity 5. No detector vendor publishes validity evidence of that kind.
Set the guidance next to the error data and the exposure states itself. The published false positives concentrate on non-native English writers 1, national origin is a protected class, and a detector score is not job-related to anything: it predicts a text's statistical profile, not job performance. When a candidate asks how the flag was produced, the answer is a proprietary number the vendor will not explain and you cannot reproduce.
What would it take to check a detector for adverse impact? A fixed, inspectable instrument, for a start, and a set of tools an independent review found neither accurate nor reliable is not one 3. The full argument, including which state rules bite first, is in whether it is legal to reject on an AI-detector result.
Replace authentication with evaluation
Stop asking who wrote the answer and start asking what the answer proves. An interview answer is a claim about work, and a follow-up that forces the candidate to operate on their own answer (recompute it, defend the number, name what would make it wrong) cannot be prepared in advance, because it depends on what they just said.
Three changes do most of the work, and none of them need a tool:
- Ask the second question. Take any strong answer and go one level down: which of these numbers would change the recommendation, and by how much? A rehearsed answer has depth of exactly one. Follow-up questions expose understanding faster than any transcript analysis.
- Give them the assistant on purpose. A round that bans AI measures a version of the job nobody has. A round where the model is open, and the candidate has to decide what to hand it and what to keep, produces evidence you can read.
- Score against something written first. Decide what counts as a good answer before the first candidate, apply it to everyone, and record why each answer landed where it did. That record survives a challenge. A detector score does not.
The general version of this trade, dropping the authentication step and moving the burden onto a work sample, is worked through in dropping the resume screen for a work sample.
Common questions
Can GPTZero or Turnitin check an interview transcript?
They will return a number, and the number won't mean what you need it to. Both are built for long-form written text, and an interview answer is short, spoken, and already distorted by transcription. The independent review of fourteen tools, Turnitin included, found them neither accurate nor reliable on the essays they were designed for, before anyone pointed them at speech. A score coming back is not the same as a score you can act on.
What false-positive rate should you assume?
Higher than the vendor's, and unevenly spread. The published academic figure is 61.3% across seven detectors on human-written TOEFL essays, and OpenAI measured its own tool at 9% on general human text before withdrawing it. Whatever threshold you set produces a different burden in every job family, because writing register varies by occupation and by whether English is the candidate's first language. There is no single number to plan around, which is the argument against planning around one.
Is a candidate using AI during a live interview a problem?
It depends on what the interview is for. If the question is one the job itself would let them use a model on, watching them do it is information rather than a violation. If the question exists to test unaided recall, say so before the round starts and pick a format that makes the answer checkable: a shared screen, a live extension of their own answer, a follow-up a prepared script cannot survive.
Is stating your AI rule up front better than checking afterwards?
Yes, and it is the only one of the two that leaves you something to show. A rule written into the invitation, in one sentence, before they start, is a fixed instrument: you can produce the text, apply it to everyone, and say what it meant. A detector result is a number you cannot reproduce or explain, and its errors fall hardest on the plainest writers. A stated rule falls on everyone the same way. It also hands you a key to read the answers against, which is what the flag was standing in for.
Does Olive check whether an answer was written by AI?
No. Olive is an assessment an employer opens for a role: the candidate works a task from their own occupation with an AI assistant available, and a human reviewer writes six findings on what happened (how the problem was framed, what evidence was demanded, what was kept, what was built in between, what was refused, what was tested). There is no score, no ranking and no match percentage, and the candidate is granted the same report the employer reads.
References
- 1. GPT detectors are biased against non-native English writers ✓ pmc.ncbi.nlm.nih.gov 61.3% average false-positive rate on 91 human-written TOEFL essays across seven detectors; 19.8% unanimous, 97.8% flagged by at least one; near-perfect accuracy on native-speaker essays; rewriting dropped the rate to 11.6%.
- 2. New AI classifier for indicating AI-written text ✓ web.archive.org 26% true-positive rate on AI text, 9% false positives on human text; unreliable below 1,000 characters; not to be used as a primary decision-making tool; withdrawn 20 July 2023 for low accuracy.
- 3. Testing of detection tools for AI-generated text ✓ edintegrity.biomedcentral.com Fourteen tools tested, including Turnitin: neither accurate nor reliable, biased toward classifying output as human-written, and significantly degraded by content obfuscation.
- 4. Racial disparities in automated speech recognition ✓ pmc.ncbi.nlm.nih.gov Average word error rate of 0.35 for Black speakers against 0.19 for white speakers across five commercial ASR systems.
- 5. Employment Tests and Selection Procedures ✓ eeoc.gov A selection procedure with disparate impact must be shown job-related and consistent with business necessity.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.