Screening
Can You Reject Interns on an AI-Detector Flag?
An AI-detector flag is not grounds to reject an intern application at any threshold, even when half the pile is flagged. A detector returns a probability with no audit trail, and its errors land unevenly: seven detectors averaged a 61.3% false-positive rate on essays by non-native English writers while classifying US student essays accurately [1]. Decline on it and you are running an unvalidated selection procedure that sorts on English fluency. Route flagged applications to a human read, and let the reader's reasons be the only reasons on record.
The takeHalf your applicants did not suddenly start cheating. What broke is the artifact: a 500-word essay stopped carrying anything about the person the moment every applicant had the same drafting tool, and the detector is a way of not saying that out loud. Watch what the flag rate actually tracks. It rises with the share of your applicants who learned English second. The honest guess is that the funnels keeping the detector are the ones that never sized the reading, and they pay for it in the applicants they never meet.
Where Olive fits
Open a role and see what the work shows
A flag says something about the prose, never about the person who submitted it; Olive assesses the person instead: a 40-to-60-minute assignment built for the occupation, worked with an AI assistant, returned as six findings a human wrote with the moment behind each one. The candidate is granted the same report, free.
Rank your shortlistCan a detector flag be grounds for rejection?
Not on its own, and not safely. A detector returns a probability with no reasoning attached, no per-cohort accuracy you can audit, and no record you could hand to counsel six months later. The moment you decline an application because of that number, the detector is a selection procedure, and the EEOC's own guidance says a selection procedure has to survive a statistical look at who it excludes 2.
Vendors do not make that check easy. Turnitin published a 1% false-positive rate; Vanderbilt did the arithmetic against the 75,000 papers it submitted in 2022, worked out that roughly 750 pieces of student writing would be wrongly labeled, and switched the detector off 3. A campus funnel is the same shape with worse consequences. An intern who never hears why is a person you declined on a number nobody can explain, including you.
What a flag actually tells you is narrow:
- The prose has low lexical variability. That is close to what most detectors measure.
- Nothing about who typed it, and nothing about whether the applicant can do the work.
If half the intern applications are flagged, the interesting number is not 50%. It is that a screen you never validated is now making half your decisions. The general question, do AI detectors work well enough to screen on, has the same answer at every level of seniority. The intern funnel just meets it at volume.
Why does the flag land hardest on your international applicants?
Because most detectors score predictability, and writing in a second language is more predictable. Seven widely used detectors were run over 91 TOEFL essays written by people and 88 essays by US eighth-graders: the US essays were classified accurately, while the TOEFL essays drew an average false-positive rate of 61.3%. At least one detector flagged 97.8% of them, and all seven agreed on 19.8% 1.
The same study fed those TOEFL essays back through a model with one instruction (enrich the word choices), and the average false-positive rate fell from 61.3% to 11.6% 1. Same authors, same arguments, different vocabulary. Read it the other way and the flag is a measure of polish, which is what a well-resourced applicant can buy and a second-language applicant often cannot.
Two groups in a campus pipeline sit on the wrong side of that:
- Applicants writing in English as a second language.
- Applicants trained into a house style (a finance cover letter, a consulting memo, a lab report) where flat, formulaic prose is the correct output.
The EEOC's guidance already lists English proficiency tests among selection procedures 2. A detector pointed at intern essays behaves like one, without the validation an English test would need.
Measure the flag against your own applicant pool
Run the numbers before the policy, not after a charge. Pull the last cycle's applications, apply the threshold you are considering, and compare flag rates across the groups you can lawfully analyze. The same detector at the same setting produces a different exclusion profile in a computer-science funnel than in a marketing one, because the two pools are not composed the same way.
How different depends on the major. Students on temporary visas earned 36% of US science and engineering master's degrees in 2019, and at doctoral level 59% in computer sciences, 58% in engineering and 51% in mathematics and statistics 4. Those are degree figures rather than applicant figures, and that is the point: composition swings by field, so one threshold is a mild filter against one pipeline and a national-origin screen against another.
This is an adverse-impact question you can actually run, and running it costs less than defending a result you never measured. If the flag rate is meaningfully worse for any group, the burden moves to you to show the screen is job-related and consistent with business necessity, and a challenger can still prevail by naming a less discriminatory alternative that predicts performance as well 2. For a detector, there is no validation study to point at.
What should you do with a flagged application instead?
Turn off automatic rejection first. One setting, and it removes most of the exposure. Then decide whether the flag earns a place in the process at all. The defensible version is narrow: no application is declined on a score, a human reads every flagged submission against the same rubric as the unflagged ones, and the reason recorded for any decline is something a person can state out loud.
Three moves, in the order they pay off:
1. Stop the auto-decline. Nothing else matters while a threshold can end an application by itself. 2. Drop the flag from the rubric. If a reader knows the score, it colors the read. If it has to stay, keep it off the reviewer's screen until after the substantive read is written down. 3. Stop asking for the artifact the detector was aimed at. A 500-word "why this internship" essay was never a work sample. It was a writing filter, and it has now been automated on both sides.
What replaces it should be small, role-shaped, and completed by everyone under the same rules, with AI use permitted and disclosed. The rules for AI on an intern take-home are worth setting before the posting opens rather than after the submissions arrive. The gain is a decision you can describe. "Declined because the detector returned 78%" is not a reason. "Declined because the analysis never checked the figure the recommendation rested on" is.
Write the policy before the next posting opens
Four lines, in the posting and in your own records. Whether AI use is permitted, and for what. What the application is assessed on. Whether any tool screens submissions, and that none of them declines anyone. Who to contact about an error. Publish it before applications open, because a rule announced after a rejection reads as a justification for the rejection.
Then keep the records that make the rule checkable: the threshold you used, the date you changed it, flag rates by the groups you analyze, and how many flagged applications a human advanced. Check whether your jurisdiction regulates automated employment decision tools, and do it before any tool touches a decline.
And keep the flag out of the rejection letter. An applicant told that a tool judged their essay to be machine-written will ask how the tool knew, and there is no answer that survives the question. What you can honestly say when explaining an AI-related rejection is bounded by what you could defend in front of the applicant.
Common questions
Is it illegal to reject an applicant because of an AI-detector flag?
No statute names detectors, which is the trap. Once a tool influences who gets declined it is a selection procedure, and the exposure is disparate impact rather than the tool itself. If the flag rate is worse for a protected group, the burden moves to you to show the screen is job-related and consistent with business necessity, and no detector vendor publishes what you would need to make that showing. Keep the flag away from the decline.
What false-positive rate would make a detector safe to reject on?
None, because the rate you would need is not the one vendors quote. A published document-level figure is an average over a corpus you did not submit; what matters is the error rate inside your applicant pool, which no vendor measures and you cannot audit. Even a genuine 1% behaves badly at campus volume: 5,000 intern applications means about 50 people wrongly labeled, none of them visible to you. The number to fix is not the threshold. It is whether a score can decline anyone.
Should you tell an applicant their submission was flagged?
If the flag reaches the decision at all, yes, and that is a good reason to keep it out. An applicant told a tool judged their writing cannot check the finding, cannot appeal it, and will reasonably ask what the tool measured. If the flag only routes work to a human reader and never declines anyone, there is nothing to disclose about the outcome, because the outcome came from the reader.
What if an intern admits they used AI to write the application?
That is a policy question you should have settled in the posting, not a finding. Decide in advance whether AI-assisted writing is permitted, say so in one line, and hold every applicant to the same rule. Teams that write it down mostly land on permitted, and scored no differently, because the essay was never the evidence they were hiring on. An admission is also the one honest signal about authorship that a detector cannot give you.
Where does Olive fit if the flag is not usable?
Olive is not a detector, and no part of it looks at whether a document was written by a model. It is an employer-purchased assessment: the candidate works a task built for their occupation with an AI assistant available, and a human reviewer writes six findings, each anchored to a moment in the session. There is no score, no ranking and no hiring recommendation, and the candidate is granted the same report the employer reads.
References
- 1. GPT detectors are biased against non-native English writers ✓ pmc.ncbi.nlm.nih.gov Seven detectors, 91 TOEFL essays and 88 US eighth-grade essays: 61.3% average false-positive rate on the TOEFL essays, 97.8% flagged by at least one detector, 19.8% unanimously, and 11.6% after a vocabulary rewrite.
- 2. Employment Tests and Selection Procedures ✓ eeoc.gov Selection procedures include English proficiency tests; disparate impact requires statistical analysis, and the employer must show job-relatedness and business necessity against a less discriminatory alternative.
- 3. Guidance on AI detection and why we're disabling Turnitin's AI detector ✓ vanderbilt.edu Turnitin's stated 1% false-positive rate against the 75,000 papers submitted in 2022, roughly 750 wrongly labeled, and the decision to switch detection off.
- 4. Science and Engineering Indicators 2022: International S&E Higher Education ✓ ncses.nsf.gov Temporary-visa students earned 36% of US S&E master's degrees in 2019, and 59% of computer science, 58% of engineering and 51% of mathematics and statistics doctorates.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.