Screening

Why Formal Writers and ESL Applicants Get Flagged as AI

Formal writing in a second language is exactly what AI detectors misread most: independent tests keep finding non-native English speakers' writing flagged as AI-generated at rates far above native speakers', and error rates swing sharply with the kind of writing checked. That pattern is measured, and the primary papers are citable. Don't write worse on purpose to sound more human; that trades away the clarity a reviewer wanted. If a detector flags you, ask in writing what produced the flag and offer drafts and version history.

The takeThis isn't a coincidence a vendor forgot to fix. A detector judges text against the prose it was trained on and reads whatever departs from that baseline as an anomaly by design, and a second language learned through study rather than immersion departs from that baseline constantly. That's a structural limit, and knowing the shape of the error is worth more than hoping a future model version behaves differently, since the underlying mismatch is likely to outlast any one detector's version number.

Where Olive fits

Open a role and see what the work shows

No screening step can prove which draft a model helped write, and Olive doesn't try. If an employer sends you an Olive assessment, it is a role-grounded assignment worked openly with an AI assistant, and the report a person writes about it is yours to read too, free, on every tier.

Rank your shortlist

What the Research Actually Found

Seven widely used detectors, tested against 91 human-written TOEFL essays by non-native English speakers, produced an average 61.3 percent false-positive rate: nearly all of those essays were flagged as AI-generated by at least one tool, while the same detectors read US eighth-graders' writing accurately 1.

That gap is not a one-off. A 2025 follow-up ran three detectors over academic abstracts, comparing native and non-native authors directly. On the mixed case, human writing polished with an LLM, one tool's over-detection rate on non-native authors ran more than double its rate on native authors, even though the same tool was highly accurate at telling fully human text from fully AI text 2. The direction repeats across a second, independent study with a different set of documents, which is what makes it a finding rather than a fluke of one paper's sample.

A newer benchmark makes the mechanism explicit: a detector's false-positive rate is less a fixed property of the tool than a setting, and it swings by the kind of writing being checked, by tens of percentage points between one genre and another at the same threshold 3. None of the benchmark's genres are hiring documents, and that is the part that transfers: the rate a tool shows on one kind of writing says very little about the rate it produces on yours.

One university ran the arithmetic on what a small advertised error rate means at real volume, thousands of documents against a single-digit false-positive claim, and turned its own detector off rather than defend the result 4. That decision is worth knowing about even outside a classroom: it's an institution with real stakes concluding that the tool wasn't trustworthy enough to keep running, not an activist argument against detection in general.

Put the four findings together and a pattern emerges that has nothing to do with whether any individual essay or resume actually involved a model. A detector judges writing against the text it was trained on, and one prominent vendor says plainly that its training data is mostly English prose written by adults 5; anything that departs from a tool's baseline, a second language's sentence rhythms included, reads as unusual to it. An unusual pattern is not evidence of AI use. It's evidence that the tool's baseline doesn't match the writer in front of it, which is a limitation of the measurement, not a fact about the person being measured.

Don't Write Worse to Sound More Human

The tempting fix is to loosen your own writing on purpose: shorter sentences, more contractions, an informal aside dropped in, on the theory that sounding more casual reads as more human. Resist it; the evidence that trade actually works is thin, and the real cost of making it is not.

The vendor of one widely used detector tells educators its results shouldn't be used to punish students, and says accuracy is higher for text resembling its training data, which is mostly English prose written by adults; a deliberately looser draft is a guess about that resemblance with no guarantee it pays 5. You'd be trading the clarity that makes your application readable for a guess about what one tool's threshold happens to reward this month, on a document a real person is also going to read and judge on its own terms.

Hiring managers are hearing the same finding from their side of the process. What stops a hiring manager from rejecting someone for 'sounding like AI' is changing what a rejection is allowed to say, since style alone can't fill in a specific, checkable reason. A process built that way leaves no room for 'it read as too polished' as a stated cause, which is the protection this research is actually arguing for on your behalf, whether or not any individual employer has caught up to it yet. Writing plainly and formally, the way you already do, is not the liability the detector made it look like.

Build a Record That Doesn't Depend on Anyone Believing You

If a flag does reach you, the response is the same one that works for any wrongful AI accusation: ask what produced it, in writing, and offer durable process evidence rather than a defense of your prose style. Naming the pattern above, writing by non-native English speakers gets misread most, is a legitimate part of that answer: a documented finding with your name nowhere near it, which is exactly what makes it useful to cite.

What to do if you're wrongly accused of using AI walks through the specific steps, drafts, version history, a short walkthrough, in more detail than belongs here.

A few habits are worth keeping as routine rather than only reaching for them once a flag arrives: save drafts as you go rather than working entirely in one buffer that keeps getting overwritten, note when you asked a tool to help with a specific paragraph, and keep any assignment brief or prompt that shows your process predates the finished document. None of this is about proving innocence in advance. It's about not starting from zero if a flag ever does show up, which for a writer in your position is common enough to plan for once rather than dread every time you submit something.

It also helps to say the pattern out loud when it's relevant, calmly and once. If an interviewer or a recruiter raises a concern about how polished your writing sounds, naming the documented direction of this research, second-language writing gets misread more, is a factual answer rather than a defensive one. You're not asking anyone to take your word for it, and you don't need them to. You're pointing at a measured, published finding that happens to describe exactly what's sitting in front of them right now.

See a sample report

Common questions

Does writing more casually actually lower my risk of being flagged?

It might move a specific tool's score, since detectors respond to surface features like sentence length and word choice, but none of the studies cited on this page show it reliably helps, and it costs you the clarity that made your writing worth reading in the first place. It's not a trade worth making on a guess.

Is this a problem with one detector, or true across most of them?

The pattern shows up across multiple independently tested tools, not one product. The size of the effect varies by tool and by the kind of writing tested, but the direction, non-native writing getting flagged more, repeats across the studies.

Should I mention that English is my second language if I get flagged?

It's reasonable context to offer once you've asked what produced the flag, alongside process evidence rather than instead of it. It explains a pattern; it doesn't replace showing your work has a history.

Do detectors get better at this over time?

Newer versions change, but the underlying mismatch, a training baseline your writing departs from, is a design property that a version bump doesn't automatically fix. Don't assume this year's tool behaves differently from last year's without evidence.

What if the employer used a detector I've never heard of?

Ask for its name directly. A tool with no public track record is one you can reasonably ask an employer to justify relying on, and 'I don't recognize this product' is a fair thing to say out loud.

References

  1. 1. GPT detectors are biased against non-native English writers Patterns (Cell Press), via PubMed Central, 2023. pmc.ncbi.nlm.nih.gov Supports the 61.3% average false-positive rate on TOEFL essays by non-native English writers.
  2. 2. The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication PeerJ Computer Science, via PubMed Central, 2025. pmc.ncbi.nlm.nih.gov Supports that, on human writing polished with an LLM, over-detection on non-native authors ran more than double the rate on native authors.
  3. 3. RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 2024. aclanthology.org Supports that a detector's false-positive rate varies sharply by the kind of writing being checked.
  4. 4. Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector Vanderbilt University (Brightspace / Center for Teaching), 2023. vanderbilt.edu Supports the arithmetic showing what a small advertised false-positive rate means at real document volume.
  5. 5. Frequently asked questions - What are the limitations of the classifier? GPTZero, 2026. gptzero.me Supports that the vendor tells educators results shouldn't punish students, and that its training data is mostly English prose written by adults.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.