TwainGPT is now VervaLearn more

Why AI Detectors Flag Non-Native English Speakers

Why AI Detectors Flag Non-Native English Speakers

AI detectors flag writing by non-native English speakers as AI-generated more often than writing by native speakers, even when a person wrote every word. This pattern appears consistently across studies of how these tools behave.

The reason traces back to the design of most detectors. They measure how predictable the language is. Writers working in a second language often rely on simpler vocabulary and more standard sentence patterns. That predictable style matches the statistical profile detectors treat as a sign of AI.

The gap is larger than most people expect. In one study, researchers pushed false positive rates sharply up or down just by changing the essays' vocabulary, without touching a single idea. Below, we break down what the research found, why it happens, and how professional editing complicates the picture.

Key Findings

  • In a Stanford study, seven AI detectors flagged 61.22% of essays by non-native English speakers as AI-generated, on average.
  • 97.80% of those essays (89 of 91) were flagged by at least one detector.
  • 19.78% (18 of 91) were flagged by all seven detectors at once.
  • When native speakers' essays were rewritten in simpler language, their false positive rate rose from 5.19% to 56.65%.
  • A 2026 study of 13 detectors found false positive rates on human-written academic text ranging from 0% to 100%, depending on the detector.

The Research at a Glance

Study What was tested Main finding
Liang et al. (2023), Stanford 7 detectors, 91 TOEFL essays and 88 US eighth-grade essays 61.22% of non-native essays falsely flagged on average
Park, Jeong & Kim (2026) 13 detectors, 135,389 manuscript pairs before and after professional editing False positive rates from 0% to 100% depending on the detector
Karr et al. (2026) Human-written abstracts across four academic fields 9–15% of unedited human abstracts flagged; non-STEM flagged more than STEM

The Stanford Study

In 2023, Stanford researchers Weixin Liang, James Zou and colleagues tested seven widely used AI detectors on two sets of human-written essays. One set contained 91 TOEFL essays written by non-native English speakers. The other contained 88 essays by US eighth-grade students.

The detectors identified the US essays with near-perfect accuracy. Results on the TOEFL essays differed sharply: on average, 61.22% were labeled AI-generated. Nearly all of them (97.80%) were flagged by at least one detector, and about one in five was flagged by every detector.

Bar chart: AI detectors falsely flagged 61.22% of non-native English essays versus 5.19% of native English essays

None of these essays had been produced with AI. The study appeared in the journal Patterns under the title GPT detectors are biased against non-native English writers.

Why It Happens

Many AI detectors calculate perplexity, a measure of how unexpected each word is given the words that precede it. Language models prefer common, high-probability words, which produces text with low perplexity. Detectors treat that low score as evidence of AI generation.

Diagram of a perplexity scale: AI-generated text and non-native writing sit on the more predictable end, native writing on the less predictable end

Non-native writers often generate the same low-perplexity profile for ordinary reasons:

  • Playing it safe. When you are less confident in a language, you stick to words and grammar you know are correct. That means choosing the most common option again and again, which is exactly what AI models do.
  • Textbook templates. Language courses and exams like the TOEFL and IELTS teach set structures and transitions such as "Furthermore," "In addition," and "In conclusion." These are also some of the phrases AI writing leans on most.
  • Even sentence rhythm. Native writers mix short and long sentences without thinking about it. Second-language writers often keep sentences similar in length and shape, and that uniform rhythm looks machine-like to detectors.
  • Translation and grammar tools. Many non-native writers draft in their first language or polish their work with translation and grammar tools. These tools smooth text toward the most standard phrasing, making it even more predictable.

The writing remains clear and grammatically correct, yet it shares statistical traits with AI output.

The Stanford researchers tested this mechanism directly. They instructed ChatGPT to revise the TOEFL essays with more advanced, native-like vocabulary. The false positive rate dropped from 61.22% to 11.77%. They then simplified the US student essays so the language resembled non-native writing. The false positive rate rose from 5.19% to 56.65%.

Chart: richer vocabulary cut the false positive rate for non-native essays from 61.22% to 11.77%, while simpler vocabulary raised it for native essays from 5.19% to 56.65%

Essay content stayed the same in both experiments. Only word choice changed. The detectors responded to stylistic features rather than to the presence or absence of AI.

Editing Makes It Harder to Interpret

Non-native writers frequently submit work that has passed through professional editors or grammar tools. A 2026 study by Hyeonchu Park, Gahye Jeong and Bugeun Kim examined 135,389 pairs of academic manuscripts recorded before and after professional English editing and ran the pairs through 13 AI detectors.

The tools produced widely divergent results. False positive rates on human-written text ranged from 0% to 100% depending on the detector, and the effect of editing on scores was inconsistent. The researchers concluded that editing style itself acts as a confounding factor, making it difficult for detectors to separate AI authorship from ordinary language polishing (Style as a Confound).

A separate 2026 study found that non-STEM writing was flagged at significantly higher rates than STEM writing, and that unedited human abstracts were flagged 9–15% of the time (Why AI Detection Fails for Academic Integrity).

What This Means

For students and researchers working in a second language, a single AI detection score supplies only limited evidence. A positive flag may simply reflect common features of non-native writing rather than actual AI use.

This limitation helps explain why several universities have disabled or declined to use AI detectors. Frequent concerns include false positives, systematic bias against non-native speakers, and incomplete transparency about how scores are calculated.

That does not mean AI detectors are unreliable. The best detectors are highly accurate on longer AI-generated passages, and the technology keeps improving with each new generation of models. Choosing a well-tested detector makes a big difference, since accuracy varies widely between products. For the fairest results, scores work best alongside drafts, version history, and a conversation with the writer.

If you write in a second language, you can check your text with an accurate AI detector before submitting it to see how detectors are likely to score it.

Frequently Asked Questions

Are AI detectors biased against non-native English speakers?

Research shows they can be. In the Stanford study, the same seven detectors that classified native speakers' essays almost perfectly flagged 61.22% of non-native speakers' essays as AI-generated. The bias comes from detectors treating simple, predictable language as a sign of AI.

Can grammar tools or professional editing trigger AI detectors?

Editing can change detector scores. In the 2026 study of 135,389 edited manuscripts, detector responses shifted inconsistently after professional editing, and the researchers could not reliably separate editing style from AI authorship. Results depend heavily on which detector is used.

Which AI detector is the least biased?

Detectors vary widely. The 2026 study of 13 detectors found false positive rates on human writing ranging from 0% to 100%. See our comparison of which AI detector is the most accurate for how individual tools perform.

How can I bypass AI detectors?

Detectors react to word choice and sentence patterns, so rewriting text with more natural, varied wording can change how they score it. Verva's AI humanizer does this for you, helping your writing sound human and bypass AI detectors while preserving its meaning.

Ready to get started?

Start humanizing your content today.