AI Detectors and Non-Native English Writers: The Bias Problem
By AI Detector 360 Editorial Team · · 6 min read

In 2023, a Stanford team ran a blunt experiment: feed 91 human-written TOEFL essays to seven commercial AI detectors and count the false accusations. The results were bad enough to change university policy, and three years later they still describe the single largest fairness problem in AI detection.
AI detectors show measurable bias against non-native English writers. In the Stanford study published in Patterns, seven commercial detectors flagged an average of 61.3% of human-written TOEFL essays as AI-generated, and 97.8% were flagged by at least one tool, while essays by native-speaking US students sailed through almost untouched.
Key takeaways
- Stanford's Patterns study found detectors flagged 61.3% of human-written TOEFL essays on average, and 97.8% were flagged by at least one of seven tools.
- The same detectors were near-perfect on essays by native-speaking US eighth graders.
- The bias is mechanical, not intentional: simpler, more predictable phrasing reads as statistically machine-like.
- ESL students can protect themselves with process evidence, and teachers can treat scores as signals rather than verdicts.
The Stanford study, in plain numbers
The design was simple on purpose. Researchers Weixin Liang, James Zou and colleagues collected 91 essays written by real people for the TOEFL exam and ran them through seven popular commercial AI detectors, alongside essays written by US eighth graders. Every essay in the study was human-written. The detectors didn't care.
On average, 61.3% of the TOEFL essays were labeled AI-generated. Nearly every one of them, 97.8%, was flagged by at least one detector, meaning a non-native speaker unlucky enough to face the wrong tool had almost no chance of a clean result. The native-speaker essays, meanwhile, were flagged only about 5% of the time. The authors' warning was unusually direct for an academic paper: unaddressed, this bias risks marginalizing non-native speakers in exactly the settings where writing is evaluated, which is to say school, work and publishing.
The flagged essays weren't borderline cases, either. These were ordinary exam responses, written under time pressure by real people demonstrating English proficiency, which is precisely the kind of prose millions of students produce every week.
No comparable study has overturned the finding since. Vendors have adjusted models and some have improved, but the underlying mechanism hasn't gone anywhere, because it's baked into how detection works.
Why detectors misread non-native writing
AI detectors score predictability. Text where each next word is statistically expected reads as machine-generated, because language models produce exactly that kind of prose; text with surprising word choices and irregular rhythm reads as human. The Stanford team measured what every language teacher already knows: writers working in a second language use a smaller active vocabulary, lean on grammatical constructions they trust, and take fewer stylistic risks. In the study's terms, their essays showed lower lexical richness and less syntactic variety, which collapses straight into "predictable" on a detector's scale.
Notice the cruel loop here. Language instruction explicitly rewards safe, correct, standard constructions; standardized tests are graded on them. Then a detector arrives and treats that hard-won correctness as evidence of a robot. The student is punished not for writing badly but for writing exactly as they were taught.
It's worth being precise about what the bias is and isn't. Detectors don't know anyone's nationality or first language; they punish statistical simplicity wherever it appears. Native speakers with plain, careful styles get caught in the same net, as our guide to AI detector false positives documents. Non-native writers just live where the net hangs lowest.
What AI detector bias means for non-native speakers in practice
The lab numbers became campus reality quickly. In August 2023, The Markup reported on international students being falsely accused of cheating by AI detection tools, and the stakes for that group are uniquely high: grades, scholarships and academic standing, and for some students the enrollment status their visa depends on.
The arithmetic compounds quietly. Turnitin discloses a roughly 4% sentence-level false positive rate on writing in general, and its published reliability claims cover documents with at least 20% flagged text; layer a 61.3% bias effect on top of baseline error, and an ESL-heavy classroom will produce wrongful flags as a matter of routine. This is part of why Vanderbilt, when it disabled Turnitin's AI detector in 2023, cited the risk to non-native English speakers alongside its false-positive math. OpenAI itself retired its classifier in July 2023 over low accuracy after warning it could misjudge non-native writing.
If your institution uses detection, you're entitled to ask two questions: what threshold triggers action, and what the documented appeal path looks like. Schools with good answers to both rarely produce horror stories.
Check your essay before you submit
See your AI likelihood score, sentence-level flags and confidence level — so a detector never surprises you.
Open the AI essay checkerPractical steps for ESL students
You can't rewrite the detectors, but you can make yourself a hard target for a wrongful accusation.
- Draft where timestamps accumulate. Google Docs or Word with AutoSave, every assignment, no exceptions. Version history is the evidence that ends disputes.
- Keep your thinking on paper. Outlines, notes, annotated readings, even notes in your first language. A visible trail from idea to essay is something no chatbot transcript can imitate.
- Write with specifics. Concrete examples from lectures, local details, your own observations. Specificity lowers statistical uniformity and raises grades at the same time; it's the rare fix with no downside.
- Know your baseline before someone else measures it. Running a draft through our AI essay checker shows you the sentence-level heatmap of what detection engines find suspicious, free up to 5,000 characters with no sign-up. AI Detector 360 attaches an explicit confidence level to every scan because a probability is evidence, not proof, and you deserve to see the uncertainty, not just a scary number.
- Never touch "humanizer" tools. Paraphrasing your honest work through evasion software turns a defensible false positive into genuine misconduct, and it erases the natural process trail that would have saved you.
- If an accusation lands, work the process. Our step-by-step defense guide covers exactly what to gather, what to cite, and how to escalate.
Fair-use guidance for teachers
If you teach non-native English speakers, the Stanford numbers should reshape how you use detection tools, without requiring you to abandon them.
Start by knowing your instrument. Every detector has a published or measurable false positive rate; our breakdown of how accurate AI detectors really are collects the current evidence, and a 2025 University of Chicago working paper found that only one commercial tool tested could meet a 0.5% false-positive policy threshold. Whatever the tool's baseline, assume it runs meaningfully worse for your ESL students, because that's what the best available research shows.
Then let the score be a question, never an answer. Compare flagged work against the student's earlier writing, check whether the citations exist, and have the conversation before forming the conclusion; that's the approach Vanderbilt recommended to its own faculty after disabling Turnitin's detector. Weigh process artifacts, drafts, version history and a short chat about the argument, above any percentage. And publish your policy: which tools you use, what thresholds mean, how a student appeals. On our side, we document what AI Detector 360 scores can and cannot claim on our methodology page, because detection only helps education when everyone can see its limits.
What honest detection should look like
The Stanford paper raised the bar for what fair tools owe their users. At minimum: published false positive rates, visible confidence levels instead of bare percentages, sentence-level transparency so a flag can be inspected rather than believed, and explicit warnings on short or simple texts where the statistics run thin. That's the standard we hold ourselves to at AI Detector 360, and it's the standard worth demanding from any tool your institution buys. A detector that hides its uncertainty isn't protecting integrity; it's outsourcing accusations to a black box.
The bias problem is real, measured and unresolved. But between honest tools, evidence-keeping students and teachers who treat scores as signals, it's a problem classrooms can work around, fairly, starting now.
Check your essay before you submit
See your AI likelihood score, sentence-level flags and confidence level — so a detector never surprises you.
Open the AI essay checkerFrequently asked questions
Were the flagged TOEFL essays actually written by AI?
No. All 91 essays in the Stanford study were written by real people taking the TOEFL exam, before ChatGPT existed as a writing shortcut for them. The detectors flagged them anyway, at an average rate of 61.3%, which is what made the result a clean demonstration of bias rather than a judgment call.
Which AI detectors are biased against non-native speakers?
The Stanford team tested seven widely used commercial detectors in 2023 and found the pattern across all of them, so this is an industry-wide property rather than one bad product. OpenAI had already warned that its own, since-retired classifier could misread text by non-native English writers.
What should I do if my essay was flagged and English isn't my first language?
Don't confess to something you didn't do. Export your document's version history immediately, gather drafts and past writing samples, and respectfully cite the Stanford Patterns study showing detectors flag non-native writers at dramatically higher rates. Then follow your institution's formal process, keeping everything in writing.
Can I make my writing less likely to be falsely flagged without cheating?
To a degree. Concrete details, personal observations and naturally varied sentence lengths all lower the statistical uniformity detectors respond to, and they improve the writing anyway. Avoid paraphrasing tools entirely, and keep timestamped drafts, because process evidence protects you in a way no style adjustment can.
Sources & further reading
Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.
Related reading

Falsely Accused of Using AI? A Step-by-Step Defense Guide for Students
Falsely accused of using AI? A step-by-step defense guide: version history, drafts, writing samples, the research on detector errors, and how to escalate.
Jul 13, 2026 · 6 min read

Does Turnitin Detect ChatGPT? How It Works and What It Misses
Does Turnitin detect ChatGPT? Yes, but with limits: a claimed sub-1% document false positive rate, 4% per sentence, and blind spots for paraphrased AI.
Jul 6, 2026 · 6 min read

How Accurate Are AI Detectors in 2026? What Studies Actually Show
How accurate are AI detectors in 2026? What independent studies from Stanford, UPenn and NBER actually found — and how to vet any vendor's accuracy claim.
Jul 6, 2026 · 6 min read