The US Constitution "Written by AI"? What Detector Fails Teach Us
By AI Detector 360 Editorial Team · · 6 min read

In the summer of 2023, a screenshot made the rounds: the opening of the US Constitution pasted into an AI detector, and a verdict in confident red — most likely AI-generated. The joke wrote itself. James Madison, secret time traveler with a ChatGPT subscription. Underneath the joke sat a genuinely useful lesson about how these tools think, and it's still the best case study we have.
The fail was real: ZeroGPT rated the US Constitution 92.15% AI-generated in widely reported 2023 testing covered by Ars Technica. It wasn't a random glitch. It was a predictable consequence of how detectors measure predictability — and of famous texts sitting inside every language model's training data. Those mechanics still shape every score you read today.
Key takeaways
- The Constitution scored 92.15% AI because models have memorized it, and memorized text reads as maximally predictable.
- Formal, uniform legal prose compounds the effect — the same profile that gets modern human writing flagged.
- In later benchmarking, ZeroGPT couldn't be tuned below a 16.9% false positive rate, so the failure was structural, not a one-off.
- The lasting lesson: a percentage without an error rate and confidence level is theater, not measurement.
The screenshot that defined an era of detector fails
The definitive account came from Ars Technica's Benj Edwards in July 2023, who ran founding-era documents through the popular detectors of the day and watched them misfire. ZeroGPT put the Constitution at 92.15% AI. Other tools of that generation produced their own embarrassments on similar texts, and the story landed in the middle of a bad season for the industry: that same month, OpenAI retired its own AI classifier for "low accuracy" — it had caught just 26% of AI-written text while false-flagging 9% of human writing — and universities were beginning to question the tools they'd rushed to adopt months earlier.
The timing mattered. Students were already being confronted over scores from these same tools. A screenshot proving a detector would convict the founding fathers gave every wrongly accused writer their exhibit A — and it hasn't left the discourse since.
Why does the US Constitution fail the AI detector test?
Two mechanisms stacked on top of each other.
Memorization masquerading as generation. Early detectors leaned heavily on perplexity: how predictable is each next word to a language model? AI text scores predictable because it was assembled from high-probability words. But the Constitution appears countless times in the corpora used to train large language models. A model completing "We the People of the United States..." isn't guessing; it's reciting. GPTZero's founder Edward Tian explained the mechanism to Ars Technica directly: text that's heavily represented in training data is text models are trained to reproduce, so it registers as exactly what a model would write. Maximum predictability, minimum perplexity — indistinguishable, to that metric, from machine output.
Genre uniformity. Eighteenth-century legal drafting is formal, repetitive and structurally even. Long enumerated clauses, parallel constructions, a narrow ceremonial vocabulary. That's a low-burstiness profile, which pushed the second classic signal in the same wrong direction. (For a full tour of both metrics and their blind spots, see perplexity and burstiness explained.)
Neither mechanism involves the detector being broken, exactly. Each metric computed correctly. The failure was in converting two correctly computed statistics into a confident verdict with no acknowledgment of the cases where those statistics mislead.
Check any text for AI — free
Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.
Try the free AI detectorThe number that proves it wasn't a fluke
If the Constitution incident were just an amusing edge case, it would have faded. What kept it relevant is that systematic evaluation later confirmed the underlying problem. In the RAID benchmark (Dugan et al., ACL 2024) — more than 10 million documents, the largest public test of its kind — ZeroGPT could not be tuned below a 16.9% false positive rate. Not on antique parchment: on ordinary human writing, roughly one flag in six was wrong at the tool's best achievable setting.
That's what an error floor means: no threshold setting fixes it. And it reframes the famous screenshot — the Constitution wasn't an exotic input that found a rare bug. It was a vivid demonstration of an everyday failure rate that testing later quantified. We've examined that tool's record separately in is ZeroGPT accurate, and the same era produced failures that hurt real people rather than parchment, like the Texas A&M–Commerce instructor who threatened a whole class's grades on ChatGPT's say-so (Rolling Stone, May 2023).
Institutions did the math and drew conclusions. Weeks after the Constitution story circulated, Vanderbilt University disabled Turnitin's AI detector, publishing its reasoning: at the vendor's claimed 1% false positive rate, roughly 750 of the 75,000 papers it submits each year could be wrongly flagged. The meme and the policy retreat were the same lesson arriving through different doors — unexplained percentages had been asked to carry more weight than they could hold.
What the Constitution fail teaches anyone reading a score
False positives are structural, not accidental. Detectors measure statistical properties that innocent writing can share: formulaic structure, constrained vocabulary, memorized phrasing. The groups this hits — second-language writers, technical writers, anyone working in rigid formats — are documented in why human writing gets flagged. A tool that can't be wrong gracefully will be wrong harmfully.
Extreme scores deserve extra suspicion, not extra trust. 92.15% felt authoritative. But an extreme score on an unusual input is often the model saying "this is far from anything I was calibrated on." The correct output for the Constitution was never a verdict; it was "this text matches known training material — statistical scoring doesn't apply."
A score without an error rate is theater. The 2023 tools reported percentages to two decimal places while publishing no false positive rates at all. That decimal point (92.15%, not 92%) did real rhetorical work, dressing a miscalibrated guess as laboratory measurement. Precision and accuracy are different things. Any detector worth using tells you both what it thinks and how often it's wrong — the record of documented mistakes across the industry is collected in can AI detectors be wrong.
Vivid failures drive better norms. It took a meme to force the conversation that statistics papers had been trying to start: within months of the screenshot, vendors began publishing false positive rates, universities rewrote policies, and "the Constitution test" became shorthand reviewers still use on every new tool. Public, falsifiable embarrassments made the whole field more honest.
How to pressure-test a suspicious score today
The Constitution test is reproducible wisdom. When any score looks off, work through the same logic:
- Consider the text's profile. Is it formal, formulaic, famous or short? Then a high score may reflect genre and memorization, not authorship. The modern equivalents of the Constitution test are everywhere: mission statements, legal boilerplate, pledge-and-motto language, oft-quoted passages. Scanning them individually invites the same artifact — batch them with surrounding original text, or don't score them at all.
- Demand confidence, not just percentage. AI Detector 360 attaches an explicit confidence level to every scan and labels results that rest on thin evidence — by design, because of exactly this history. Our methodology page documents how those levels are calibrated and where our engines are weakest.
- Read the evidence map. A sentence-level heatmap shows whether suspicion is spread evenly or concentrated in a few conventional passages — the difference between a real signal and a Constitution-style artifact.
- Get a second, transparent opinion. Run the text through AI Detector 360's free AI detector (no sign-up, up to 5,000 characters) and compare not the numbers but the reasoning each report shows.
Three years on, detection tools are meaningfully better, and the best now operate with false positive rates the 2023 generation couldn't approach. What hasn't changed is the lesson Madison's prose taught the whole industry: a confident percentage, standing alone, is the least trustworthy sentence a machine can write.
Check any text for AI — free
Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.
Try the free AI detectorFrequently asked questions
Did an AI detector really say the US Constitution was written by AI?
Yes. In July 2023, widely reported testing showed ZeroGPT rating the US Constitution 92.15% AI-generated, and other detectors of that era produced similar false positives on founding documents. The story was covered in depth by Ars Technica and became the defining example of AI detector false positives.
Why do old and famous documents get flagged as AI?
Because they saturate the training data of large language models. A model finds memorized text extremely predictable, and predictability is exactly what perplexity-based detection reads as machine generation. Formal, uniform 18th-century legal prose compounds the effect.
Does the Constitution fail mean AI detectors are useless?
No — it means unaccompanied scores are. The incident exposed tools that reported extreme confidence without confidence levels, sample context or published error rates. Modern detection used honestly, with calibrated confidence and corroborating evidence, remains a useful screening instrument.
Would the Constitution still be flagged by AI detectors today?
Less often and less confidently. Leading detectors have added training data, memorized-text handling and calibration since 2023, and the strongest now hold false positive rates near or below 1% in independent testing. But the underlying mechanism — memorized, formulaic text scoring as predictable — hasn't disappeared, so cautious interpretation still applies.
Sources & further reading
Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.
Related reading

Perplexity and Burstiness: The Science Behind AI Text Detection
What is burstiness in AI detection, and what does perplexity measure? The two statistics behind AI text detectors, explained with examples and honest limits.
Jul 29, 2026 · 6 min read

Can AI Detectors Be Wrong? Yes — Here's How Often
Can AI detectors be wrong? Yes: documented failures, real error rates from independent studies, and a checklist for when to trust or challenge a score.
Jul 15, 2026 · 6 min read

AI Detector False Positives: Why Human Writing Gets Flagged
Why do AI detectors flag human writing? The real causes of AI detector false positives, who gets flagged most often, and what to do when it happens to you.
Jul 6, 2026 · 6 min read