AI Detector 360

Do AI Detectors Catch Claude, Gemini and Llama — or Just GPT?

By AI Detector 360 Editorial Team · · 6 min read

Four vintage typewriters side by side on a long table, each loaded with paper

In June 2026, GPTZero, the detector famously built by a student to catch ChatGPT essays, was acquired by Superhuman, Grammarly's parent company, with 19 million registered users and roughly $30 million in annual recurring revenue behind it. Not bad for a tool whose founding premise was catching one chatbot. The catch: the whole detection industry grew up aiming at GPT, and GPT stopped being the only model that matters years ago.

Mostly, yes: mainstream AI detectors work on Claude, Gemini and Llama text, not just ChatGPT, because they target statistical regularities that all current language models share. But coverage is uneven. Tools trained mostly on GPT-era corpora can drift on other model families, and error rates typically rise in the months right after a major new release.

Key takeaways

  • Detectors read statistical fingerprints that every current large language model leaves, so the core signal transfers across GPT, Claude, Gemini and Llama.
  • Coverage frays at the edges: fine-tuned open models and brand-new releases are where detectors drift most.
  • The RAID benchmark spans many generators and 12 adversarial attacks; robustness separated tools far more than which model wrote the text.
  • Vendors almost never publish per-model error rates, so treat any specific claim about Claude or Gemini detection as unverified.

An industry raised on GPT

ChatGPT created this market in a single winter. The first wave of detectors trained on what existed: GPT-3.5 and GPT-4 outputs, set against human corpora for contrast. Classifiers learn the distribution you feed them, so those early tools genuinely were GPT specialists, the way a sommelier raised on Bordeaux is a Bordeaux specialist first.

For a while, specialization cost nothing, because the market was the model. Through 2023, if AI text turned up in a classroom or a newsroom, it was overwhelmingly ChatGPT output. Then Claude landed in workplaces, Gemini shipped inside Google's products, and Llama derivatives spread through every startup that wanted to run models on its own hardware. The monoculture ended. The detectors trained on it didn't automatically notice.

The follow-up question wrote itself. If a detector learned GPT's accent, would it recognize Claude's? Skepticism was healthy, and the early record fed it: OpenAI's own classifier, built by the people with the best imaginable access to GPT text, caught only 26% of AI writing before being retired in July 2023. If the house tool struggled on its own models, cross-model detection sounded like wishful thinking.

The 2026 picture is better than that history suggests, for one structural reason.

Do AI detectors work on Claude, Gemini and Llama?

The strongest detection signals are architectural, not brand-specific. Every mainstream model produces text by repeatedly choosing probable next tokens, which leaves output measurably more predictable and more evenly paced than human prose. Those two properties, perplexity and burstiness, transfer across model families because they come from how language models sample, not from any one vendor's training run. Our guide to how AI detectors work walks the full signal stack.

What doesn't transfer is house style. As of mid-2026 the families really do write differently: Claude leans toward carefully hedged, structured prose; Gemini has recognizable formatting habits; and Llama is less a style than an ecosystem, since anyone can fine-tune it into anything. Style-level features need retraining as families evolve, which is why detectors built across many generators age better than GPT-only classifiers.

The genuinely hard case isn't any flagship model. It's the long tail: a Llama fine-tune trained on one company's support tickets, or a model prompted into an unusual persona, drifts away from every distribution a detector has seen. Detection still works there more often than not, but confidence should drop, and an honest tool shows that drop instead of papering over it.

Detection signalSurvives a new model family?
Low perplexity (predictable word choice)Mostly
Even sentence rhythmMostly
House phrasing and formatting habitsPartly; needs retraining
Vendor watermarksNo, one vendor only
Provenance metadataOnly where platforms preserve it

AI Detector 360's engines train on output from multiple model families for exactly this reason, and because engines disagree more on unfamiliar generators, the multi-engine report shows you that disagreement instead of averaging it into false confidence.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

What the benchmarks actually show

The best independent evidence is the RAID benchmark (Dugan et al., ACL 2024): more than 10 million documents spanning a wide range of generators, domains and sampling settings, plus 12 adversarial attacks. Two things matter for this question. First, RAID exists across many generators at all, which lets anyone check cross-model claims instead of taking vendor decks on faith. Second, what separated detectors most wasn't which model wrote the text; it was whether anyone attacked it, because commercial tools degraded sharply under paraphrase and homoglyph substitution regardless of the source model.

There's also a useful human calibration point. In a 2025 ACL study, Russell, Karpinska and Iyyer found that expert annotators who frequently use ChatGPT identified AI-generated text with 99.3% accuracy, staying robust against evasion tactics that beat automated tools. The tells those experts used weren't GPT-specific; they were LLM-specific. That's the encouraging read for cross-model detection: the signal lives in the text itself. The gap lives in the classifiers.

What benchmarks can't give you is a fresh number the week a new model ships. Nobody has that, whatever their homepage implies.

The Turnitin and Gemini question

"Does Turnitin detect Gemini" deserves its own answer because so many people search it. As of mid-2026, Turnitin describes its detector as targeting AI-generated writing in general rather than any single vendor's output, and its headline reliability claim, under 1% document-level false positives, is about protecting human writers rather than about per-model catch rates. What it does not publish is a Gemini-specific or Claude-specific accuracy figure. Neither does anyone else.

So the honest answer runs: assume Turnitin attempts to catch text from every mainstream model, and assume nobody outside the company knows how evenly it succeeds. The pattern repeats across the industry: the question people search is per-model, the answer vendors give is model-agnostic, and the gap between the two is unaudited. The deeper dive into what Turnitin catches, what it misses and how its scores get misread is in does Turnitin detect ChatGPT.

Model drift is the real weak spot

Every detector is a snapshot of the generators that existed when it trained. New releases shift vocabulary, rhythm and formatting, and the months right after a major launch are when false negatives quietly climb, until engines retrain on the new distribution. Paraphrasing stacks on top of that drift: run any model's output through a rewriter and you're attacking a detector's weakest flank, which is why paraphrased and humanized text gets separate treatment on this site.

The practical consequences, in order. Prefer tools that disclose training breadth and attach a confidence level instead of a bare percentage. Rescan contested text after a few weeks when the verdict matters, because engines update. And ignore product names when choosing: despite what it's called, AI Detector 360's ChatGPT detector screens for output from every major model family, with the calibration documented on our methodology page.

If you review other people's writing for a living, build the drift assumption into your process rather than your gut. Note which detector and version produced any score you acted on. Expect engines to disagree more on brand-new model output, and read that disagreement as information about the tools, not as evidence against the writer.

The 2026 question was never whether detectors can see past GPT. They can. It's whether they keep up with whatever ships next quarter, and the only honest answer is to check the date on the benchmark before you trust the badge.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

Frequently asked questions

Does Turnitin detect Claude-written essays?

Turnitin positions its detector as model-agnostic and does not publish per-model accuracy figures, so there is no verified public number for Claude specifically as of mid-2026. Independent research suggests core detection signals generalize across model families but weaken on newly released models and on paraphrased text.

Can GPTZero detect Gemini and Llama text?

GPTZero describes itself as detecting AI writing broadly rather than one vendor's output, and its statistical approach is not tied to GPT alone. Like every detector, though, its published claims are strongest on established model families, so treat any specific per-model accuracy figure as unverified until an independent benchmark reports it.

Are open-source models like Llama harder to detect?

Sometimes. Base outputs carry the same statistical fingerprints as other large language models, but Llama is widely fine-tuned by third parties, which shifts style in ways a detector may never have trained on. Heavily fine-tuned or unusually prompted variants are the harder targets, not the stock model.

Do AI detectors need updates for every new model release?

In practice, yes. A detector is a snapshot of the generators it trained on, and accuracy typically dips in the weeks after a major release until engines retrain. That drift window is a good argument for multi-engine tools, and for rescanning contested text after engines update.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading