AI Detector 360

Can Detectors Catch GPT-5-Class Writing?

By AI Detector 360 Editorial Team · · 9 min read

Laboratory bench holding a cloth-covered object surrounded by calibration instruments in cool light

OpenAI's own AI text classifier caught 26% of AI-written text before the company retired it in July 2023 for low accuracy. That statistic usually gets quoted as a story about one weak product, which misses the interesting part: it was built by the lab that built the generator, with full access to the model, and it still failed. Every serious question about catching frontier-model writing starts from that asymmetry.

Detectors can flag GPT-5-class writing, but with lower and less stable accuracy than on older model output. Frontier models sit closer to human statistical distributions, so the signal detectors measure gets thinner, and retraining lags each release by months. Long unedited text still separates reasonably well. Short or paraphrased text does not.

Key takeaways

  • Each generation of model moves closer to human statistical patterns, which narrows the exact gap detectors depend on.
  • Detector retraining always trails a model launch, so the months right after a release are the least reliable window.
  • Sample length and editing history change your odds far more than which detector brand you happen to use.
  • Trained human readers still outperform automated tools on frontier-model text, which is a genuinely uncomfortable finding for our industry.

What "GPT-5-class" actually means for a detector

Detectors don't recognize models. They recognize statistical habits. Nearly every text detector in production measures some version of predictability: given everything written so far, how surprising is the next word? Machine-generated text has historically been less surprising, because generation procedures favor high-probability continuations, and that flatness is what a classifier learns to see. The full mechanics are in our guide to how AI detectors work.

"GPT-5-class" is therefore shorthand for a capability level rather than a product name. It describes any model whose output distribution has moved close enough to human writing that the flatness signal gets faint: longer-range coherence, more varied sentence rhythm, fewer stock transitions, less of the particular blandness that made 2023-era output easy to spot at a glance.

That is worth stating carefully because vendors sometimes market model-specific detection as if they had a fingerprint database. They don't. What they have is a classifier retrained on samples from newer models, which is a different and much shakier thing.

Why it's harder to detect GPT-5 writing than GPT-3.5 output

Picture two overlapping bell curves: human text on one side, machine text on the other. Detection is the business of drawing a line between them. Every capability improvement that makes a model sound more human is, mechanically, a shift of the machine curve toward the human one, and the overlap region is where every error lives.

Three consequences follow, and they compound.

The first is that thresholds get more expensive. To keep false positives low as the curves converge, a vendor has to move the line, which means catching less generated text. To keep catch rates up, it has to accept more wrongly flagged humans. That tradeoff was always present, but on frontier output the price per unit of accuracy rises steeply.

The second is that adversarial handling gets cheaper for the evader. The RAID benchmark (Dugan et al., ACL 2024) ran more than ten million documents through twelve attacks and found commercial detectors degrading sharply under paraphrase and homoglyph substitution. When the underlying separation is already thin, a light paraphrase pass does not need to work hard.

The third is a measurement problem. Benchmarks are built on the models that existed when they were built. A detector that scores well on a public benchmark has proved something about the past, and vendors rarely say out loud how much of their headline accuracy comes from older generations in the test set.

The retraining lag nobody puts on the pricing page

Here is the workflow behind every "now supports the latest models" announcement. A model ships. A vendor collects generated samples at volume, which takes time because you need diversity of prompt, genre and length, not just quantity. The classifier gets retrained. Then it gets validated, because a retrain that raises catch rates while quietly doubling false positives is worse than doing nothing.

Realistically that cycle runs weeks to months, and the validation step is the part under commercial pressure to be short. Which produces a predictable pattern: the least reliable time to run any detector is the window immediately after a major release, exactly when public curiosity about detecting that model peaks.

There is a second, subtler lag underneath the first. Even after a vendor retrains, the population of text people actually submit keeps shifting, because writers adopt new tools unevenly and then edit the output to varying degrees. A classifier tuned on clean, unedited samples from a new model meets a real world of half-generated drafts, dictated notes cleaned up by a chatbot, and human prose polished by an assistant that was itself a frontier model. None of those are the categories the training data was labeled with.

This is also why "supports the newest models" is close to unfalsifiable as a claim. There is no public, continuously updated benchmark that scores commercial detectors against each frontier release within weeks of launch, so nobody outside the vendor can check the assertion at the moment it matters most.

Treat any detector's confidence during the first months after a frontier launch as provisional. If a decision with real consequences is riding on a score from that window, get a second independent reading and weight process evidence more heavily than usual.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

What the numbers do and don't support

Two research results frame the honest position, and they point in opposite directions.

On the pessimistic side: a 2025 NBER working paper by Jabarian and Imas found that of the commercial detectors they tested, exactly one met a strict policy cap of 0.5% false positives, at per-detection costs of roughly two to six cents. One out of a field. That is not an industry-wide failure, but it is a long way from the marketing.

On the optimistic side, and stranger: an ACL 2025 study by Russell, Karpinska and Iyyer found that expert annotators who use ChatGPT frequently identified AI-generated writing with 99.3% accuracy, and stayed robust against evasion tactics that reliably defeat automated tools. Humans who read a lot of model output develop an ear that current classifiers can't match.

Scale turns those percentages into people. Turnitin reported running more than 200 million papers through its AI detector in the first year, with 11% flagged at 20% or more likely AI writing and 3% at 80% or more. Apply even a 1% document-level error to a 200-million-document pipeline and you are describing two million documents whose scores are wrong in one direction or the other. Vanderbilt did the local version of that arithmetic in August 2023 and concluded that 1% of its 75,000 annual papers was about 750 students, then switched the feature off.

A framework for reading a frontier-model score

Since no single number carries a decision, use conditions instead. This is roughly how we'd weight a score on suspected frontier-model text:

ConditionEffect on how much the score is worth
600+ words, unedited, ordinary proseHighest reliability available today
150–300 wordsWeak; treat as a prompt to look closer
Under 150 wordsEffectively noise
Known paraphrase or humanizer passLow score means nothing; high score still means something
Two independent engines agreeMeaningfully stronger than either alone
Formal, formulaic or second-language proseElevated false-positive risk; discount accordingly
Intact C2PA or provenance metadataBeats statistical inference outright

Watch how those rows interact in a real decision. A features editor at a regional magazine commissions a 900-word essay from a freelancer she has published twice before. Her in-house check returns 78%. Two rows of the table apply: the piece is long enough to take seriously, and the freelancer's subject is a formal policy topic where prose naturally runs uniform. She asks for the drafting history and gets a document with four hours of revisions, an abandoned second section, and a paragraph rewritten six times. That evidence outweighs 78% comfortably, and the honest reading is not "the detector was wrong" but "the detector answered a narrower question than the one the editor needed answered."

Change one variable and the conclusion flips. Same score, same writer, but the submission is a 140-word sidebar delivered as a single pasted block with no history. Now the score is worth almost nothing in either direction, and the editor's real options are to ask for a rewrite in a shared document or to accept the piece on trust. Neither of those is a detection problem. Both are process problems that detection was never going to solve.

That last row of the table is the quiet strategic point. Content Credentials embedded by OpenAI since February 2024, by Adobe Firefly, by Microsoft's imaging tools and by Google's 2026 image models are cryptographic assertions rather than probability estimates. They are also routinely stripped when files pass through platforms, which is why provenance is a strong positive signal and a worthless negative one.

If you are setting policy rather than reading one score, this sentence is a defensible starting point to paste into a syllabus, style guide or contributor agreement:

Automated detection results are treated here as a prompt for review, never as a finding. No grade, payment or publication decision is made on a detector score alone.

What we can't verify, including about ourselves

We build a detector, so read this section with appropriate suspicion.

Nobody outside a frontier lab knows how a given detector performs on that lab's newest unreleased model, because the samples don't exist yet. Nobody can tell you the true false-positive rate of any commercial detector on your specific genre, because published rates come from research corpora that probably don't include grant applications, discharge summaries or your company's release notes. And no vendor, ours included, can honestly claim a stable accuracy figure that survives the next model launch, because the target moves.

What we do instead is report uncertainty rather than hide it. Every AI Detector 360 scan carries an explicit confidence level, flags short samples rather than scoring them with false precision, and shows a sentence-level heatmap so you can see which passages drive a score instead of receiving one number to trust or reject. Multi-engine scoring means disagreement between engines surfaces as disagreement, not as an averaged-away middle. How those confidence bands are set is documented on our methodology page.

We do not claim 100% accuracy on frontier-model text. No one honestly can. Our companion piece on how often detectors are wrong puts our own error modes on the same page as everyone else's.

What this changes in practice

For teachers: stop treating the score as the case. A student's drafting history, a five-minute conversation about their argument, and an in-class writing sample together outperform any classifier on frontier output, and the ACL findings suggest your own trained ear is better than you think.

For editors and hiring managers: raise your minimum sample length before you take a score seriously, and stop screening 200-word cover letters. If you need a signal, ask for a live revision or a short verbal walkthrough of a submitted draft.

For writers worried about being flagged: the answer hasn't changed and isn't glamorous. Draft in a versioned editor, keep the history, and check your own text before it matters. AI Detector 360's free AI detector handles 5,000 characters with no sign-up and five scans a day, and the ChatGPT detector view breaks results down by passage so you can see whether a flag clusters in one paragraph or spreads across the piece. For a fuller picture of what accuracy claims are worth, our accuracy explainer walks through the study-by-study record.

The honest summary is that detection still works, less well than it used to, on longer text than people expect, and never well enough to be the last word. That was true of GPT-3.5. It is more true now, and it will be more true again next year.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

Frequently asked questions

Do AI detectors get updated when a new model comes out?

Most do, but not instantly. Detectors are trained on samples of generated text, so a vendor needs volume from the new model before retraining is meaningful, and then needs to validate that the update did not raise false positives. The practical result is a lag measured in weeks to months after each major release.

Is newer AI writing genuinely harder to detect?

On average yes, for a structural reason. Newer models are optimized to produce text that looks more like human writing, which is the same statistical property detectors measure. As the generated distribution moves closer to the human distribution, the separation any detector relies on narrows.

Does a low score mean text definitely was not written by a frontier model?

No. A low score means the tool found no strong statistical evidence of generation, which is a weaker claim than absence. Short samples, heavy editing and paraphrasing all push scores down without changing who wrote the text.

Are paid detectors better than free ones on newer models?

Sometimes, and the difference is usually about retraining cadence rather than the underlying method. A 2025 NBER working paper found only one tested commercial detector met a strict 0.5% false-positive cap, so paying is no guarantee. Check what error rates a vendor publishes rather than what tier you are on.

What is the most reliable signal available on frontier-model text?

Length plus consistency. Long unedited passages give any detector more evidence to work with, and agreement across independent engines is more informative than any single number. Where provenance metadata such as C2PA survives, it beats statistical inference outright.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading