AI Detector 360

What Does an AI Detection Score Actually Mean?

By AI Detector 360 Editorial Team · · 6 min read

Laptop screen with an abstract percentage-style text analysis visualization on a bright desk

A number like "62% AI" looks like a measurement, the way a thermometer reads 62 degrees. It isn't one. It's a model's opinion, expressed on a scale that means different things in different tools — and misreading it is how most detection injustices start.

An AI detection score is a statistical estimate, not a measurement of how much text was AI-written. Depending on the tool, "62% AI" can mean a 62% probability the document is machine-generated, or that 62% of its sentences look AI-like. Those are very different claims, and reading a score correctly starts with knowing which one your tool makes.

Key takeaways

  • Most detection percentages express probability or flagged-share, not the measured proportion of AI authorship.
  • The same essay can legitimately score 15% in one tool and 70% in another because the tools define their numbers differently.
  • A score without a confidence level and a sample-size check is a headline without an article.
  • Thresholds turn scores into verdicts, so the threshold — not the score — is where fairness lives.

One percentage, three possible meanings

When a detector prints a percentage, it's expressing one of three distinct ideas:

Probability of authorship. "We estimate a 62% chance a language model produced this text." The number describes the whole document's likelihood, not any proportion of it. A fully human text can sit at 62% simply because its style is ambiguous to the model.

Share of flagged text. "Our sentence-level model flagged 62% of the sentences." This one looks like a proportion of AI writing, and it still isn't — each flagged sentence is itself just a probability call, with its own error rate. Turnitin, for instance, discloses that its sentence-level false positive rate is around 4%, which means flagged-share numbers carry built-in noise even when the document-level claim is conservative.

Calibrated confidence. "Among past texts with these signals, roughly 62% turned out to be AI." The most useful framing — and the rarest, because honest calibration forces vendors to admit uncertain cases exist.

Most public confusion comes from reading meaning #1 or #2 as if it were a measured fact of composition: "the detector found that 62% of this essay was AI-written." No detector finds that. None can.

Watch the three definitions produce three numbers from one document. Imagine an 800-word, 40-sentence essay. Tool A weighs the whole document and reports "62% likely AI": one probability, no location information. Tool B flags 10 individual sentences and reports "25% AI" — a flagged-share, silent about how confident each flag was. Tool C reports "moderate confidence of AI involvement, concentrated in the middle section." Same essay, numbers ranging from 25 to 62, and not one of them contradicts another; they're answers to different questions. Now imagine a professor comparing Tool A's 62% against a syllabus threshold written with Tool B's arithmetic in mind. That mismatch, multiplied across thousands of classrooms, is the quiet engine of unfair AI accusations.

What does an AI detection score mean in the major tools?

Same percentage sign, different claims underneath:

ToolWhat its number expresses
TurnitinShare of qualifying prose its model predicts was AI-generated, shown only for documents meeting minimum length rules
GPTZeroProbability the document (and each sentence) is AI-generated, with mixed and human classes
ZeroGPTPercentage of the text its model flags as GPT-like
AI Detector 360AI likelihood from an engine ensemble, always paired with an explicit confidence level and per-sentence heatmap

This table is also the answer to the most common support question in this industry: "why did the same text score differently in two tools?" Different definitions, different training data, different thresholds. The RAID benchmark (ACL 2024) documented just how much detectors disagree on identical documents — cross-tool disagreement is the expected state, not a scandal. It's also why comparing your essay's score across five free scanners tells you less than reading one full report carefully.

If you want the mechanics behind these numbers — perplexity, burstiness, classifier training — start with our guide on how AI detectors work.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

Thresholds: where a number becomes a verdict

A score only does something when it crosses a line someone chose. That line is a policy decision disguised as a technical setting.

Set the flag threshold at 50% and you catch more AI but flag more innocent writers. Set it at 90% and the reverse. Every operating point implies a false positive rate, which is why a 2025 NBER working paper by Jabarian and Imas argued institutions should start from the error rate they can tolerate — say, no more than 0.5% of human texts wrongly flagged — and only then ask which tools can operate under that cap. (In their testing, exactly one commercial detector could.)

Real-world score distributions show why threshold placement is so consequential. In its detector's first year, Turnitin reported that of 200 million-plus papers scanned, 11% came back at 20% or more likely AI writing while only 3% hit 80% or more. The mass of scores sits in the ambiguous middle, not at the extremes — which means small threshold moves sweep enormous numbers of borderline documents from "fine" to "flagged." Turnitin itself withholds scores on documents below its length minimums — a tacit admission that under some conditions the honest number is no number.

So when a syllabus or an editorial policy says "anything over X% gets investigated," the real question is: what false positive rate does X imply, on this population of writers? If nobody can answer, the threshold is a vibe. Our guide to what AI percentage counts as cheating works through fair threshold-setting in academic settings specifically.

Reading a score like an analyst

Five checks turn a raw percentage into an informed judgment:

  1. Know the definition. Probability, flagged-share or calibrated confidence? The tool's docs should say; ours are in the report itself and on the methodology page.
  2. Check the sample size. Statistical signals need text. Below a few hundred words, scores swing on coincidence — the thresholds are covered in how much text detectors need.
  3. Look for the confidence level. A 78% score at low confidence is weaker evidence than 62% at high confidence. A tool that shows no confidence level is hiding its uncertainty, not lacking any.
  4. See where the evidence sits. A sentence heatmap distinguishes "uniformly suspicious throughout" from "three formal paragraphs dragged the average up." When you scan with AI Detector 360, the heatmap and engine breakdown make that distribution visible instead of collapsing it into one number.
  5. Corroborate before acting. Detection error rates are real and documented in both directions — the numbers are collected in how accurate AI detectors are. Scores plus process evidence make a case; scores alone make an accusation.

The mindset shift that fixes most misreadings

Stop asking "how much of this was AI?" and start asking "how strong is the evidence of AI involvement?" The first question sounds precise but no tool on earth can answer it. The second is what the score actually addresses — and it naturally invites the follow-ups that make decisions fair: how strong, based on what signals, with what error rate, on how much text?

Base rates belong in that list too. If genuine AI misuse is rare in your context, even a small false positive rate means a meaningful share of flags will be wrong. Vanderbilt University made exactly this arithmetic public in 2023: a 1% false positive rate against its 75,000 yearly papers implied roughly 750 wrongful flags, and the university disabled its detector rather than absorb them. A score means less in a population where the thing it detects barely occurs.

That's the standard we hold our own reports to, and the standard worth demanding from any detector whose number is about to affect a person. A percentage is the beginning of an inquiry. Treat any tool that sells it as the end of one with suspicion.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

Frequently asked questions

Does 40% AI mean 40% of my essay was written by AI?

Usually not. In most tools a percentage expresses either the probability that the text (or each sentence) is AI-generated, or the share of sentences the model flagged — not a measured proportion of AI authorship. A 40% score can appear on a fully human essay whose style happens to look machine-like to the model.

Is a 0% AI score proof that a text is human?

No. Zero means the detector found no statistical evidence of machine generation, which paraphrased or carefully edited AI text can achieve. Research has repeatedly shown detection accuracy collapsing under paraphrase attacks, so a clean score lowers suspicion without eliminating it.

Why do two AI detectors give different scores for the same text?

Because they answer different questions with different models. One tool's 70% may mean "70% probability of any AI involvement," another's "70% of sentences flagged." Add different training data, thresholds and calibration, and disagreement is expected — it's a definition mismatch, not necessarily a malfunction.

What AI detection score should trigger action?

There's no universal cutoff. Sensible practice scales with stakes — a high score with high confidence on a long text justifies a closer look and a conversation, while any score on a short sample justifies almost nothing. Institutions should set thresholds based on the false positive rate they can tolerate, not a round number.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading