AI Detector 360

Turnitin's AI Score vs Similarity Score: Know the Difference

By AI Detector 360 Editorial Team · · 6 min read

Two transparent rulers crossing over a printed essay page in cool macro daylight

In August 2023, Vanderbilt University switched off Turnitin's AI detector, reasoning that even a 1% false-positive rate would wrongly flag roughly 750 of the 75,000 papers it grades in a year. It kept the similarity checker running. That split decision tells you everything about Turnitin's two headline numbers: they are not the same kind of evidence, and they should never be read the same way.

The Turnitin AI score estimates how likely it is that a document contains machine-generated writing; the similarity score counts how much of it matches existing sources word for word. One is a probability judgment with a real error rate, the other is closer to a measurement. They fail differently, and they justify different actions.

Key takeaways

  • The similarity score measures overlap with a database of existing sources; the AI score is a statistical estimate of how the text was produced.
  • A high similarity score comes with receipts you can inspect, while a high AI score comes with a probability and nothing to check behind it.
  • Turnitin claims under 1% document-level false positives for its AI detector but acknowledges roughly 4% at the sentence level.
  • Neither number is a verdict on its own: similarity needs a review of the matches, and AI scores need process evidence like drafts.

Turnitin AI score meaning: likelihood, not a measurement

The AI writing indicator estimates what share of a submission was likely machine-generated. Under the hood, a classifier scores segments of the document for the statistical patterns typical of language-model output, then aggregates them into the percentage instructors see. The output looks like a measurement. It behaves like a prediction.

The error bars are public, if you know where to look. Turnitin claims under 1% false positives at the document level, a figure that applies to documents scoring at least 20% AI, and acknowledges roughly 4% error at the sentence level. Independent testing has been rougher: in April 2023, the Washington Post ran 16 mixed student samples through the tool and it got over half of them at least partly wrong.

None of that makes the score worthless. It makes it probabilistic, which is a specific thing: a claim about likelihood that will be wrong for a predictable fraction of honest students. And unlike a similarity match, it comes with no receipts. There is no source document to open, no highlighted parallel passage, just a model's opinion about statistical texture. We break down the detection mechanics, and what slips past them, in does Turnitin detect ChatGPT.

Visibility adds a wrinkle. In most institutional configurations, students never see the AI indicator at all; it renders for instructors and administrators only. So the number that can start a misconduct conversation is one the accused student typically has not seen and cannot reproduce, which is worth remembering when tempers rise.

Check your essay before you submit

See your AI likelihood score, sentence-level flags and confidence level — so a detector never surprises you.

Open the AI essay checker

What the similarity score actually counts

The similarity score is the older, better-understood half of the pair. Turnitin compares your submission against web pages, published work and previously submitted papers, then reports the percentage of your text that matches. Unlike the AI score, every point of it is inspectable: the report shows which passages matched which sources, side by side.

That makes it a measurement, but not a verdict. High similarity is often innocent, inflated by quotations, bibliographies and standard methods phrasing, which is why instructors can exclude quotes and references before judging the number. Low similarity proves little in the other direction, since mosaic plagiarism and translated plagiarism can slide under the matcher. Prior submissions surface too, including a student's own earlier papers, which is how most self-plagiarism conversations begin.

And here is the connection people miss: similarity says nothing about AI. Freshly generated text is new text; it usually matches nothing. A 2% similarity essay can be 100% machine-written. The two scores are answering different accusations, which is exactly why Turnitin displays both.

Where each score goes wrong

Side by side, the failure modes barely overlap:

AI scoreSimilarity score
What it claimsText was likely machine-generatedText overlaps existing sources
False alarm looks likeFormulaic or second-language prose flaggedQuotes and templates counted as matches
Miss looks likeParaphrased AI slips throughMosaic or translated plagiarism
What you can inspectHighlighted sentences, nothing behind themThe matched sources themselves
Sensible responseAsk for drafts and process evidenceOpen the matches and check citation

Scale turns those failure rates into people. Turnitin ran over 200 million papers through AI detection in its first year, from April 2023 to April 2024, reporting 11% with at least 20% likely AI writing and 3% at 80% or more. Even a small false-positive percentage inside numbers that size is a stadium's worth of wrongly flagged students, which was precisely Vanderbilt's objection. The AI-score false alarms also cluster unfairly: formulaic, formal and non-native English writing gets flagged at elevated rates, a pattern we document in why human writing gets flagged.

How to read the two numbers together

The scores earn their keep in combination:

  • High AI, low similarity. The classic generated-essay signature: fluent new text that matches nothing. Worth a conversation and a look at drafts, with the false-positive caveats above kept in view.
  • Low AI, high similarity. An old-fashioned citation problem. Open the matches; the AI question is probably a distraction here.
  • High AI, high similarity. Rare and strange, usually boilerplate-heavy documents like templated reports, or generated text that reproduced common phrasing. Read the matches first, because they are the checkable half.
  • Low AI, low similarity. No signal from either system. That is not a certificate of authorship, since paraphrased AI and translated plagiarism can produce exactly this pattern, but nothing on screen justifies suspicion.

Reading the pair is the actual skill. It is the difference between using Turnitin as an instrument and using it as an oracle.

What each score can and cannot justify

A similarity score can nearly settle its own question. Open the matches, check the citations, and you know whether the overlap is quotation, sloppiness or theft. The evidence is right there.

An AI score can justify starting a conversation, comparing the submission with in-class writing, or asking for drafts and version history. It cannot carry a misconduct case alone, and Turnitin itself frames the indicator as a starting point for dialogue rather than proof. A probability with a known error rate convicts no one. Where your school draws its usage lines is a separate question, covered in is using AI for homework cheating.

For instructors, the practical sequence is short. Check the length of the flagged text, since fragments produce noisy scores. Read the flagged passages and ask whether they sound like the student's in-class voice. Then request the document's history before anyone says the word integrity: honest students usually have drafts, and fabricating a writing process on demand is far harder than producing a suspicious percentage.

If you're the student in this story, act before submission, not after the flag. Keep drafts. Run your paper through the AI essay checker to see the same statistical signals a grader's tool will see, sentence by sentence. AI Detector 360 scans combine multiple detection engines, mark each sentence on a heatmap with an explicit confidence level, and export a PDF report you can hand to an instructor; how we calibrate those confidence labels is public on our methodology page. And if the worst happens anyway, our defense guide for falsely accused students walks through the appeal step by step.

Two numbers, two questions, two very different standards of proof. Read them separately and both are useful. Confuse them and someone gets hurt.

Check your essay before you submit

See your AI likelihood score, sentence-level flags and confidence level — so a detector never surprises you.

Open the AI essay checker

Frequently asked questions

Is a 20% similarity score on Turnitin bad?

Not by itself. Similarity counts matched text, which includes quotes, bibliographies and standard phrasing, so a properly cited paper can score 20% or more while being completely honest. What matters is what the matches are. Instructors are expected to open them, not just read the number.

Does the similarity score detect ChatGPT or other AI writing?

No. Similarity compares your text against a database of existing sources, and freshly generated AI prose usually matches nothing. An essay can score 2% similarity and still be entirely machine-written. AI detection is a separate system with a separate score and separate failure modes.

Can students see the Turnitin AI score?

Usually not. In most institutional setups as of mid-2026, the AI writing indicator is shown to instructors and administrators, while students typically see only the similarity report. That asymmetry is one more reason to run your own independent check before submitting.

What Turnitin AI percentage gets students in trouble?

There is no official disciplinary threshold. Turnitin itself gates its reliability claims to documents scoring at least 20%, and institutions define their own review processes. A percentage alone should trigger a conversation and a look at drafts, never an automatic sanction.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading