AI Detector 360

Stylometry: The Old Science Behind AI Detection

By AI Detector 360 Editorial Team · · 8 min read

Archive table holding a handwritten letter, cotton gloves and a magnifying glass in library light

Stylometry is the statistical study of writing style: counting the small, unconscious habits in a text to work out who produced it. That definition is accurate, widely repeated, and the single biggest reason people misread AI detection scores, because it frames the whole enterprise as identifying an author. AI detection does not identify an author. It compares you to a crowd.

Stylometry authorship analysis measures things writers do not choose: function-word frequency, sentence-length distribution, punctuation rhythm. It settled the disputed Federalist Papers in the 1960s, and it underwrites modern AI detection. The difference is that classical stylometry weighs one named candidate against another, while a detector weighs your text against the entire output of machines.

Key takeaways

  • Stylometry works by counting the words writers never think about, because deliberate style is easy to fake and unconscious habit is not.
  • The Federalist Papers case succeeded under conditions AI detection almost never has: two named candidates and thousands of words of verified writing from each.
  • Modern detectors inherited the statistical machinery but changed the question from who wrote this to how machine-like this reads.
  • That shift is what produces false positives on real people whose ordinary prose happens to sit near the machine average.

What stylometry authorship analysis measures

The founding insight is counterintuitive. If you want to identify a writer, ignore everything they are consciously doing. Vocabulary choice, topic, argument structure, favorite metaphors: all of it is imitable, and all of it changes with genre and audience. A person writing a cover letter and the same person writing a group chat message look like different authors by those measures.

What stays stable is the connective tissue. How often you write "upon" instead of "on". Whether you prefer "while" or "whilst". Your ratio of commas to semicolons. The distribution of your sentence lengths, not the average but the spread. Whether you open clauses with "there is". Nobody has opinions about these. They sit below the level where writers make decisions, which is exactly why they persist across topics and why they resist deliberate disguise.

Call it an idiolect. It is closer to handwriting than to voice, and the classical claim of the field is that it is measurable, individual and durable.

The Federalist test that founded the field

The case everyone cites is the right one to cite. Of the eighty-five Federalist essays published under the pen name Publius, authorship of twelve was contested between Alexander Hamilton and James Madison for well over a century. Both men had claimed them. Historians argued from politics and memory and got nowhere.

In the early 1960s the statisticians Frederick Mosteller and David Wallace approached it as a measurement problem. They gathered large bodies of undisputed Hamilton and Madison writing, then hunted for words whose rates differed reliably between the two and had nothing to do with subject matter. The star discriminator was "upon". Hamilton used it constantly; Madison barely at all. Add enough such markers and the twelve disputed essays clustered decisively with Madison.

What makes it a landmark is not the answer. It is that the answer was reached from words nobody would think to fake, under conditions almost ideally suited to the method.

Four conditions made the Federalist result possible: a closed candidate list of two, thousands of words of verified writing from each, a shared genre and era, and no adversary trying to defeat the test. Modern AI detection typically has none of the four.

Function words, the fingerprint you cannot feel

Why do these tiny words carry so much? Because they are frequent and semantically empty. A frequent word gives you a stable rate to measure over a reasonable sample. A semantically empty word does not shift when the topic shifts, so its rate reflects the writer rather than the subject.

Content words behave the opposite way. Write about maritime law and "vessel" spikes, which tells you about the document rather than about you. That is why stylometric feature sets center on determiners, prepositions, conjunctions, auxiliary verbs and punctuation, plus structural measures like sentence-length variance and paragraph rhythm.

Modern extensions add character n-grams, part-of-speech sequences and syntactic tree shapes. The philosophy has not changed since 1964: find the signal in what the writer is not attending to.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

What modern AI detection inherited

Every text detector on the market is a descendant of this tradition, whether or not its marketing says so. The measurements have moved from hand-counted word rates to model-derived probabilities, but the logic is continuous.

Perplexity is a stylometric feature in modern clothing: instead of asking how often you write "upon", it asks how surprising each of your word choices is to a language model. Burstiness is sentence-length variance, restated. The sentence-level heatmap in a good detector is doing what Mosteller and Wallace did by hand, at a resolution they could not have attempted. Our explainer on perplexity and burstiness walks through the modern versions properly.

The inheritance also explains why detectors report distributions rather than verdicts, or at least why the honest ones do. Stylometry was never a yes-or-no method. It produced likelihood ratios, and the analyst stated their uncertainty out loud.

Where AI detection breaks from the tradition

DimensionClassical stylometryAI text detection
Candidate setClosed, usually two or three named peopleOpen, effectively every model ever released
Reference samplesThousands of verified words per candidateUsually none from the actual writer
Question askedWhich of these people wrote itHow machine-like does this read
AdversaryRarely presentRoutinely present, with paraphrase tools
Sample lengthLong documentsOften a few hundred words
OutputLikelihood ratio with stated uncertaintyA percentage, frequently stripped of context

Read down the right-hand column and the failure modes stop being surprising. The task got harder in every row, and the reporting got less careful in the last one.

The candidate-set change is the deepest. Mosteller and Wallace asked a closed question: Hamilton or Madison? A detector asks an open one: does this resemble the aggregate output of language models more than it resembles the aggregate output of humans? Aggregate comparisons have no room for the individual. A person whose natural prose sits near the machine average is not identified as themselves; they are absorbed into the wrong crowd.

That is the entire mechanism behind false positives, and it is structural rather than a bug someone will patch. We take it apart in why human writing gets flagged.

The bias the method carried with it

Aggregate comparison has a predictable victim: anyone whose writing sits further from the reference model's idea of ordinary prose. In 2023, Liang, Zou and colleagues published a study in Patterns that measured it. Seven detectors flagged an average of 61.3% of TOEFL essays written by non-native English speakers as AI-generated, and one flagged 97.8% of them. The same tools were near-perfect on US eighth-grade essays by native speakers.

Sit with those numbers. A second-language writer using careful, correct, slightly formal English produced the statistical profile the tools had learned to call machine-written. The bias was not about ability. It was about distance from a reference distribution built mostly from fluent native prose.

There is a second, quieter group in the same position: writers trained in genres that reward regularity. Legal drafting, technical documentation, clinical notes, standardized exam prose. All of it is deliberately low-variance, because variance is a defect in those genres. A detector reading sentence-length spread as a humanity signal will find these writers suspiciously consistent, and it will be measuring their professional competence.

Classical stylometry avoided this by comparing you to you. Give the method a body of a person's verified writing and it asks whether the new document matches that person's habits, which is bias-resistant in a way population comparison never can be. Give it no reference sample and it has only the crowd.

The RAID benchmark (Dugan et al., ACL 2024) added the other half of the picture: across more than 10 million documents and 12 adversarial attacks, commercial detectors degraded sharply under paraphrase and homoglyph attacks. Population comparison is fragile in both directions. It absorbs innocent writers and it releases determined ones.

Using it without overclaiming

The practical lesson from sixty years of authorship work is that the method is strong when you supply what it needs and weak when you do not. So supply it.

If you run a writing program, collect in-class writing samples early, before any dispute exists. That gives you a personal baseline, which converts an open-set problem into something closer to the Federalist problem: does this document match this student's established habits? It is more work than pasting text into a scanner, and it is worth incomparably more.

A paragraph worth keeping on file for exactly that moment:

Before any discussion of a flagged submission, we compare the document against the writing samples collected from this student at the start of term, and we ask the student to talk us through their drafting process. An automated score is used to decide where to look, never to decide what happened.

AI Detector 360 is built for the "where to look" half of that sentence: multi-engine scoring with explicit confidence levels, a sentence-level heatmap showing which passages drove a result, and downloadable PDF reports so a reviewer can audit the reasoning instead of trusting a number. You can test it on your own writing with the free AI detector at 5,000 characters and no sign-up, or run coursework through our AI essay checker. What our confidence levels mean is documented on the methodology page, and the underlying mechanics are in how AI detectors work.

The open question

Nobody has published a reliable answer to how much AI-assisted editing it takes to erase a person's stylometric fingerprint. Copy editing normalizes precisely the function-word and punctuation habits the method depends on, which means the professionally edited writing we most trust is also the writing that behaves strangest under measurement. If that sounds like an uncomfortable finding for the detection industry, it is, and we would rather name it than wait for someone else to.

Nor is there a settled minimum sample length. Vendors publish thresholds. None of them rests on the kind of independent replication that would let you rely on the number.

Meanwhile the strongest attribution result in the recent literature is not a machine at all. In the 2025 ACL study by Russell, Karpinska and Iyyer, expert annotators who use ChatGPT frequently identified AI-generated text with 99.3% accuracy, holding up against evasion that beat automated tools. Sixty years after two statisticians counted the word "upon" through eighty-five essays, the best detector is still a careful reader who knows the genre. The counting just tells them where to start.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

Frequently asked questions

Is stylometry accepted as evidence in court?

Forensic linguists do testify in some jurisdictions, and their conclusions are generally presented as expert opinion with stated uncertainty rather than as identification. Standards vary widely by country and by court, so treat any blanket claim about admissibility with suspicion.

How much text does authorship attribution need?

More than most people expect. Classical studies worked from thousands of words per candidate author, because the signal lives in frequency distributions that short samples cannot establish. Anything under a few hundred words gives you noise with a decimal point on it.

Can stylometry identify which person in a group wrote something?

Sometimes, when you have substantial verified writing from every candidate and the candidate list is genuinely closed. Remove either condition and reliability drops sharply, which is why open-set attribution on the internet remains an unsolved problem.

Does heavy editing destroy an author's stylometric fingerprint?

It degrades it, and nobody has published a clean threshold for how much editing is enough. Copy editing tends to normalize exactly the function-word and punctuation habits that carry the signal, which is one reason edited professional prose behaves oddly in these systems.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading