AI Detector 360

How to Check a PDF or Word Doc for AI Writing

By AI Detector 360 Editorial Team · · 6 min read

Stack of printed reports in clear folders beside a document scanner tray on an office desk

In August 2023, Vanderbilt University did the arithmetic on Turnitin's advertised sub-1% false-positive rate and didn't like the answer: across roughly 75,000 papers a year, about 750 wrongly flagged documents. It turned the detector off. Keep that math in mind, because scanning documents is precisely where AI detection meets consequences: theses, client reports, grant proposals and filings, each with a name attached.

An AI detector for documents follows a short pipeline: upload the PDF or DOCX directly or extract clean text, OCR anything scanned, glance at the file metadata, scan with a tool that reports sentence-level results and confidence, and save the report before you act on it. Extraction quality matters more than most people expect.

Key takeaways

  • Direct PDF and DOCX upload beats copy-paste, which mangles hyphenation, headers and footnotes into false signals.
  • Scanned PDFs have no text layer at all; OCR them first or the scan reads nothing.
  • Metadata like author fields and editing time is context, never proof, and it's trivially wrong on shared machines.
  • Save the report with tool, date and confidence level, because a score you can't reproduce later settles nothing.

Choosing an AI detector for documents

The requirements are different from checking a paragraph of pasted text. You want native file ingestion (PDF and DOCX at minimum), sentence-level results rather than one blended score, an explicit confidence label, and an exportable report, because document checks tend to end up in front of other people.

Volume features matter more than they look on day one. Batch handling saves an afternoon when a backlog lands, language support decides whether your non-English reports get a real scan or a shrug, and an upload path that preserves structure decides whether month-three you still bothers. None of this shows up in accuracy marketing, and all of it decides whether the tool gets used.

Cost discipline matters at volume, and so does skepticism. A 2025 NBER working paper by Jabarian and Imas priced commercial detection at $0.02 to $0.06 per document checked, and found that only one of the commercial detectors they tested met a strict 0.5% false-positive policy cap. Both numbers argue the same thing: shop carefully, and never assume the marketing page equals the error rate.

For scale reference, AI Detector 360's document scanner takes PDF and DOCX uploads on the homepage's documents card, prices scans at one credit per 100 words, and includes 300 credits a month on a free account; the paid tiers on our pricing page start at $9.99 for 4,000 credits. For a one-off check of a short document, pasting up to 5,000 characters into the free AI detector needs no account at all.

Get the text out clean

If your tool only takes pasted text, extraction becomes your problem, and PDFs fight back. Copy-paste from a designed PDF typically drags in headers, footers, page numbers and footnote markers, and it breaks hyphenated words at line endings. To a detector, that mess reads as strange, choppy prose and skews the statistics.

FileCommon problemFix before scanning
Scanned PDFNo text layer, nothing to extractRun OCR first
LaTeX or designed PDFBroken hyphenation, ligatures, footnotesPaste to plain text, rejoin words
Slide deck exportFragments and speaker notes interleavedScan notes and body separately
Contract or templateBoilerplate reads as formulaicScan drafted sections, not the template
Legacy .docEncoding artifactsConvert to .docx first

Thirty seconds of cleanup buys you a meaningfully more accurate scan. Or skip the whole problem by uploading the file to a scanner that parses it natively.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

OCR comes before detection

A scanned page is a photograph of words. The quickest diagnosis: try to select text in the file. If nothing highlights, there is no text layer, and an AI detector has literally nothing to read.

Run OCR first, then scan the recognized text. And discount the result accordingly, because OCR introduces its own errors (misread characters, merged words, phantom line breaks) that distort the statistical patterns detectors measure. A flagged scan of OCR output is a weaker signal than the same flag on born-digital text, and an honest reviewer treats it that way.

Most modern PDF tools bundle usable OCR, and mainstream word processors open scanned PDFs with recognition built in. The specific software matters less than the habit: recognize, spot-check a paragraph against the page image, then scan.

What file metadata can and can't tell you

Office documents carry gossip. A DOCX records an author name, a company, revision counts and total editing time; a PDF records the software that produced it and creation versus modification dates. Some of it is genuinely useful context: a 6,000-word report showing four minutes of total editing time invites a question or two.

Now the other side. Metadata is circumstantial, easily wrong and easily edited. Files pass through converters that rewrite producer strings, templates inherit their creator's name, shared machines log the wrong author, and a paste from any source resets the story. Treat metadata the way a good detector treats its own score: as evidence with a known error rate, never a verdict.

A worked example. A consultant's report arrives with a blank author field, created and modified in the same minute. Damning? No: that's what every export from a cloud editor looks like. The same fields turn interesting only against a baseline, like ten earlier reports from the same author that all carry editing history while the eleventh doesn't. Our breakdown of how detectors get things wrong applies to every layer of this workflow, this one included.

Read the results sentence by sentence

Document-level and sentence-level results tell different stories, and the gap is documented: Turnitin claims under 1% false positives at the document level while acknowledging roughly 4% at the sentence level. Whatever tool you use, expect individual highlighted sentences to be the noisiest part of the report.

Documents amplify this because they're full of text nobody drafted: methods sections, legal disclaimers, reference lists, standard clauses. All of it reads as formulaic because it is. A sentence-level heatmap, like the one in every AI Detector 360 report, lets you separate "the analysis section reads as generated, with high confidence" from "the disclaimer looks like every disclaimer ever written." The first is worth a conversation. The second is a shrug.

Length cuts the same way. A 300-word executive summary inside a 40-page report is too short to score on its own, and a document assembled by six authors will read as six different voices, none of them necessarily synthetic. When authorship is the real question, scan sections separately and keep the sample sizes honest.

Save the finding before you act on it

A score glimpsed once and never reproduced settles nothing. Before the conversation, the deliverable or the escalation, record four things: the tool and date, the exact text scanned, the score, and the confidence level. Exporting the PDF report and archiving it alongside the original file takes a minute and survives every later argument about what the scan "really said."

One more habit worth institutionalizing: mind where the documents go. Client reports and unpublished theses are confidential, and pasting them into whatever free tool ranks first is a data-handling decision someone may have to defend later. Use scanners whose retention practices you've actually read.

Then match the follow-up to the context. Editorial teams auditing content at scale should fold this into the workflow in our AI detection guide for SEO teams. Recruiters staring at a suspicious cover letter should read our take on AI-written resumes first, because short documents produce the least reliable scores of all. In every case the rule is the same: a documented, reproducible result invites a conversation. A bare percentage invites a fight.

Check any text for AI — free

Paste up to 5,000 characters into our free scanner, no sign-up. Full multi-engine reports with sentence heatmaps start at $0.

Try the free AI detector

Frequently asked questions

Can I upload a PDF directly to an AI detector?

On tools built for documents, yes. AI Detector 360 accepts PDF and DOCX uploads and scans the extracted text with the same engines as pasted text. For tools without file upload, export the text to plain text first and repair broken line breaks before scanning.

Do AI detectors work on scanned paper documents?

Not until you OCR them. A scanned page is an image with no text layer, so a detector either refuses it or returns nothing meaningful. Run OCR first, scan the recognized text, and treat results with extra caution because recognition errors distort the statistics detectors rely on.

How many credits does a long report cost to check?

On AI Detector 360, one credit covers 100 words, so a 5,000-word report costs 50 credits. A free account includes 300 credits per month, the $9.99 Starter plan carries 4,000, and the $24.99 Pro plan carries 15,000, enough for roughly 1.5 million words of document checking.

Does document formatting change the AI score?

It can, indirectly. Broken hyphenation, embedded headers, footnote markers and template boilerplate all distort the text statistics detectors measure, usually toward machine-like uniformity. Clean extraction, or a scanner that parses the file format natively, removes most of that noise.

Sources & further reading

Fair-use note: AI detection scores — from any tool, including ours — are probabilistic estimates, not proof. Never make academic, employment or legal decisions on a score alone.

Related reading