Become a Writer Today

AI Detector Accuracy

Evidence report, v1. All sources accessed 2026-07-20. Only published third-party results; BAWT's own test run is pending.

AI writing detectors are being used to reject submissions, cancel book deals and fail students. The vendors advertise false-positive rates around 1 percent or lower. Independent tests paint a different picture: results range from detectors that really do hold near-zero false positives to tools that flag a century-old novel as machine-written. This page puts the vendor claims and the independent measurements side by side, with every figure linked to the source it was read from.

If a detector score has been used against you, read our companion guide: what to do when you are falsely accused of using AI.

What the vendors claim

GPTZero's own benchmarking page states an accuracy rate of 99 percent for AI versus human text and says it keeps its false positive rate at no more than 1 percent (GPTZero, published 2025-01-30, accessed 2026-07-20).Turnitin's public claim, as carried in the comparison below, is a document-level false-positive rate of less than 1 percent.

Conflicted sources
Vendor accuracy numbers are self-studies. The vendor picks the test corpus, the text lengths and the score threshold, and no outside party audits the result before publication. They are listed here for contrast, not as evidence of real-world performance. Every other study on this page is independent of the companies whose tools it measures.

What independent tests find

independentpublished 2026-05-26 · accessed 2026-07-20

Authors Guild detector test

The Authors Guild

Test corpus: Ten Authors Guild articles published in 2022 or earlier, before generative AI tools were widely available, so every AI flag is a false positive.

DetectorFalse positives on human writing
Pangram0 percent across all ten articles
Originality.ai0 percent on eight articles; 1 percent on the other two
Grammarly0 percent on eight articles; 7 percent and 9 percent on the other two
ZeroGPT5 percent to 76 percent, with multiple articles scored above 50 percent AI
Sidekicker.aievery article flagged as predominantly AI-written, scores 71 percent to 100 percent

Finding: Some commonly used consumer-facing AI detection tools are wildly inaccurate. The report notes that polished, edited prose written by experienced human writers shares many characteristics with AI output.

Source

independent, peer-reviewedpublished 2023 · accessed 2026-07-20

GPT detectors are biased against non-native English writers

Liang, Yuksekgonul, Mao, Wu and Zou (Stanford), Patterns / arXiv:2304.02819

Test corpus: 91 human-written TOEFL essays by non-native English speakers and 88 US 8th-grade essays by native speakers, run through seven widely used GPT detectors.

DetectorFalse positives on human writing
Seven detectors (aggregate), TOEFL essaysaverage false positive rate 61.22 percent; more than half of the essays misclassified as AI-generated
Seven detectors (aggregate), US 8th-grade essaysnear-perfect accuracy (near-zero misclassification)
Unanimous flags18 of 91 TOEFL essays (19.78 percent) flagged as AI by all seven detectors; 89 of 91 (97.80 percent) flagged by at least one

Finding: Detectors consistently misclassify non-native English writing as AI-generated while identifying native writing accurately. The authors caution against using GPT detectors in evaluative settings.

Source

independent summary of an independent study, with vendor responsespublished 2026 · accessed 2026-07-20

AI detector false-positive rates: vendor claims vs independent tests

GradPilot (summarising Jabarian and Imas, University of Chicago Booth, BFI Working Paper 2025-116)

Test corpus: The Booth working paper tested 1,992 human and 1,992 AI texts across multiple genres and lengths.

DetectorFalse positives on human writing
GPTZero (vendor claim)around 1 percent claimed; 0.05 percent claimed on GPTZero's own re-run of the Booth data
GPTZero (Booth measurement)at or below 1 percent on medium-to-long texts; up to about 2.4 percent on short texts
Turnitin (vendor claim)less than 1 percent at document level
Turnitin (third-party reads)about 4 percent at sentence level; about 1.4 percent on 300+ word second-language English writing
Pangram (Booth measurement)essentially zero on medium and long passages; the only detector meeting a strict 0.5 percent policy cap
Originality.ai (Booth measurement)at or below 1 percent on medium-to-long texts; up to about 2 to 3 percent on short texts

Finding: Vendor headline numbers are document-level and best-case. The sentence-level and second-language gaps get buried in vendor marketing. GPTZero disputed the Booth methodology in January 2026, arguing the researchers used the wrong API field.

Source

How to read the spread

Three patterns hold across the independent tests. First, the gap between the best and worst tools is enormous: in the Authors Guild test the same ten pre-2022 articles scored 0 percent AI on Pangram and up to 100 percent AI on Sidekicker.ai. Quoting "AI detectors are 99 percent accurate" without naming the tool is meaningless. Second, headline vendor numbers are best-case: document-level scores on long native-English text. Sentence-level scores and short passages run worse, which matters when an editor highlights one paragraph of your manuscript. Third, the errors are not evenly distributed. The Stanford study found seven detectors averaged a 61.22 percent false-positive rate on essays by non-native English speakers while scoring native-speaker essays almost perfectly. Plain, polished, conventional prose is the profile detectors misread most, and the Authors Guild reached the same conclusion about experienced human writers.

BAWT's quarterly creative-writing run

First run pendingMethodology committed 2026-07-20

The published studies above test essays and articles. Nobody regularly tests creative writing, the category this site's readers actually submit. So we are committing our own instrument in advance, before any tool is scored:

The corpus leans deliberately on early-20th-century plain prose (Stein, Anderson, Cather, Lewis, Fitzgerald): the published evidence says clean, uncluttered prose is exactly what detectors misread. The first run has not happened yet. Until it does, this page carries only the third-party results above.

Where this matters for your submissions

Magazines and contests are writing AI rules directly into their guidelines, and publishers are acting on detector scores. See which markets ban, allow or require disclosure of AI at our AI submission policies tracker, and if you self-publish on Amazon, the KDP AI disclosure wizard walks the generated-versus-assisted line. If you have already been accused, start with the defense guide.

Newsletter

The writer's edit

One practical writing tip, a new tool, and a fresh contest deadline, every Thursday. No spam, unsubscribe anytime.