Evidence report, v1. All sources accessed 2026-07-20. Only published third-party results; BAWT's own test run is pending.
AI writing detectors are being used to reject submissions, cancel book deals and fail students. The vendors advertise false-positive rates around 1 percent or lower. Independent tests paint a different picture: results range from detectors that really do hold near-zero false positives to tools that flag a century-old novel as machine-written. This page puts the vendor claims and the independent measurements side by side, with every figure linked to the source it was read from.
If a detector score has been used against you, read our companion guide: what to do when you are falsely accused of using AI.
GPTZero's own benchmarking page states an accuracy rate of 99 percent for AI versus human text and says it keeps its false positive rate at no more than 1 percent (GPTZero, published 2025-01-30, accessed 2026-07-20).Turnitin's public claim, as carried in the comparison below, is a document-level false-positive rate of less than 1 percent.
The Authors Guild
Test corpus: Ten Authors Guild articles published in 2022 or earlier, before generative AI tools were widely available, so every AI flag is a false positive.
| Detector | False positives on human writing |
|---|---|
| Pangram | 0 percent across all ten articles |
| Originality.ai | 0 percent on eight articles; 1 percent on the other two |
| Grammarly | 0 percent on eight articles; 7 percent and 9 percent on the other two |
| ZeroGPT | 5 percent to 76 percent, with multiple articles scored above 50 percent AI |
| Sidekicker.ai | every article flagged as predominantly AI-written, scores 71 percent to 100 percent |
Finding: Some commonly used consumer-facing AI detection tools are wildly inaccurate. The report notes that polished, edited prose written by experienced human writers shares many characteristics with AI output.
Liang, Yuksekgonul, Mao, Wu and Zou (Stanford), Patterns / arXiv:2304.02819
Test corpus: 91 human-written TOEFL essays by non-native English speakers and 88 US 8th-grade essays by native speakers, run through seven widely used GPT detectors.
| Detector | False positives on human writing |
|---|---|
| Seven detectors (aggregate), TOEFL essays | average false positive rate 61.22 percent; more than half of the essays misclassified as AI-generated |
| Seven detectors (aggregate), US 8th-grade essays | near-perfect accuracy (near-zero misclassification) |
| Unanimous flags | 18 of 91 TOEFL essays (19.78 percent) flagged as AI by all seven detectors; 89 of 91 (97.80 percent) flagged by at least one |
Finding: Detectors consistently misclassify non-native English writing as AI-generated while identifying native writing accurately. The authors caution against using GPT detectors in evaluative settings.
GradPilot (summarising Jabarian and Imas, University of Chicago Booth, BFI Working Paper 2025-116)
Test corpus: The Booth working paper tested 1,992 human and 1,992 AI texts across multiple genres and lengths.
| Detector | False positives on human writing |
|---|---|
| GPTZero (vendor claim) | around 1 percent claimed; 0.05 percent claimed on GPTZero's own re-run of the Booth data |
| GPTZero (Booth measurement) | at or below 1 percent on medium-to-long texts; up to about 2.4 percent on short texts |
| Turnitin (vendor claim) | less than 1 percent at document level |
| Turnitin (third-party reads) | about 4 percent at sentence level; about 1.4 percent on 300+ word second-language English writing |
| Pangram (Booth measurement) | essentially zero on medium and long passages; the only detector meeting a strict 0.5 percent policy cap |
| Originality.ai (Booth measurement) | at or below 1 percent on medium-to-long texts; up to about 2 to 3 percent on short texts |
Finding: Vendor headline numbers are document-level and best-case. The sentence-level and second-language gaps get buried in vendor marketing. GPTZero disputed the Booth methodology in January 2026, arguing the researchers used the wrong API field.
Three patterns hold across the independent tests. First, the gap between the best and worst tools is enormous: in the Authors Guild test the same ten pre-2022 articles scored 0 percent AI on Pangram and up to 100 percent AI on Sidekicker.ai. Quoting "AI detectors are 99 percent accurate" without naming the tool is meaningless. Second, headline vendor numbers are best-case: document-level scores on long native-English text. Sentence-level scores and short passages run worse, which matters when an editor highlights one paragraph of your manuscript. Third, the errors are not evenly distributed. The Stanford study found seven detectors averaged a 61.22 percent false-positive rate on essays by non-native English speakers while scoring native-speaker essays almost perfectly. Plain, polished, conventional prose is the profile detectors misread most, and the Authors Guild reached the same conclusion about experienced human writers.
The published studies above test essays and articles. Nobody regularly tests creative writing, the category this site's readers actually submit. So we are committing our own instrument in advance, before any tool is scored:
detectorCorpus.json, and the extraction is scripted so anyone can reproduce it.The corpus leans deliberately on early-20th-century plain prose (Stein, Anderson, Cather, Lewis, Fitzgerald): the published evidence says clean, uncluttered prose is exactly what detectors misread. The first run has not happened yet. Until it does, this page carries only the third-party results above.
Magazines and contests are writing AI rules directly into their guidelines, and publishers are acting on detector scores. See which markets ban, allow or require disclosure of AI at our AI submission policies tracker, and if you self-publish on Amazon, the KDP AI disclosure wizard walks the generated-versus-assisted line. If you have already been accused, start with the defense guide.
One practical writing tip, a new tool, and a fresh contest deadline, every Thursday. No spam, unsubscribe anytime.