humai← Guide

· 9 min read

Do AI detectors actually work? What the research says

A source-backed look at how accurate — and how unreliable — AI writing detectors really are. Verified numbers from Stanford, OpenAI, Turnitin and Vanderbilt, a table, and citations you can reuse.

Frequently asked questions

Do AI detectors actually work?

The data says they are unreliable. A Stanford study (Patterns, 2023) found an average 61.3% false-positive rate on TOEFL essays written by non-native English speakers, OpenAI retired its own classifier for low accuracy, and Turnitin only claims a sub-1% document-level false-positive rate for documents that are already at least 20% AI. A detector result is a probability, not proof.

What is the false-positive rate of AI detectors?

It depends on the tool and the text. In the Stanford study the average false-positive rate on non-native essays reached 61.3%. OpenAI's classifier mislabelled human text as AI in 9% of cases. Turnitin claims below 1% at the document level, but only for documents with 20% or more AI content, and about 4% at the level of a single sentence.

Why did OpenAI shut down its own AI detector?

OpenAI launched its AI Text Classifier on 31 January 2023 and retired it on 20 July 2023, citing low accuracy. The tool correctly flagged only 26% of text that was actually AI-written and wrongly labelled human text as AI in 9% of cases.

Can an AI detector wrongly flag my own writing?

Yes. False positives are common, especially for people who write plainly, formally or in a second language. The Stanford study found that roughly 61% of genuine human TOEFL essays were flagged as AI. A red score is not evidence — treat it as a weak signal, not a verdict.

Can I cite these statistics?

Yes. You may quote the figures on this page as long as you link back to it as the source of the round-up, and ideally to the original source listed in the table (Stanford/Patterns, OpenAI, Turnitin, Vanderbilt, OPI PIB). Every number here has its primary source cited.