Guide
Do AI detectors actually work?
Short answer: not reliably. Detectors produce false positives on human writing, especially by non-native English speakers, and OpenAI shut its own down. Here is the evidence, and what to do about the tells that real readers notice.
The Stanford result that changed the conversation
In 2023 a team at Stanford tested seven commercial GPT detectors on 91 TOEFL essays written by non-native English speakers. On average the detectors flagged 61.3% of those human-written essays as AI-generated, and more than 91% of the essays were flagged by at least one tool. The same detectors were close to perfect on essays by native-speaking US eighth-graders. The reason is mechanical: detectors lean on "perplexity", a measure of how predictable the word choices are. Careful, textbook English is predictable, so it looks like a machine wrote it.
That matters for any author who learned English as a second language, writes in a plain style, or works in a technical field where the vocabulary is fixed. It also matters if you are publishing in Hindi, Malayalam or Tamil, where detector training data is thin and results are closer to coin flips.
OpenAI gave up on its own detector
OpenAI launched an "AI classifier" in January 2023 and quietly retired it in July 2023 citing its "low rate of accuracy". In OpenAI's own evaluation the tool correctly identified only about 26% of AI-written text while labelling 9% of human text as AI. If the company that built the model could not build a dependable detector for it, be sceptical of anyone selling certainty.
What the vendors themselves admit
- Turnitin claims under 1% false positives on documents that are more than 20% AI-written, but its own blog acknowledges higher false-positive rates on short passages and mixed documents, and told educators not to use the score as the sole basis for any decision.
- A 2023 survey of the field, Towards Possibilities and Impossibilities of AI-generated Text Detection, concludes that as models improve, reliable detection gets harder, not easier, and that simple paraphrasing defeats most detectors.
- Universities have started publishing their own caution notes. The University of San Diego's law library catalogues the false-positive and false-negative problems and warns against treating any score as evidence.
Why this matters for authors
Amazon does not run a public detector on your manuscript; its disclosure rule relies on you telling the truth. Readers, however, run a detector in their heads, and theirs is better than the software. They notice the tidy triads, the "it's not X, it's Y" reversals, the paragraph that ends on a moral, the word "delve". Those are the patterns worth removing, not because a tool will catch them but because a reader will stop trusting the voice.
What to do instead of chasing a detector score
- Disclose honestly where a platform asks, and put a plain line on the copyright page. It removes the risk entirely.
- Fix the tells readers notice. Run a chapter through our free AI tells checker, which counts the patterns rather than guessing a probability, and edit the flagged sentences.
- Put yourself in the text. A specific memory, a named place, a number you actually measured. Detectors cannot fake that, and neither can a model that never met you.
- Do not "humanise" with paraphrasers. They swap words for synonyms and make the prose worse. Rewrite the sentence yourself or cut it.
Neubook Write does the tells pass for you. After each chapter is drafted and edited, it counts the same patterns the checker looks for and rewrites only the sentences that trip them, then shows you the before and after count.
Start a book, freeSources
- GPT detectors are biased against non-native English writers (Liang et al., 2023)
- OpenAI: New AI classifier for indicating AI-written text (retired July 2023)
- TechCrunch: OpenAI scuttles AI-written text detector over low rate of accuracy
- Turnitin: Understanding false positives within our AI writing detection capabilities
- K-12 Dive: Turnitin admits higher false positives in some cases
- Towards Possibilities and Impossibilities of AI-generated Text Detection: A Survey
- University of San Diego LRC: The problems with AI detectors
- Wikipedia: Artificial intelligence content detection