Three years after generative AI became a mainstream concern, determining whether a piece of text was written by a human or a machine remains a frustrating puzzle. AI content detectors have proliferated, but they are not always accurate, and some have actually become less reliable over time. My latest evaluation of 11 content detectors and five AI chatbots shows that a few tools perform well, but the entire category should be approached with caution. More importantly, the tests reveal that freely available chatbots can often match or exceed the accuracy of dedicated detectors.
In this evaluation, I used five blocks of text: two were written by a person, and three were created by ChatGPT. Each detector was asked to classify every block as either human or AI-generated. Any answer that matched the source was counted as correct. This method has been used consistently over the past two years, which allows direct comparisons across versions and changing model capabilities.
The earliest evaluation of this kind, conducted in January 2023, produced poor results. The best detector of that group identified only 66% of the samples correctly. The next major round, in February 2025, used ten detectors, and three of them achieved perfect scores. A few months later, in April, five detectors reached that level. Yet in this latest round, which is roughly half a year later, the number of perfect performers dropped back to three. That demonstrates that no clear upward trend exists; accuracy can shift in either direction as both AI language models and detection algorithms evolve.
Plagiarism in the age of AI
Passing off AI-generated words as your own fits the standard definition of plagiarism. Merriam-Webster defines the term as stealing and passing off the ideas or words of another as one's own, or using another's production without crediting the source. While using an AI writing tool does not involve theft in a strict legal sense, presenting its output without attribution is still plagiarism. That is why editors, teachers, and publishers continue to look for practical and reliable detection methods.
One important risk is that non-native speakers can be caught in false positives. A person whose English writing follows unusual or overly regular patterns may see their genuine work flagged as machine-generated. This issue was present in earlier tests and remains an unresolved concern in 2025.
Comparing dedicated AI content detectors
This round covered 11 dedicated detectors: BrandWell, Copyleaks, GPT-2 Output Detector, GPTZero, Grammarly, Originality.ai, QuillBot, Undetectable.ai, Writer.com, ZeroGPT, and Pangram. One previously included tool, Writefull, was removed because it discontinued its GPT detector. Another, Monica, was dropped because it limited text samples to 250 words and then required a paid upgrade. In place of those, Pangram joined the test and immediately delivered a perfect result.
Two of the five text samples were written by a human, and three were generated by AI. The test required each detector to make one determination per sample. Anything above a 70% confidence level was treated as a strong verdict, and if that verdict matched the true source, the test was passed.
The final tally revealed considerable variation. Pangram, QuillBot, and ZeroGPT correctly identified all five samples. Copyleaks, GPTZero, and Originality.ai each earned an 80% score, getting one sample wrong. The GPT-2 Output Detector scored 60%. BrandWell, Grammarly, and Writer.com were below passing, while Undetectable.ai scored just 20%, making it the least accurate tool in this group.
Pangram: A strong newcomer
Pangram stood out not only because it was new, but also because it had perfect accuracy. The company was founded by engineers who previously worked at Google and Tesla. Its focus is AI detection, rather than the more common 'humanizer' feature that many other products offer. Users receive five free scans per day, which is enough for occasional checks. The scan process is somewhat slow, but the accuracy makes the wait worthwhile.
QuillBot and ZeroGPT: Proven reliability
QuillBot had an uneven history in early tests, sometimes returning different results for identical text across repeated scans. That inconsistency disappeared in the previous round, and QuillBot performed perfectly again this time. ZeroGPT, which was once a bare-bones website with no clear ownership, has turned into a full software service with company details and pricing. It retained its perfect score from previous rounds, showing that the service can maintain quality as it scales.
High-profile misses
Copyleaks published a press release just before this testing round, describing itself as the most accurate AI detector. Its actual performance did not match that claim. Copyleaks flagged a human-written sample as 100% AI-generated, a significant error for a company that sells plagiarism and integrity tools to institutions. Originality.ai, another commercial detector, also misclassified the same human-written block. This was particularly notable because Originality.ai had correctly identified that exact text in the previous round.
GPTZero, which has grown into a company with a mission of 'protecting what is human', delivered an 80% score. But it got a different sample wrong than in its previous test. It now correctly identified a human text that it had missed, but it missed an AI text that it had correctly identified in the previous round. That change shows how much variance can happen between releases.
Other weak performers
Grammarly's AI content checker has not improved, even though the company has promoted it as being out of beta. In this test, Grammarly earned a 40% accuracy score. Writer.com, which offers AI writing tools for corporate teams, also scored only 40%; it classified every single text block, including three AI-generated samples, as human. The GPT-2 Output Detector remains technically frozen in the past, since it was built around an older OpenAI model and appears not to have been updated in a long time.
Undetectable.ai suffered the largest drop in accuracy. In the previous round, it had received a perfect score. This time it rated human writing as 60% likely to be AI, and it rated all three AI-written samples as likely to be human. Given that the service markets itself as a way to make AI content undetectable, these results are puzzling, but the detection side of the service is clearly not consistent.
Chatbots as AI detectors
Because general-purpose chatbots already understand language deeply, they may be better equipped for this task than many specialized tools. I gave the same five text samples to ChatGPT, ChatGPT Plus, Microsoft Copilot, Google Gemini, and Grok, using a simple prompt that asked whether each sample was written by a human or an AI.
ChatGPT Plus, Copilot, and Gemini all achieved perfect scores. The free tier of ChatGPT missed one human-written text sample, but it correctly identified another human sample and even recognized the author of that text from writing style alone, without being given any personal information. That was a surprising and slightly eerie result. Grok, in contrast, failed three of the five samples and tended to label almost everything as human-written.
The implication is clear: for many users, a chatbot is enough to check whether text was likely generated by AI. That means there is less need to buy separate detection software, especially if you only need to review a few documents at a time.
Practical takeaways
No detector is reliable enough to serve as the sole arbiter of authorship. The technology is improving in fits and starts, but it is still prone to serious errors. In this round, some detectors that had previously earned perfect scores declined sharply. Meanwhile, several chatbots proved that low-cost and widely accessible AI tools are capable of strong performance.
If you must decide whether a piece of writing was produced by a person, you should combine automated signals with judgment and context. If you feel the need to verify the accuracy of any tool, create a small benchmark from known human and AI text samples and run it before trusting the tool on real content.
Source: ZDNET News