Do ChatGPT Detectors Work? Let's Find Out!

Do ChatGPT Detectors Work? Let's Find Out!

🎙 Matt Wolfe 👥 1.0M 📅 January 10, 2023 ⏱ 11 min 👁 17K 📄 expert opinion 🧭 2026-08-28
Available in: English (current) Français

Keywords

AI detectorChatGPTGPT-3Originality AIDetect GPT

Summary

In this video, Matt Wolfe tests three AI content detectors: GPT-2 Output Detector, Detect GPT (a Chrome extension), and Originality AI (a paid service). He uses text samples he knows are AI-generated (from ChatGPT and GPT-3) and human-written (from his own blog) to evaluate each tool’s accuracy. The GPT-2 Output Detector, though designed for GPT-2, correctly identifies ChatGPT output but fails on some GPT-3 text. Detect GPT performs well on longer texts but struggles with shorter ones. Originality AI, despite being paid, misclassifies some AI-generated content as original and flags some human content as AI. Based on his limited testing, Wolfe concludes that Detect GPT is the most accurate among the three, while acknowledging the limitations of his informal experiment. He also mentions a Reddit thread claiming Originality AI is a scam, which motivated his testing. The video includes a promotional segment for FutureTools.io and a newsletter.

147 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides practical, hands-on information about the current state of AI detectors, which is valuable for content creators concerned about AI content detection. The argumentation is based on direct experimentation, which adds credibility, but the sample size is small and the testing is not systematic. The creator acknowledges the limitations of his approach and does not overstate his conclusions. However, the analysis lacks depth in explaining why certain detectors succeed or fail, and the comparison is not controlled (e.g., text length varies).

Scientific Rigor, Source Quality, Title Accuracy

The video does not cite scientific sources but references the tools tested and a Reddit thread (now deleted) that raised concerns about Originality AI. The description provides links to the tools and the creator’s website. The title accurately reflects the content. The methodology is transparent but not rigorous; the creator does not provide statistical analysis or peer-reviewed backing. The video is more of an anecdotal review than a scientific study.

168 words

Title / Content Match

The title accurately reflects the content: the creator tests three AI detectors and shares his findings.

Quality & Reliability

6/10

The video is an informal, hands-on comparison of three AI detectors. The methodology is transparent but limited (small sample, no statistical rigor), and the conclusions are based on personal testing rather than peer-reviewed research. The creator does not provide deep technical analysis but offers practical insights.

Key Moments

Cited Sources

  • FutureTools.io — The creator's website listing AI tools, used to find the detectors.
  • GPT-2 Output Detector — One of the three detectors tested.
  • Detect GPT — One of the three detectors tested.
  • Originality AI — One of the three detectors tested.
  • Matt Wolfe's blog — Source of human-written content used for testing.
  • Mubert — Outro music generated by Mubert.

Concurring Sources

Dissenting Sources

  • Originality AI — The video suggests Originality AI may not be reliable, as it misclassified some AI-generated content as human. This contradicts the tool's marketing claims of high accuracy.

External References

Contribution & Novelties

The video offers a practical, user-level comparison of three AI detectors, providing real-world examples of their performance. It highlights the limitations of current detectors, especially with short texts and certain types of AI-generated content. This is useful for content creators and marketers who need to understand the reliability of these tools.

Pour aller plus loin :

106 words

Radar Profile

The radar profile shows moderate scores across all dimensions, with a slightly higher score in information quantity and quality, reflecting the practical but limited nature of the content. The low technical level and moderate reliability indicate that the video is more of an informal review than a rigorous scientific analysis.

Reliability 5/10

💬 No comments were provided for analysis.