
100 Hours Testing Claude Code vs ChatGPT Codex (honest results)
Keywords
Summary
152 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video provides substantial value through its hands-on, comparative approach. The creator’s 100 hours of testing lend credibility to his observations, and he supports his claims with specific examples and metrics (e.g., token usage, time, cost). The argumentation is structured logically, moving from feature comparison to practical tests, and finally to a nuanced verdict. He acknowledges the subjectivity of his ‘gut feeling’ about each tool’s ‘feel’, which adds honesty. The inclusion of real-world use cases (report, landing page, dashboard) makes the information directly applicable. However, the argumentation relies on anecdotal evidence rather than rigorous benchmarks, and the ‘honest results’ are presented without statistical backing.
Scientific Rigor, Source Quality, Title Accuracy
The video demonstrates a good level of scientific rigor for a practical comparison. The methodology is transparent: same prompts, same conditions, side-by-side testing. The creator cites specific features and versions (e.g., Opus-4.7, GPT-5.5) and provides concrete metrics. However, the sources cited are primarily the tools themselves and the creator’s own experience; no external benchmarks or academic studies are referenced. The title accurately reflects the content, and the video’s structure (timestamps, chapters) aids clarity. The creator’s expertise is evident, but the lack of independent verification limits the overall scientific robustness.
208 words
Title / Content Match
The title accurately reflects the content: a detailed, honest comparison of Claude Code and ChatGPT Codex based on extensive testing.
Quality & Reliability
7/10
The video is a practical, hands-on comparison based on 100 hours of testing, with transparent methodology (same prompts, side-by-side) and specific metrics (time, tokens, cost). However, it relies heavily on subjective impressions and lacks rigorous statistical validation or independent benchmarks. The creator's expertise is evident but the analysis is anecdotal.
Chapters
Cited Sources
- AI Automation Society (Free Course) — Mentioned as a free resource for AI automation.
- AI Automation Society Plus (Full Courses) — Mentioned as a paid resource for full courses and support.
- Podcast Application — Mentioned as a way to apply for the creator's podcast.
- Uppit AI (Work with me) — Mentioned as a way to work with the creator.
- Glaido (Voice to Text) — Mentioned as a tool for voice-to-text, with a free month offer.
- Hostinger VPS (Claude Code Hosting) — Mentioned as a hosting solution for Claude Code, with a discount code.
Concurring Sources
- Claude Code documentation — Confirms features like hooks, sub-agents, and slash commands as described in the video.
- OpenAI Codex documentation — Confirms features like work trees, in-app browser, and computer use as described in the video.
Contribution & Novelties
The video’s main contribution is a practical, side-by-side comparison of two leading AI coding agents, based on extensive real-world testing. It goes beyond feature lists to provide insights into workflow differences, token efficiency, and philosophical stances on third-party integrations. The ‘honest results’ framing and the inclusion of specific metrics (time, tokens, cost) offer a valuable perspective for developers.
Pour aller plus loin :
- Claude Code documentation — Official documentation for Claude Code, detailing features like hooks, sub-agents, and slash commands.
- OpenAI Codex documentation — Official documentation for OpenAI Codex, covering its features and usage.
- Model Context Protocol (MCP) — The open protocol for connecting AI tools to external data and services, mentioned as a key feature in both tools.
- Open Claw — An open-source agent harness that can be used with ChatGPT subscriptions, illustrating the philosophical difference discussed in the video.
- Anthropic Agent SDK — The SDK that powers Claude Code, allowing developers to build custom agents.
157 words
Radar Profile
The radar profile shows high scores in information quantity and technical level, indicating a detailed and technically rich video. The lower scores in information quality and reliability reflect the subjective nature of the comparison and lack of external validation. The profile suggests a practical, hands-on review rather than a rigorous scientific study.
💬 Positif. Sur les 30 commentaires analysés, le climat est très positif, avec des remerciements pour la comparaison détaillée et des discussions constructives sur les forces respectives des deux outils, certains partageant leurs propres expériences et suggestions d'amélioration.