100 Hours Testing Claude Code vs ChatGPT Codex (honest results)

100 Hours Testing Claude Code vs ChatGPT Codex (honest results)

🎙 Nate Herk 👥 964K 📅 May 26, 2026 ⏱ 26 min 👁 135K 📄 expert opinion 🧭 2026-08-28
Available in: English (current) Français

Keywords

Claude CodeCodexAI codingcomparisonworkflow

Summary

In this video, Nate Herk presents a comprehensive comparison of two leading AI coding agents: Anthropic’s Claude Code and OpenAI’s ChatGPT Codex. Based on 100 hours of hands-on testing, he evaluates them across features, pricing, and three specific use cases: a research report PDF, a landing page, and an interactive dashboard. He highlights Claude Code’s strengths in customization (hooks, sub-agents, slash commands) and creative output, while Codex excels in unified workflow (work trees, in-app browser, computer use) and efficient token usage. The video includes a live breakdown of metrics, showing Codex using significantly fewer tokens and less time for similar tasks. The creator concludes that the best tool depends on the specific use case, recommending both for different workflows. He also notes a philosophical difference: OpenAI allows third-party harnesses like Open Claw, while Anthropic restricts such usage. The video is practical, well-structured, and provides actionable insights for developers choosing between these tools.

152 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides substantial value through its hands-on, comparative approach. The creator’s 100 hours of testing lend credibility to his observations, and he supports his claims with specific examples and metrics (e.g., token usage, time, cost). The argumentation is structured logically, moving from feature comparison to practical tests, and finally to a nuanced verdict. He acknowledges the subjectivity of his ‘gut feeling’ about each tool’s ‘feel’, which adds honesty. The inclusion of real-world use cases (report, landing page, dashboard) makes the information directly applicable. However, the argumentation relies on anecdotal evidence rather than rigorous benchmarks, and the ‘honest results’ are presented without statistical backing.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates a good level of scientific rigor for a practical comparison. The methodology is transparent: same prompts, same conditions, side-by-side testing. The creator cites specific features and versions (e.g., Opus-4.7, GPT-5.5) and provides concrete metrics. However, the sources cited are primarily the tools themselves and the creator’s own experience; no external benchmarks or academic studies are referenced. The title accurately reflects the content, and the video’s structure (timestamps, chapters) aids clarity. The creator’s expertise is evident, but the lack of independent verification limits the overall scientific robustness.

208 words

Title / Content Match

The title accurately reflects the content: a detailed, honest comparison of Claude Code and ChatGPT Codex based on extensive testing.

Quality & Reliability

7/10

The video is a practical, hands-on comparison based on 100 hours of testing, with transparent methodology (same prompts, side-by-side) and specific metrics (time, tokens, cost). However, it relies heavily on subjective impressions and lacks rigorous statistical validation or independent benchmarks. The creator's expertise is evident but the analysis is anecdotal.

Chapters

Cited Sources

Concurring Sources

  • Claude Code documentation — Confirms features like hooks, sub-agents, and slash commands as described in the video.
  • OpenAI Codex documentation — Confirms features like work trees, in-app browser, and computer use as described in the video.

Contribution & Novelties

The video’s main contribution is a practical, side-by-side comparison of two leading AI coding agents, based on extensive real-world testing. It goes beyond feature lists to provide insights into workflow differences, token efficiency, and philosophical stances on third-party integrations. The ‘honest results’ framing and the inclusion of specific metrics (time, tokens, cost) offer a valuable perspective for developers.

Pour aller plus loin :

  • Claude Code documentation — Official documentation for Claude Code, detailing features like hooks, sub-agents, and slash commands.
  • OpenAI Codex documentation — Official documentation for OpenAI Codex, covering its features and usage.
  • Model Context Protocol (MCP) — The open protocol for connecting AI tools to external data and services, mentioned as a key feature in both tools.
  • Open Claw — An open-source agent harness that can be used with ChatGPT subscriptions, illustrating the philosophical difference discussed in the video.
  • Anthropic Agent SDK — The SDK that powers Claude Code, allowing developers to build custom agents.

157 words

Radar Profile

The radar profile shows high scores in information quantity and technical level, indicating a detailed and technically rich video. The lower scores in information quality and reliability reflect the subjective nature of the comparison and lack of external validation. The profile suggests a practical, hands-on review rather than a rigorous scientific study.

Reliability 6/10

💬 Positif. Sur les 30 commentaires analysés, le climat est très positif, avec des remerciements pour la comparaison détaillée et des discussions constructives sur les forces respectives des deux outils, certains partageant leurs propres expériences et suggestions d'amélioration.