I Tested Opus 5 vs. Fable 5. What You Need to Know.

I Tested Opus 5 vs. Fable 5. What You Need to Know.

🎙 Nate Herk | AI Automation 👥 964K 📅 July 24, 2026 ⏱ 30 min 👁 93K 📄 expert opinion 🧭 2026-08-28
Available in: English (current) Français

Keywords

Claude Opus 5Fable 5benchmarkreal-world testcost per tokenverificationAI agentsworkflow

Summary

The video presents a practical comparison of Claude Opus 5 and Fable 5, two AI models, across a series of real-world tasks including codebase bug fixing, landing page creation, audience research, video outline generation, and a computer-use snake game. The creator, Nate Herk, runs identical prompts through both models within the Claude Code harness, documenting cost, time, and token usage for each. Key findings show that Opus 5 often outperforms Fable 5 on coding tasks and is generally cheaper per token, but Fable 5 tends to produce more visually appealing and creative outputs. The video highlights the importance of verification in AI agent workflows and notes that Opus 5 has improved in this area. The creator also discusses token efficiency, noting that Opus 5 sometimes uses more tokens, negating its cost advantage. The conclusion suggests that the best model depends on the specific use case, with Opus 5 excelling in detail-oriented tasks and Fable 5 in creative and big-picture work. The video includes a breakdown of total costs and tokens across all experiments, and the creator offers free resources for viewers.

181 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides substantial value by moving beyond benchmark claims and testing models in realistic scenarios, offering concrete data on cost, time, and token usage. The argumentation is solid, as the creator explains the methodology and acknowledges limitations, such as the lack of a controlled harness and the subjectivity of some evaluations. The use of a third-party reviewer (Codex) for coding tasks adds credibility. However, the conclusions are based on a limited number of tests and personal preferences, which may not generalize to all users. The creator’s emphasis on verification and prompt engineering is a valuable takeaway, but the overall argument would be stronger with more rigorous experimental controls and a larger sample size.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates a reasonable level of scientific rigor, with a clear methodology for each test and transparent reporting of costs and tokens. The creator references Anthropic’s release blog for Opus 5 improvements, but does not provide direct links to external sources. The description includes links to the creator’s own resources and tools, which are not independent sources. The title accurately reflects the content, and the video’s structure with timestamps helps viewers navigate. The main weakness is the lack of a formal experimental design and the potential for bias in subjective evaluations. The creator’s honesty about the ‘whoopsies’ in one test adds to the credibility, but also highlights the need for more careful execution.

243 words

Title / Content Match

The title accurately reflects the content: a direct comparison of Opus 5 and Fable 5 based on practical testing, with a focus on what users need to know for real-world applications.

Quality & Reliability

7/10

The video provides a hands-on, practical comparison of two AI models across multiple real-world workflows, with detailed cost, time, and token data. However, the methodology is not fully controlled (e.g., one test used the same model twice), and conclusions are partly subjective. The creator is transparent about limitations, but the lack of a formal experimental design and reliance on personal experience limit the scientific rigor.

Chapters

Cited Sources

Concurring Sources

  • Anthropic's Opus 5 release blog — The creator references this blog for improvements in Opus 5's verification capabilities, which aligns with his findings.

Contribution & Novelties

The video contributes practical, hands-on insights into the real-world performance of two leading AI models, moving beyond benchmark scores. It provides a detailed cost and token analysis that is rarely available in such depth, helping users make informed decisions about model selection for specific tasks. The emphasis on verification as a critical factor in AI agent workflows is a valuable addition to the discourse.

Pour aller plus loin :

119 words

Radar Profile

The radar profile shows a balanced performance across all dimensions, with slightly higher scores in quantity of information and technical level, reflecting the video's practical focus and detailed data. The lower score in quality of information is due to the subjective nature of some evaluations and the lack of a fully controlled methodology.

Reliability 7/10

💬 Très positif. Sur les 30 commentaires analysés, la grande majorité exprime une forte appréciation pour l'approche pratique et la transparence du créateur, avec des remerciements et des demandes de contenu supplémentaire.