I Battle Tested Sakana Fugu's Fable Killer

I Battle Tested Sakana Fugu's Fable Killer

🎙 Nate Herk 👥 964K 📅 June 23, 2026 ⏱ 12 min 👁 115K 📄 expert opinion 🧭 2026-08-28
Available in: English (current) Français

Keywords

Fugu Ultraorchestrationmulti-agentClaude Opus 4.8cost comparison

Summary

Nate Herk tests Sakana Fugu Ultra, an AI orchestration API that routes tasks to various frontier models (Opus, GPT, Gemini). He compares it against Claude Opus 4.8 across 38 tasks, evaluating quality, speed, and cost. The video explains how Fugu works, its dashboard demo, and its integration with Claude Code. Results show that Fugu Ultra is 4.5 times slower and 5 times more expensive than Opus 4.8, with no significant quality difference in his tests. He concludes that for his use cases, Fugu Ultra does not justify the cost, but he acknowledges the potential of orchestration as a future trend. He also discusses the difference between Fugu and OpenRouter’s Fusion API, and provides a spectrum of orchestration approaches. The video includes a sponsorship segment and promotes his free community resources.

130 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable, hands-on testing data on a newly released AI tool, which is useful for practitioners considering adoption. The argumentation is clear and structured: he explains the concept of orchestration, presents his methodology, and shares quantitative results (time and cost). He also acknowledges the limitations of his test (not heavy software development) and offers a balanced perspective, noting that Fugu might be beneficial for teams. However, the argumentation relies heavily on anecdotal evidence and a single test run, which limits its scientific rigor. The creator’s expertise in AI automation adds credibility, but the lack of detailed code examples or raw outputs weakens the technical depth.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates a reasonable level of scientific rigor by using a controlled test (38 tasks, same prompts, AI-generated assessments) and transparently reporting metrics. However, the test is not peer-reviewed and the grading is done by another AI (Codex), which could introduce bias. The creator cites no external sources beyond his own resources and the Sakana AI website. The title is accurate and not clickbait, as it directly reflects the content. The description includes links to his courses and tools, but these are promotional rather than scientific references. Overall, the rigor is moderate, suitable for a practical review but not for academic citation.

225 words

Title / Content Match

The title accurately reflects the content: the creator tests Sakana Fugu's 'Fable Killer' claim and reports his findings.

Quality & Reliability

7/10

The creator provides a transparent account of his testing methodology, including the number of tasks, the comparison model, and the metrics (quality, speed, cost). However, the test is not a rigorous scientific benchmark; it relies on a single user's experience and AI-generated assessments, which limits generalizability. The video clearly distinguishes between the model's claims and his own findings.

Chapters

Cited Sources

Concurring Sources

Dissenting Sources

  • Sakana Fugu announcement — The announcement claims Fugu matches Fable and Mythos performance, but the creator's tests show no significant quality advantage over Opus 4.8, and higher cost and latency.

Contribution & Novelties

The video provides a practical, real-world evaluation of Sakana Fugu Ultra, an orchestration API, which is a relatively new concept. It offers a comparative analysis against a single model (Claude Opus 4.8) on quality, speed, and cost, which is valuable for practitioners. The creator also introduces the ‘orchestration spectrum’ and contrasts Fugu with OpenRouter’s Fusion API, providing a conceptual framework. The main novelty is the hands-on test data, though the methodology is not exhaustive.

Pour aller plus loin :

  • Mixture of Experts — Relevant to the concept of combining multiple models.
  • Multi-agent system — Directly related to Fugu’s architecture.
  • OpenRouter — The platform mentioned for comparison, though the specific Fusion API page is not linked.
  • Sakana AI — The company behind Fugu, though the specific product page is not linked.

130 words

Radar Profile

The radar profile shows balanced scores across information quantity, quality, technical level, and reliability, with a slight dip in technical depth. This indicates a well-rounded but not deeply technical review, suitable for a general audience interested in AI tools.

Reliability 7/10

💬 Positif. Sur les 30 commentaires analysés, la majorité exprime de la gratitude pour le test et la transparence, certains partagent des expériences similaires, et quelques-uns soulèvent des questions techniques. Aucun commentaire négatif ou haineux n'est présent.