Key Takeaways
- Sakana Fugu and OpenRouter Fusion are both multi-agent orchestrators. Every other competitor in this article is a single frontier model. These are not the same category of product.
- Fugu Ultra leads Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on 10 of 11 Sakana-reported benchmarks. The one loss is MRCRv2 long-context recall, where GPT-5.5 edges ahead.
- Fable 5 and Mythos are not in Fugu’s agent pool due to US export controls. The “parity” claim is Sakana’s, not a head-to-head result.
- Fugu costs roughly 4x less than OpenRouter Fusion for the same prompts, and also offers flat-rate subscription plans Fusion does not.
- Fugu is not available in the EU or EEA at launch. If your team is based in Europe, this comparison ends here.
- All benchmark numbers below are vendor-reported. No independent third-party lab has replicated them as of July 2026.
Before You Compare: Get the Category Right
Most confusion in “Fugu vs X” comparisons comes from treating all five options as the same type of product. They are not.
| Product | Type | What It Actually Is |
|---|---|---|
| Sakana Fugu | Multi-agent orchestrator | Coordinates Opus 4.8, GPT-5.5, Gemini and others behind one API |
| OpenRouter Fusion | Multi-agent synthesizer | Queries multiple models and blends their replies after the fact |
| Claude Fable 5 | Single frontier LLM | Anthropic’s high-end long-horizon model (export-controlled) |
| Claude Mythos Preview | Single frontier LLM | Anthropic’s most capable model (export-controlled, gated) |
| Claude Opus 4.8 | Single frontier LLM | Anthropic’s previous high-performance model (still available) |
| Claude Opus 5 | Single frontier LLM | Anthropic’s current near-frontier model, launched July 24, 2026 |
| GPT-5.5 | Single frontier LLM | OpenAI’s current production flagship |
Fugu does not replace Fable 5, Mythos, Opus, or GPT-5.5. It orchestrates some of them. Choosing Fugu means choosing a coordination layer, not a better underlying model.
Fugu vs Fusion: Two Different Control Flows
Fugu and Fusion look similar on the surface. Both accept one request and return one answer using multiple models. But the internal logic runs in opposite directions.
- Fugu decides upfront which models to call and in what order. It is a conductor: it assigns tasks before execution.
- Fusion synthesizes after the fact. It queries multiple models in parallel and blends the outputs once they arrive.
One Hacker News commenter framed it well: ask GPT to derive the math, ask Opus to check it for security issues, ask Gemini to resolve the disagreement. That is Fugu’s model. Fusion is closer to a voting booth. The distinction matters when you are choosing between them for structured, multi-step tasks.
Sakana Fugu vs Fable 5
This is the comparison that drove the most interest at launch, and it requires the most context to read correctly.
What Sakana Actually Claims
Sakana does not claim Fugu beats Fable 5. The official language is “shoulder-to-shoulder.” That is a deliberate choice of words.
Fable 5 is not in Fugu’s agent pool. It was pulled from non-US access due to US export controls ten days before Fugu launched. The “parity” claim is based on Sakana’s benchmarks comparing Fugu against Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, then citing Fable 5’s publicly reported scores from prior third-party evaluations alongside those numbers.
That is not a head-to-head. It is a side-by-side on separate test runs.
Benchmark Comparison: Fugu Ultra vs Fable 5
| Benchmark | Fugu Ultra | Fable 5 | Notes |
|---|---|---|---|
| SWE-Bench Pro | 73.7 | 86.0 | Fable 5 wins clearly. 12+ point gap. |
| TerminalBench 2.1 | 82.1 | 80.4 | Fugu Ultra leads |
| LiveCodeBench | 93.2 | n/a | Fable 5 not independently published here |
| GPQA-Diamond | 95.5 | n/a | Fable 5 not independently published here |
| Humanity’s Last Exam | 50.0 | 53.3 | Fable 5 leads |
The honest read: Fugu Ultra holds its own on most coding and reasoning benchmarks. On SWE-Bench Pro, the hardest real-world software engineering benchmark in this comparison, Fable 5 leads by nearly 13 points. That gap is real and not trivial.
Who Should Choose Fugu Over Fable 5?
Choose Fugu if:
- You are based outside the US and cannot access Fable 5 due to export controls
- You want vendor diversification rather than dependence on one Anthropic model
- Your primary use cases are code review, multi-step research, or reasoning tasks rather than complex software engineering
Choose Fable 5 if:
- You are a US-based team with access and your workload is heavily weighted toward software engineering
- You need a single model you can audit, with transparent query routing
- SWE-Bench Pro scores are directly relevant to your production use case
Sakana Fugu vs Mythos Preview
Mythos Preview is Anthropic’s most capable model. It is gated, not publicly available, and restricted to a small set of trusted organizations through Anthropic’s Project Glasswing.
What the Comparison Actually Means
Mythos is in the same position as Fable 5 in this comparison: not in Fugu’s pool, not independently benchmarked against Fugu directly, and not accessible to most organizations.
Sakana claims Fugu Ultra performs “shoulder-to-shoulder” with Mythos Preview on frontier benchmarks. That claim cannot be verified through a head-to-head test because Mythos is not publicly accessible.
Benchmark Comparison: Fugu Ultra vs Mythos Preview
| Benchmark | Fugu Ultra | Mythos Preview | Notes |
|---|---|---|---|
| SWE-Bench Pro | 73.7 | Not published | Cannot compare directly |
| Humanity’s Last Exam | 50.0 | Not published | Cannot compare directly |
| GPQA-Diamond | 95.5 | Not published | Cannot compare directly |
There is no meaningful benchmark comparison available to verify or refute Sakana’s claim about Mythos parity.
Who Should Choose Fugu Over Mythos?
This is not a real choice for most organizations. Mythos Preview is gated and not available on demand.
If you do not have Mythos access, Fugu Ultra is a credible alternative for complex research and reasoning tasks based on what Sakana has published.
If you do have Mythos access, run your own evaluation on 20 to 50 representative tasks before making any routing decision based on marketing claims from either side.
Sakana Fugu vs Claude Opus 4.8
This is the most apples-to-apples comparison in the article. Opus 4.8 is publicly available, independently benchmarked, and it is one of the models inside Fugu’s own agent pool.
Benchmark Comparison: Fugu Ultra vs Claude Opus 4.8
| Benchmark | Fugu Ultra | Opus 4.8 | Winner |
|---|---|---|---|
| SWE-Bench Pro | 73.7 | 69.2 | Fugu Ultra |
| LiveCodeBench | 93.2 | n/a | Fugu Ultra (no Opus score published) |
| GPQA-Diamond | 95.5 | 92.0 | Fugu Ultra |
| TerminalBench 2.1 | 82.1 | n/a | Fugu Ultra (no Opus score published) |
| Humanity’s Last Exam | 50.0 | 49.8 | Fugu Ultra (narrow margin) |
| CTI-REALM (cybersecurity) | n/a | 69.6 | Opus 4.8 edges ahead |
Fugu Ultra leads Opus 4.8 across most published benchmarks. The cybersecurity benchmark is the one area where Opus 4.8 outperforms. Note that Fugu-Cyber, Sakana’s separate gated cyber-defense endpoint launched July 21, 2026, reports 72.1% on CTI-REALM, which would flip that comparison.
The Cost and Speed Reality
Fugu Ultra outperforms Opus 4.8 on most benchmarks, but it costs more time and money per query.
A real-world head-to-head test on a Crossy Road clone showed Fugu Ultra finishing in 22 minutes at $7.32. Opus 4.8 took 79 minutes and $37.85. Fugu was faster and cheaper on that specific task, but the user preferred Opus 4.8’s output quality. That is the practical tension: the benchmark edge does not always translate into a better deliverable for every task type.
Who Should Choose Fugu Over Opus 4.8?
Choose Fugu Ultra if:
- Your tasks are complex, multi-step, and benefit from agent verification
- You want to reduce reliance on a single Anthropic model
- Code review depth matters more to you than raw response speed
Choose Opus 4.8 if:
- You need full query-level transparency and auditable model routing
- Your workload is latency-sensitive
- You want to stay within the Anthropic ecosystem with simpler billing
Sakana Fugu vs Claude Opus 5
Claude Opus 5 launched on July 24, 2026, the same day Sakana released Fugu Ultra v1.1. It is the most directly relevant Anthropic model to compare against Fugu right now, because it is publicly available, aggressively priced, and positioned squarely at the same “near-frontier at reasonable cost” audience Fugu is targeting.
What Opus 5 Is
Opus 5 is not a minor update. Anthropic describes it as a step-change over Opus 4.8, and the benchmark numbers back that up.
Key specs:
- Price: $5 input / $25 output per 1M tokens (same as Opus 4.8, half the cost of Fable 5)
- Fast mode: $10 / $50 per 1M tokens, roughly 2.5x faster
- Context window: 1M tokens, 128K max output
- Knowledge cutoff: May 2026, the most current of any Claude model
- API model ID:
claude-opus-5 - Effort toggle: Low, medium, or high per request, letting you trade cost against depth
It is now the default model on Claude Max and the strongest model available on Claude Pro.
Benchmark Comparison: Fugu Ultra vs Claude Opus 5
| Benchmark | Fugu Ultra | Claude Opus 5 | Notes |
|---|---|---|---|
| SWE-Bench Pro | 73.7 | n/a | No published Opus 5 score yet |
| Frontier-Bench v0.1 | n/a | 43.3% | More than doubles Opus 4.8’s score |
| ARC-AGI-3 | n/a | 30.2% | 3x the next-best model at launch |
| CursorBench 3.2 | n/a | Within 0.5% of Fable 5 | At half Fable 5’s cost per task |
| OSWorld 2.0 | n/a | Beats Fable 5’s best result | At roughly one-third of Fable 5’s cost |
| GPQA-Diamond | 95.5 | n/a | No published Opus 5 score yet |
| Humanity’s Last Exam | 50.0 | n/a | No published Opus 5 score yet |
Direct benchmark overlap between Fugu Ultra and Opus 5 is currently zero. They launched on the same day using completely different benchmark suites.
Sakana tested Fugu Ultra on SWE-Bench Pro, LiveCodeBench, GPQA-Diamond, TerminalBench 2.1, and Humanity’s Last Exam. Anthropic tested Opus 5 on Frontier-Bench v0.1, ARC-AGI-3, CursorBench 3.2, OSWorld 2.0, and ARC-AGI-2. None of those benchmarks appear on both cards.
Filling a comparison table with numbers across those two suites would mean mixing incompatible test environments, scaffolds, and dates. That produces a misleading number, not a useful one. Until an independent lab runs both models on the same benchmark under the same conditions, any direct score comparison between Fugu Ultra and Opus 5 is speculation. What you can compare right now is pricing, architecture, and the use cases each model is explicitly designed for.
What is clear: Opus 5 is a significant leap over Opus 4.8 on agentic and autonomous tasks. On benchmarks where Fugu Ultra led Opus 4.8 comfortably, the gap against Opus 5 will be narrower, and possibly reversed on some dimensions.
The Pricing Angle
This is where the comparison gets interesting for business decision-makers.
| Fugu Ultra (pay-as-you-go) | Claude Opus 5 (standard) | Claude Opus 5 (fast mode) | |
|---|---|---|---|
| Input per 1M tokens | $5 | $5 | $10 |
| Output per 1M tokens | $30 | $25 | $50 |
| Context above 272K | $10 / $45 | Same rate | Same rate |
Fugu Ultra’s output rate ($30/1M) is 20% higher than Opus 5’s standard rate ($25/1M). For output-heavy workloads, Opus 5 is actually cheaper than Fugu Ultra on a per-token basis, while also offering faster response times and full routing transparency.
That is a meaningful shift from the Opus 4.8 comparison. Fugu Ultra no longer has a clear cost advantage over its most direct Anthropic competitor.
Who Should Choose Fugu Over Opus 5?
Choose Fugu Ultra if:
- Your tasks genuinely benefit from multi-agent verification across different model providers
- Vendor diversification matters to your organization, not just output quality
- You want a flat-rate subscription rather than token-based billing
Choose Opus 5 if:
- You want near-frontier Anthropic intelligence at the lowest per-token cost
- Full routing transparency and a single auditable model matter to your team
- You need fast mode (2.5x speed at doubled rate) for latency-sensitive workflows
- You are already in the Anthropic ecosystem and want the simplest upgrade path from Opus 4.8
Sakana Fugu vs GPT-5.5
GPT-5.5 is OpenAI’s current production flagship. It is one of the models in Fugu’s agent pool, which makes this comparison structurally interesting: Fugu uses GPT-5.5 as one of its workers while simultaneously competing with it as an overall product.
Benchmark Comparison: Fugu Ultra vs GPT-5.5
| Benchmark | Fugu Ultra | GPT-5.5 | Winner |
|---|---|---|---|
| SWE-Bench Pro | 73.7 | 58.6 | Fugu Ultra |
| GPQA-Diamond | 95.5 | 93.6 | Fugu Ultra |
| LiveCodeBench | 93.2 | n/a | Fugu Ultra |
| TerminalBench 2.1 | 82.1 | n/a | Fugu Ultra |
| MRCRv2 (long-context recall) | 93.6 | 94.8 | GPT-5.5 |
The MRCRv2 result is worth calling out. GPT-5.5 is the only model to beat Fugu Ultra in Sakana’s own benchmark table. If your work is heavily dependent on long-context retrieval, GPT-5.5 has the demonstrated edge on Sakana’s own data.
Early real-world testing from users also flagged that Fugu Ultra’s performance on frontend work was “a bit jagged,” with one ThreeJS task reportedly coming back “notably worse than GPT-5.5.” Benchmark leads do not cover every task type.
Who Should Choose Fugu Over GPT-5.5?
Choose Fugu Ultra if:
- Your primary workloads are software engineering, scientific reasoning, or code review
- You want multi-vendor resilience rather than dependence on OpenAI alone
- You are comfortable with higher latency in exchange for deeper analysis
Choose GPT-5.5 if:
- Long-context retrieval is a core workload (MRCRv2 edge holds)
- You need fast, predictable responses for high-frequency API calls
- Frontend or creative coding work is a significant part of your use case
Sakana Fugu vs OpenRouter Fusion
This is the most practically relevant comparison for teams evaluating multi-agent orchestration tools, because Fugu and Fusion are the same product category.
Architecture Difference
| Sakana Fugu | OpenRouter Fusion | |
|---|---|---|
| Decision timing | Upfront: Fugu decides which models run before execution | After the fact: Fusion queries models and synthesizes replies |
| Control flow | Conductor-directed, dynamic, role-assigned | Parallel query, blend-on-return |
| Model pool | Fixed (Fugu) or flexible with opt-out (standard Fugu) | Configurable via OpenRouter |
| Availability | US only at launch. Not available in EU or EEA. | Broader geographic access |
| Pricing model | Flat subscription ($20/$100/$200/month) or pay-as-you-go | Pay-as-you-go only |
| Relative cost | Roughly 4x cheaper than Fusion for the same prompts | Roughly 4x more expensive |
| Response speed | Faster (minutes) | Slower (up to 5 to 10 minutes per answer) |
| API compatibility | OpenAI-compatible | OpenRouter-native |
The Cost Gap Is Decisive for High-Volume Work
The 4x cost difference is not a rounding error. At scale, it determines whether an agentic loop is economically viable.
Fusion runs entirely pay-per-use. Fugu offers flat-rate subscriptions. For teams running thousands of agent calls per day, a predictable monthly cap changes the entire cost model.
On high-volume agent loops and coding tasks on a budget, Fugu wins on price and delivers Fable-5-class output for roughly a quarter of Fusion’s cost.
The speed difference also matters in practice. Fusion can take 5 to 10 minutes to return an answer. Fugu’s response times are significantly faster, though still slower than calling a single frontier model directly.
Who Should Choose Fugu Over OpenRouter Fusion?
Choose Fugu if:
- You want flat-rate subscription pricing rather than pure pay-per-use
- Response speed is important alongside answer quality
- Your team is US-based and does not need EU/EEA access
Choose OpenRouter Fusion if:
- You need geographic flexibility including EU or EEA access
- You want to configure your own model panel rather than use Sakana’s fixed pool
- You prefer to pay only for what you use with no subscription commitment
Full Benchmark Summary
All numbers are vendor-reported or provider-reported as of June 2026. Independent third-party replication is not yet available for Fugu’s scores.
| Benchmark | Fugu Ultra | Fable 5 | Opus 5 | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|
| SWE-Bench Pro | 73.7 | 86.0 | n/a | 69.2 | 58.6 | 54.2 |
| LiveCodeBench | 93.2 | n/a | n/a | n/a | n/a | 88.5 |
| GPQA-Diamond | 95.5 | n/a | n/a | 92.0 | 93.6 | 94.3 |
| TerminalBench 2.1 | 82.1 | 80.4 | n/a | n/a | n/a | n/a |
| Humanity’s Last Exam | 50.0 | 53.3 | n/a | 49.8 | n/a | n/a |
| MRCRv2 (long-context) | 93.6 | n/a | n/a | n/a | 94.8 | n/a |
| Frontier-Bench v0.1 | n/a | n/a | 43.3% | n/a | n/a | n/a |
| ARC-AGI-3 | n/a | n/a | 30.2% | 1.5% | n/a | n/a |
| OSWorld 2.0 | n/a | n/a | Beats Fable 5 | n/a | n/a | n/a |
Bold = highest published score in that row. n/a = no publicly available score for that benchmark. Opus 5 and Fugu Ultra use different benchmark suites; direct overlap is limited as of July 2026.
The Decision Framework: Which One Is Right for You?
Answer the three questions below. They cut through the comparison noise faster than any benchmark table.
Question 1: Do you need geographic compliance? If your team or your users are in the EU or EEA, Fugu is not available at launch. Stop here and evaluate Fusion or a direct frontier model API.
Question 2: Do you need full routing transparency? If your legal, compliance, or security team requires knowing exactly which model processed each query, Fugu is a black box. Choose a direct frontier model API instead.
Question 3: Is your primary workload software engineering or everything else? If it is software engineering at the hardest level (SWE-Bench Pro-class tasks), Fable 5 leads by a significant margin and is worth pursuing if you have US access. For most other complex reasoning, research, and multi-step coding tasks, Fugu Ultra is competitive with or ahead of the individually available frontier models.
Quick Verdict Table
| If you need… | Best choice |
|---|---|
| Best raw coding performance, US access only | Fable 5 |
| Best publicly available orchestrator on a budget | Sakana Fugu |
| Long-context retrieval as a primary workload | GPT-5.5 |
| Balanced frontier performance, single model, transparent routing | Claude Opus 4.8 |
| Near-frontier Anthropic intelligence at the lowest per-token cost | Claude Opus 5 |
| Multi-model orchestration with EU access | OpenRouter Fusion |
| Most advanced model available, gated access only | Mythos Preview |
Limitations Across All Options
No product in this comparison is without trade-offs. Here is the honest summary:
Sakana Fugu: Higher latency than any single model. Routing is opaque. Not available in the EU. Benchmarks are vendor-reported only. Daily and token limits reported by early users.
Fable 5: US access only due to export controls. Not publicly available to all organizations. No OpenAI-compatible API.
Mythos Preview: Gated access through Project Glasswing. Not available on demand. Benchmarks are sparse and not independently verified.
Claude Opus 4.8: Strong all-round performance but trails Fugu Ultra on most published coding and reasoning benchmarks in Sakana’s evaluation. Largely superseded by Opus 5 at the same price point.
Claude Opus 5: Launched the same day as Fugu Ultra v1.1. Significant step up from Opus 4.8 on agentic tasks. Direct benchmark overlap with Fugu Ultra is limited as of July 2026. Thinking is on by default, which can increase token costs if you migrated from Opus 4.8 without adjusting prompts.
GPT-5.5: Leads on long-context recall (MRCRv2) but trails Fugu Ultra on SWE-Bench Pro, GPQA-Diamond, and TerminalBench. Real-world creative and frontend coding reported as a weak spot by some users.
OpenRouter Fusion: Significantly more expensive than Fugu for the same prompts. Slower response times (up to 10 minutes). Synthesizes after the fact rather than orchestrating upfront.
Choosing an AI Model With a Partner Who Knows All of Them
Reading benchmark comparisons is useful. Knowing which model actually works on your specific tasks, with your specific data, for your specific team, is a different question.
Phos AI Labs is an embedded AI consulting partner for mid-market companies. We are CCA-F certified by Anthropic, members of the Anthropic Claude Partner Network, and one of the first ten firms globally in the OpenAI Select Partner Network. We have direct access to both OpenAI and Anthropic engineering teams, which means we evaluate tools against current capability, not what was published six months ago.
Whether the right answer is Sakana Fugu, Claude Opus 5, or a combination that changes by use case, we will tell you the honest answer before you commit.
- Model evaluation across all options: We test your representative tasks against Fugu Ultra, Claude Opus 5, GPT-5.5, and others before recommending anything
- AI strategy and foundations: We build the decision framework your team needs before any API key is activated
- Team training: We train your team on whichever tools actually fit how they work
- AI Implementation: We stay until the right AI tools are part of how the business runs
Ready to Make the Right Model Decision for Your Business?
400+ engagements. Clients include Zapier, Coca-Cola, Medtronic, Dataiku, and American Express.
If you are evaluating Sakana Fugu, Claude, or any other AI model and want an honest answer about what fits your workload and budget, start with a conversation at Phos AI Labs.
Frequently Asked Questions
Is Sakana Fugu better than Fable 5?
On most benchmarks Sakana published, Fugu Ultra performs comparably to Fable 5. The one clear exception is SWE-Bench Pro, where Fable 5 scores 86.0 and Fugu Ultra scores 73.7. Fable 5 is not in Fugu’s agent pool and the two have not been tested head-to-head. Sakana’s own language is “shoulder-to-shoulder,” not “beats.”
Is Sakana Fugu better than Mythos Preview?
There is no head-to-head data available. Mythos Preview is gated and not publicly accessible. Sakana claims parity on frontier benchmarks, but that cannot be independently verified.
Is Sakana Fugu better than Claude Opus 4.8?
On Sakana’s reported benchmarks, Fugu Ultra leads Opus 4.8 across most coding and reasoning tests. Opus 4.8 holds an edge on cybersecurity benchmarks (CTI-REALM) and offers full routing transparency that Fugu does not. Both are priced in a comparable range.
Is Sakana Fugu better than GPT-5.5?
On most benchmarks, yes, according to Sakana’s data. GPT-5.5 beats Fugu Ultra on MRCRv2 long-context recall (94.8 vs 93.6). Some real-world testing has also flagged Fugu Ultra as weaker on frontend and creative coding tasks compared to GPT-5.5.
How does Sakana Fugu compare to OpenRouter Fusion?
Both are multi-agent orchestrators, but they work differently. Fugu decides which models to run upfront. Fusion queries models in parallel and blends the results afterward. Fugu costs roughly 4x less for the same prompts and is faster. Fusion offers broader geographic access including the EU.
Can I use Sakana Fugu in Europe?
No. Fugu is not available in the EU or EEA at launch, likely due to GDPR compliance requirements. If your team is based in Europe, OpenRouter Fusion or a direct frontier model API is the current alternative.
How does Sakana Fugu compare to Claude Opus 5?
Claude Opus 5 launched July 24, 2026, at $5 input / $25 output per 1M tokens, making it 20% cheaper on output than Fugu Ultra’s pay-as-you-go rate. It offers near-Fable-5 performance, a per-request effort toggle, and full routing transparency. Direct benchmark overlap with Fugu Ultra is limited since both launched simultaneously using different test suites. For teams that want a single auditable model at near-frontier quality, Opus 5 is the stronger choice. For teams that want multi-vendor orchestration and vendor diversification, Fugu Ultra remains the argument.
Does Sakana Fugu replace GPT-5.5 or Opus 4.8?
No. Fugu coordinates GPT-5.5 and Opus 4.8 as part of its agent pool. It does not replace them. If Fugu loses access to those underlying models, its capability changes accordingly.
All benchmark figures are vendor-reported or provider-reported as of June to July 2026. Independent third-party evaluation of Sakana Fugu’s scores is pending. Pricing reflects each provider’s published rates as of July 2026 and is subject to change.
Related articles
- What Is Sakana Fugu? Multi-Agent AI Model Explained
- A 12-Month AI Roadmap for Your $20M Services Company
- Seven Agency AI Workflows That Free Senior Team Time
- Agentic AI: The Business Guide to Autonomous AI Systems
- Agentic AI Capabilities: What These Systems Can Do Today
- Agentic AI: The Complete Business Guide for 2026