Blog

Sakana Fugu vs Other AI Models (2026)

Sakana Fugu vs Fable 5, Mythos, GPT-5.5, Claude Opus 4.8, and OpenRouter Fusion. Benchmarks, pricing, and a clear verdict for each comparison.

Phos Team ·
ai tools AI Strategy

Key Takeaways

  • Sakana Fugu and OpenRouter Fusion are both multi-agent orchestrators. Every other competitor in this article is a single frontier model. These are not the same category of product.
  • Fugu Ultra leads Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on 10 of 11 Sakana-reported benchmarks. The one loss is MRCRv2 long-context recall, where GPT-5.5 edges ahead.
  • Fable 5 and Mythos are not in Fugu’s agent pool due to US export controls. The “parity” claim is Sakana’s, not a head-to-head result.
  • Fugu costs roughly 4x less than OpenRouter Fusion for the same prompts, and also offers flat-rate subscription plans Fusion does not.
  • Fugu is not available in the EU or EEA at launch. If your team is based in Europe, this comparison ends here.
  • All benchmark numbers below are vendor-reported. No independent third-party lab has replicated them as of July 2026.

Before You Compare: Get the Category Right

Most confusion in “Fugu vs X” comparisons comes from treating all five options as the same type of product. They are not.

ProductTypeWhat It Actually Is
Sakana FuguMulti-agent orchestratorCoordinates Opus 4.8, GPT-5.5, Gemini and others behind one API
OpenRouter FusionMulti-agent synthesizerQueries multiple models and blends their replies after the fact
Claude Fable 5Single frontier LLMAnthropic’s high-end long-horizon model (export-controlled)
Claude Mythos PreviewSingle frontier LLMAnthropic’s most capable model (export-controlled, gated)
Claude Opus 4.8Single frontier LLMAnthropic’s previous high-performance model (still available)
Claude Opus 5Single frontier LLMAnthropic’s current near-frontier model, launched July 24, 2026
GPT-5.5Single frontier LLMOpenAI’s current production flagship

Fugu does not replace Fable 5, Mythos, Opus, or GPT-5.5. It orchestrates some of them. Choosing Fugu means choosing a coordination layer, not a better underlying model.

Fugu vs Fusion: Two Different Control Flows

Fugu and Fusion look similar on the surface. Both accept one request and return one answer using multiple models. But the internal logic runs in opposite directions.

  • Fugu decides upfront which models to call and in what order. It is a conductor: it assigns tasks before execution.
  • Fusion synthesizes after the fact. It queries multiple models in parallel and blends the outputs once they arrive.

One Hacker News commenter framed it well: ask GPT to derive the math, ask Opus to check it for security issues, ask Gemini to resolve the disagreement. That is Fugu’s model. Fusion is closer to a voting booth. The distinction matters when you are choosing between them for structured, multi-step tasks.


Sakana Fugu vs Fable 5

This is the comparison that drove the most interest at launch, and it requires the most context to read correctly.

What Sakana Actually Claims

Sakana does not claim Fugu beats Fable 5. The official language is “shoulder-to-shoulder.” That is a deliberate choice of words.

Fable 5 is not in Fugu’s agent pool. It was pulled from non-US access due to US export controls ten days before Fugu launched. The “parity” claim is based on Sakana’s benchmarks comparing Fugu against Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, then citing Fable 5’s publicly reported scores from prior third-party evaluations alongside those numbers.

That is not a head-to-head. It is a side-by-side on separate test runs.

Benchmark Comparison: Fugu Ultra vs Fable 5

BenchmarkFugu UltraFable 5Notes
SWE-Bench Pro73.786.0Fable 5 wins clearly. 12+ point gap.
TerminalBench 2.182.180.4Fugu Ultra leads
LiveCodeBench93.2n/aFable 5 not independently published here
GPQA-Diamond95.5n/aFable 5 not independently published here
Humanity’s Last Exam50.053.3Fable 5 leads

The honest read: Fugu Ultra holds its own on most coding and reasoning benchmarks. On SWE-Bench Pro, the hardest real-world software engineering benchmark in this comparison, Fable 5 leads by nearly 13 points. That gap is real and not trivial.

Who Should Choose Fugu Over Fable 5?

Choose Fugu if:

  • You are based outside the US and cannot access Fable 5 due to export controls
  • You want vendor diversification rather than dependence on one Anthropic model
  • Your primary use cases are code review, multi-step research, or reasoning tasks rather than complex software engineering

Choose Fable 5 if:

  • You are a US-based team with access and your workload is heavily weighted toward software engineering
  • You need a single model you can audit, with transparent query routing
  • SWE-Bench Pro scores are directly relevant to your production use case

Sakana Fugu vs Mythos Preview

Mythos Preview is Anthropic’s most capable model. It is gated, not publicly available, and restricted to a small set of trusted organizations through Anthropic’s Project Glasswing.

What the Comparison Actually Means

Mythos is in the same position as Fable 5 in this comparison: not in Fugu’s pool, not independently benchmarked against Fugu directly, and not accessible to most organizations.

Sakana claims Fugu Ultra performs “shoulder-to-shoulder” with Mythos Preview on frontier benchmarks. That claim cannot be verified through a head-to-head test because Mythos is not publicly accessible.

Benchmark Comparison: Fugu Ultra vs Mythos Preview

BenchmarkFugu UltraMythos PreviewNotes
SWE-Bench Pro73.7Not publishedCannot compare directly
Humanity’s Last Exam50.0Not publishedCannot compare directly
GPQA-Diamond95.5Not publishedCannot compare directly

There is no meaningful benchmark comparison available to verify or refute Sakana’s claim about Mythos parity.

Who Should Choose Fugu Over Mythos?

This is not a real choice for most organizations. Mythos Preview is gated and not available on demand.

If you do not have Mythos access, Fugu Ultra is a credible alternative for complex research and reasoning tasks based on what Sakana has published.

If you do have Mythos access, run your own evaluation on 20 to 50 representative tasks before making any routing decision based on marketing claims from either side.


Sakana Fugu vs Claude Opus 4.8

This is the most apples-to-apples comparison in the article. Opus 4.8 is publicly available, independently benchmarked, and it is one of the models inside Fugu’s own agent pool.

Benchmark Comparison: Fugu Ultra vs Claude Opus 4.8

BenchmarkFugu UltraOpus 4.8Winner
SWE-Bench Pro73.769.2Fugu Ultra
LiveCodeBench93.2n/aFugu Ultra (no Opus score published)
GPQA-Diamond95.592.0Fugu Ultra
TerminalBench 2.182.1n/aFugu Ultra (no Opus score published)
Humanity’s Last Exam50.049.8Fugu Ultra (narrow margin)
CTI-REALM (cybersecurity)n/a69.6Opus 4.8 edges ahead

Fugu Ultra leads Opus 4.8 across most published benchmarks. The cybersecurity benchmark is the one area where Opus 4.8 outperforms. Note that Fugu-Cyber, Sakana’s separate gated cyber-defense endpoint launched July 21, 2026, reports 72.1% on CTI-REALM, which would flip that comparison.

The Cost and Speed Reality

Fugu Ultra outperforms Opus 4.8 on most benchmarks, but it costs more time and money per query.

A real-world head-to-head test on a Crossy Road clone showed Fugu Ultra finishing in 22 minutes at $7.32. Opus 4.8 took 79 minutes and $37.85. Fugu was faster and cheaper on that specific task, but the user preferred Opus 4.8’s output quality. That is the practical tension: the benchmark edge does not always translate into a better deliverable for every task type.

Who Should Choose Fugu Over Opus 4.8?

Choose Fugu Ultra if:

  • Your tasks are complex, multi-step, and benefit from agent verification
  • You want to reduce reliance on a single Anthropic model
  • Code review depth matters more to you than raw response speed

Choose Opus 4.8 if:

  • You need full query-level transparency and auditable model routing
  • Your workload is latency-sensitive
  • You want to stay within the Anthropic ecosystem with simpler billing

Sakana Fugu vs Claude Opus 5

Claude Opus 5 launched on July 24, 2026, the same day Sakana released Fugu Ultra v1.1. It is the most directly relevant Anthropic model to compare against Fugu right now, because it is publicly available, aggressively priced, and positioned squarely at the same “near-frontier at reasonable cost” audience Fugu is targeting.

What Opus 5 Is

Opus 5 is not a minor update. Anthropic describes it as a step-change over Opus 4.8, and the benchmark numbers back that up.

Key specs:

  • Price: $5 input / $25 output per 1M tokens (same as Opus 4.8, half the cost of Fable 5)
  • Fast mode: $10 / $50 per 1M tokens, roughly 2.5x faster
  • Context window: 1M tokens, 128K max output
  • Knowledge cutoff: May 2026, the most current of any Claude model
  • API model ID: claude-opus-5
  • Effort toggle: Low, medium, or high per request, letting you trade cost against depth

It is now the default model on Claude Max and the strongest model available on Claude Pro.

Benchmark Comparison: Fugu Ultra vs Claude Opus 5

BenchmarkFugu UltraClaude Opus 5Notes
SWE-Bench Pro73.7n/aNo published Opus 5 score yet
Frontier-Bench v0.1n/a43.3%More than doubles Opus 4.8’s score
ARC-AGI-3n/a30.2%3x the next-best model at launch
CursorBench 3.2n/aWithin 0.5% of Fable 5At half Fable 5’s cost per task
OSWorld 2.0n/aBeats Fable 5’s best resultAt roughly one-third of Fable 5’s cost
GPQA-Diamond95.5n/aNo published Opus 5 score yet
Humanity’s Last Exam50.0n/aNo published Opus 5 score yet

Direct benchmark overlap between Fugu Ultra and Opus 5 is currently zero. They launched on the same day using completely different benchmark suites.

Sakana tested Fugu Ultra on SWE-Bench Pro, LiveCodeBench, GPQA-Diamond, TerminalBench 2.1, and Humanity’s Last Exam. Anthropic tested Opus 5 on Frontier-Bench v0.1, ARC-AGI-3, CursorBench 3.2, OSWorld 2.0, and ARC-AGI-2. None of those benchmarks appear on both cards.

Filling a comparison table with numbers across those two suites would mean mixing incompatible test environments, scaffolds, and dates. That produces a misleading number, not a useful one. Until an independent lab runs both models on the same benchmark under the same conditions, any direct score comparison between Fugu Ultra and Opus 5 is speculation. What you can compare right now is pricing, architecture, and the use cases each model is explicitly designed for.

What is clear: Opus 5 is a significant leap over Opus 4.8 on agentic and autonomous tasks. On benchmarks where Fugu Ultra led Opus 4.8 comfortably, the gap against Opus 5 will be narrower, and possibly reversed on some dimensions.

The Pricing Angle

This is where the comparison gets interesting for business decision-makers.

Fugu Ultra (pay-as-you-go)Claude Opus 5 (standard)Claude Opus 5 (fast mode)
Input per 1M tokens$5$5$10
Output per 1M tokens$30$25$50
Context above 272K$10 / $45Same rateSame rate

Fugu Ultra’s output rate ($30/1M) is 20% higher than Opus 5’s standard rate ($25/1M). For output-heavy workloads, Opus 5 is actually cheaper than Fugu Ultra on a per-token basis, while also offering faster response times and full routing transparency.

That is a meaningful shift from the Opus 4.8 comparison. Fugu Ultra no longer has a clear cost advantage over its most direct Anthropic competitor.

Who Should Choose Fugu Over Opus 5?

Choose Fugu Ultra if:

  • Your tasks genuinely benefit from multi-agent verification across different model providers
  • Vendor diversification matters to your organization, not just output quality
  • You want a flat-rate subscription rather than token-based billing

Choose Opus 5 if:

  • You want near-frontier Anthropic intelligence at the lowest per-token cost
  • Full routing transparency and a single auditable model matter to your team
  • You need fast mode (2.5x speed at doubled rate) for latency-sensitive workflows
  • You are already in the Anthropic ecosystem and want the simplest upgrade path from Opus 4.8

Sakana Fugu vs GPT-5.5

GPT-5.5 is OpenAI’s current production flagship. It is one of the models in Fugu’s agent pool, which makes this comparison structurally interesting: Fugu uses GPT-5.5 as one of its workers while simultaneously competing with it as an overall product.

Benchmark Comparison: Fugu Ultra vs GPT-5.5

BenchmarkFugu UltraGPT-5.5Winner
SWE-Bench Pro73.758.6Fugu Ultra
GPQA-Diamond95.593.6Fugu Ultra
LiveCodeBench93.2n/aFugu Ultra
TerminalBench 2.182.1n/aFugu Ultra
MRCRv2 (long-context recall)93.694.8GPT-5.5

The MRCRv2 result is worth calling out. GPT-5.5 is the only model to beat Fugu Ultra in Sakana’s own benchmark table. If your work is heavily dependent on long-context retrieval, GPT-5.5 has the demonstrated edge on Sakana’s own data.

Early real-world testing from users also flagged that Fugu Ultra’s performance on frontend work was “a bit jagged,” with one ThreeJS task reportedly coming back “notably worse than GPT-5.5.” Benchmark leads do not cover every task type.

Who Should Choose Fugu Over GPT-5.5?

Choose Fugu Ultra if:

  • Your primary workloads are software engineering, scientific reasoning, or code review
  • You want multi-vendor resilience rather than dependence on OpenAI alone
  • You are comfortable with higher latency in exchange for deeper analysis

Choose GPT-5.5 if:

  • Long-context retrieval is a core workload (MRCRv2 edge holds)
  • You need fast, predictable responses for high-frequency API calls
  • Frontend or creative coding work is a significant part of your use case

Sakana Fugu vs OpenRouter Fusion

This is the most practically relevant comparison for teams evaluating multi-agent orchestration tools, because Fugu and Fusion are the same product category.

Architecture Difference

Sakana FuguOpenRouter Fusion
Decision timingUpfront: Fugu decides which models run before executionAfter the fact: Fusion queries models and synthesizes replies
Control flowConductor-directed, dynamic, role-assignedParallel query, blend-on-return
Model poolFixed (Fugu) or flexible with opt-out (standard Fugu)Configurable via OpenRouter
AvailabilityUS only at launch. Not available in EU or EEA.Broader geographic access
Pricing modelFlat subscription ($20/$100/$200/month) or pay-as-you-goPay-as-you-go only
Relative costRoughly 4x cheaper than Fusion for the same promptsRoughly 4x more expensive
Response speedFaster (minutes)Slower (up to 5 to 10 minutes per answer)
API compatibilityOpenAI-compatibleOpenRouter-native

The Cost Gap Is Decisive for High-Volume Work

The 4x cost difference is not a rounding error. At scale, it determines whether an agentic loop is economically viable.

Fusion runs entirely pay-per-use. Fugu offers flat-rate subscriptions. For teams running thousands of agent calls per day, a predictable monthly cap changes the entire cost model.

On high-volume agent loops and coding tasks on a budget, Fugu wins on price and delivers Fable-5-class output for roughly a quarter of Fusion’s cost.

The speed difference also matters in practice. Fusion can take 5 to 10 minutes to return an answer. Fugu’s response times are significantly faster, though still slower than calling a single frontier model directly.

Who Should Choose Fugu Over OpenRouter Fusion?

Choose Fugu if:

  • You want flat-rate subscription pricing rather than pure pay-per-use
  • Response speed is important alongside answer quality
  • Your team is US-based and does not need EU/EEA access

Choose OpenRouter Fusion if:

  • You need geographic flexibility including EU or EEA access
  • You want to configure your own model panel rather than use Sakana’s fixed pool
  • You prefer to pay only for what you use with no subscription commitment

Full Benchmark Summary

All numbers are vendor-reported or provider-reported as of June 2026. Independent third-party replication is not yet available for Fugu’s scores.

BenchmarkFugu UltraFable 5Opus 5Opus 4.8GPT-5.5Gemini 3.1 Pro
SWE-Bench Pro73.786.0n/a69.258.654.2
LiveCodeBench93.2n/an/an/an/a88.5
GPQA-Diamond95.5n/an/a92.093.694.3
TerminalBench 2.182.180.4n/an/an/an/a
Humanity’s Last Exam50.053.3n/a49.8n/an/a
MRCRv2 (long-context)93.6n/an/an/a94.8n/a
Frontier-Bench v0.1n/an/a43.3%n/an/an/a
ARC-AGI-3n/an/a30.2%1.5%n/an/a
OSWorld 2.0n/an/aBeats Fable 5n/an/an/a

Bold = highest published score in that row. n/a = no publicly available score for that benchmark. Opus 5 and Fugu Ultra use different benchmark suites; direct overlap is limited as of July 2026.


The Decision Framework: Which One Is Right for You?

Answer the three questions below. They cut through the comparison noise faster than any benchmark table.

Question 1: Do you need geographic compliance? If your team or your users are in the EU or EEA, Fugu is not available at launch. Stop here and evaluate Fusion or a direct frontier model API.

Question 2: Do you need full routing transparency? If your legal, compliance, or security team requires knowing exactly which model processed each query, Fugu is a black box. Choose a direct frontier model API instead.

Question 3: Is your primary workload software engineering or everything else? If it is software engineering at the hardest level (SWE-Bench Pro-class tasks), Fable 5 leads by a significant margin and is worth pursuing if you have US access. For most other complex reasoning, research, and multi-step coding tasks, Fugu Ultra is competitive with or ahead of the individually available frontier models.

Quick Verdict Table

If you need…Best choice
Best raw coding performance, US access onlyFable 5
Best publicly available orchestrator on a budgetSakana Fugu
Long-context retrieval as a primary workloadGPT-5.5
Balanced frontier performance, single model, transparent routingClaude Opus 4.8
Near-frontier Anthropic intelligence at the lowest per-token costClaude Opus 5
Multi-model orchestration with EU accessOpenRouter Fusion
Most advanced model available, gated access onlyMythos Preview

Limitations Across All Options

No product in this comparison is without trade-offs. Here is the honest summary:

Sakana Fugu: Higher latency than any single model. Routing is opaque. Not available in the EU. Benchmarks are vendor-reported only. Daily and token limits reported by early users.

Fable 5: US access only due to export controls. Not publicly available to all organizations. No OpenAI-compatible API.

Mythos Preview: Gated access through Project Glasswing. Not available on demand. Benchmarks are sparse and not independently verified.

Claude Opus 4.8: Strong all-round performance but trails Fugu Ultra on most published coding and reasoning benchmarks in Sakana’s evaluation. Largely superseded by Opus 5 at the same price point.

Claude Opus 5: Launched the same day as Fugu Ultra v1.1. Significant step up from Opus 4.8 on agentic tasks. Direct benchmark overlap with Fugu Ultra is limited as of July 2026. Thinking is on by default, which can increase token costs if you migrated from Opus 4.8 without adjusting prompts.

GPT-5.5: Leads on long-context recall (MRCRv2) but trails Fugu Ultra on SWE-Bench Pro, GPQA-Diamond, and TerminalBench. Real-world creative and frontend coding reported as a weak spot by some users.

OpenRouter Fusion: Significantly more expensive than Fugu for the same prompts. Slower response times (up to 10 minutes). Synthesizes after the fact rather than orchestrating upfront.


Choosing an AI Model With a Partner Who Knows All of Them

Reading benchmark comparisons is useful. Knowing which model actually works on your specific tasks, with your specific data, for your specific team, is a different question.

Phos AI Labs is an embedded AI consulting partner for mid-market companies. We are CCA-F certified by Anthropic, members of the Anthropic Claude Partner Network, and one of the first ten firms globally in the OpenAI Select Partner Network. We have direct access to both OpenAI and Anthropic engineering teams, which means we evaluate tools against current capability, not what was published six months ago.

Whether the right answer is Sakana Fugu, Claude Opus 5, or a combination that changes by use case, we will tell you the honest answer before you commit.

  • Model evaluation across all options: We test your representative tasks against Fugu Ultra, Claude Opus 5, GPT-5.5, and others before recommending anything
  • AI strategy and foundations: We build the decision framework your team needs before any API key is activated
  • Team training: We train your team on whichever tools actually fit how they work
  • AI Implementation: We stay until the right AI tools are part of how the business runs

Ready to Make the Right Model Decision for Your Business?

400+ engagements. Clients include Zapier, Coca-Cola, Medtronic, Dataiku, and American Express.

If you are evaluating Sakana Fugu, Claude, or any other AI model and want an honest answer about what fits your workload and budget, start with a conversation at Phos AI Labs.


Frequently Asked Questions

Is Sakana Fugu better than Fable 5?

On most benchmarks Sakana published, Fugu Ultra performs comparably to Fable 5. The one clear exception is SWE-Bench Pro, where Fable 5 scores 86.0 and Fugu Ultra scores 73.7. Fable 5 is not in Fugu’s agent pool and the two have not been tested head-to-head. Sakana’s own language is “shoulder-to-shoulder,” not “beats.”

Is Sakana Fugu better than Mythos Preview?

There is no head-to-head data available. Mythos Preview is gated and not publicly accessible. Sakana claims parity on frontier benchmarks, but that cannot be independently verified.

Is Sakana Fugu better than Claude Opus 4.8?

On Sakana’s reported benchmarks, Fugu Ultra leads Opus 4.8 across most coding and reasoning tests. Opus 4.8 holds an edge on cybersecurity benchmarks (CTI-REALM) and offers full routing transparency that Fugu does not. Both are priced in a comparable range.

Is Sakana Fugu better than GPT-5.5?

On most benchmarks, yes, according to Sakana’s data. GPT-5.5 beats Fugu Ultra on MRCRv2 long-context recall (94.8 vs 93.6). Some real-world testing has also flagged Fugu Ultra as weaker on frontend and creative coding tasks compared to GPT-5.5.

How does Sakana Fugu compare to OpenRouter Fusion?

Both are multi-agent orchestrators, but they work differently. Fugu decides which models to run upfront. Fusion queries models in parallel and blends the results afterward. Fugu costs roughly 4x less for the same prompts and is faster. Fusion offers broader geographic access including the EU.

Can I use Sakana Fugu in Europe?

No. Fugu is not available in the EU or EEA at launch, likely due to GDPR compliance requirements. If your team is based in Europe, OpenRouter Fusion or a direct frontier model API is the current alternative.

How does Sakana Fugu compare to Claude Opus 5?

Claude Opus 5 launched July 24, 2026, at $5 input / $25 output per 1M tokens, making it 20% cheaper on output than Fugu Ultra’s pay-as-you-go rate. It offers near-Fable-5 performance, a per-request effort toggle, and full routing transparency. Direct benchmark overlap with Fugu Ultra is limited since both launched simultaneously using different test suites. For teams that want a single auditable model at near-frontier quality, Opus 5 is the stronger choice. For teams that want multi-vendor orchestration and vendor diversification, Fugu Ultra remains the argument.

Does Sakana Fugu replace GPT-5.5 or Opus 4.8?

No. Fugu coordinates GPT-5.5 and Opus 4.8 as part of its agent pool. It does not replace them. If Fugu loses access to those underlying models, its capability changes accordingly.


All benchmark figures are vendor-reported or provider-reported as of June to July 2026. Independent third-party evaluation of Sakana Fugu’s scores is pending. Pricing reflects each provider’s published rates as of July 2026 and is subject to change.

Related articles

The fastest way to know whether we're the right fit, is a conversation.

STEP 1/2 · ABOUT YOU