Skip to content

Claude Opus 5 Is Live — Anthropic's New Frontier Flagship

July 25, 2026

Claude Opus 5 is in the model picker today. It's Anthropic's most capable frontier thinker — near-Fable-5 intelligence one band below Fable's, with adaptive thinking, a 1M-token context window, and 128K max output. It takes over from Opus 4.8 as the recommended pick at the top of the SOTA tier, in the same token band. For harness scores and evidence packages across the model lineup, see Can it Run Alfrada?.

What you can do

  • Pick Claude Opus 5 from the SOTA section of the model selector, ready for any session. It's the new recommended frontier pick on cloud and hybrid setups.
  • Long-horizon agentic coding — multi-step refactors and multi-day projects that run across many tool calls and many turns, where the model needs to remember what it decided three steps ago.
  • Enterprise workflows — strategic planning, financial analysis, and deep research where you want the rigour of a frontier model without stepping up to Fable pricing.
  • High-fidelity multimodal reasoning — screenshots, charts, and dense documents, held together with the full 1M-token context.
  • Adaptive thinking with visible summaries — it decides how deeply to reason per turn and shows you a readable summary of its thinking, running at high effort by default.
  • Swarm-capable — put it in charge of a Swarm that delegates the legwork to faster, cheaper models, or drop it in as the specialist on the hardest sub-task.
  • Auto-routable — tagged SOTA tier with the highest quality score of any Opus, so Auto reaches for it when a turn genuinely needs frontier-tier reasoning.

How it scored

We ran Opus 5 through both benchmark cases in the production harness on the same prompts as the July sweep, and put the artifacts in front of the same judge panel.

  • First on the calibrated ranking. Opus 5 leads the family-balanced judge z-score the study ranks by (+1.00), and comes second on the raw 0–10 composite at 8.77 against GPT-5.6 Sol's 8.84. That gap is inside run-to-run noise, so read the two as tied on score. Its own self-verdict is published but excluded from the ranking, so it's scored by the same seven judge signals as every other model.
  • The judges liked its arithmetic. Every load-bearing figure in the case 1 financial model reproduced exactly on recomputation, at full precision.
  • They were harder on its sourcing, and that's the counterweight. Several staged screenshots didn't support the claims citing them, and the case 2 simulation shipped no source code, so its printed seed can't be independently replayed. It carries more catalogued load-bearing defects than any other model on the board — Sol, right behind it on score, carries none.
  • It was willing to say no. On case 1 it reached a conditional no-go, one of only two models out of eleven to break from the consensus conditional GO, on arithmetic its judges independently verified.
  • It's slow. 58 minutes on case 1 and 33 on case 2 — the longest case 1 on the board. Budget accordingly on latency-sensitive work.

Where this shows up

  • You've been running Opus 4.8 on hard agentic work — Opus 5 is the direct upgrade in the same band, at the same price, with a higher quality rating.
  • You want Fable-class judgement on a multi-day project but the Apex multiplier doesn't fit the budget. Opus 5 sits one band down at a lower token rate.
  • You're running a research or coding Swarm — Opus 5 leads, plans, and synthesizes while cheaper models gather and draft.

Try it

  • "Take over this half-finished migration: read the repo, figure out what's left, and finish it — flag anything you had to decide on the way."
  • "Work through the financials on this acquisition target and tell me where the deal breaks."
  • "Run a research swarm: you lead and synthesize, workers gather sources and draft sections."

Heads up

  • SOTA band — 12× tokens. Same band as Opus 4.8; billed at the SOTA multiplier on your plan quota.
  • Opus 5 replaces Opus 4.8. Opus 4.8 is retired from the picker, Auto, and Swarm. Existing conversations pinned to Opus 4.8 roll forward to Opus 5 automatically — no action needed, and pricing is unchanged.
  • EU-only geofencing supported. Like Opus 4.8, Opus 5 routes through geo-prefixed Bedrock profiles, so requests can be pinned to EU-hosted infrastructure when EU Data Residency is on.
  • Check the benchmark before you commit a workflow. Can it Run Alfrada? has the harness scores, file-by-file judge verdicts, evidence packages, cost, and speed comparisons across the model lineup.

Built for Alfrada OS.