Anthropic logo

Anthropic

Claude Opus 4.8

Frontier

Claude Opus 4.8 is a top-tier frontier AI model, developed by Anthropic and released in May 2026. It is a proprietary closed-weight model, available through the Anthropic API. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 5 benchmarks. Standout scores: 61.4 pts on AA Intelligence Index (#1 of 9); 69.2% on SWE-bench Pro (#2 of 9); 74.6% on Terminal-Bench (#2 of 8).

Claude Opus 4.8 was released on May 28, 2026 as a focused point release on top of Opus 4.7, with Anthropic describing it as a ‘modest but tangible upgrade at the same price.’ The release followed a rapid cadence: Opus 4.6, 4.7, and 4.8 all shipped within the same calendar quarter, reflecting Anthropic’s shift toward continuous incremental delivery. Despite modest benchmark improvements, Opus 4.8 took the top position on the Artificial Analysis Intelligence Index at 61.4 points — the highest independent composite score of any model at the time of release — making it the most well-rounded frontier model by independent third-party measurement, even as it was outpaced by Claude Fable 5 on pure coding tasks two weeks later.

Note: Modest but tangible upgrade over 4.7 at same price. #1 AA Intelligence Index at 61.4.

Released
May 2026
Context
1M tokens
Weights
Closed
Tier
Frontier
Provider
Anthropic

Claude Opus 4.8 Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
SWE-bench Verified88.6%3 of 13SRReal GitHub issue resolution (500 verified issues)
SWE-bench Pro69.2%2 of 9SRMulti-language, standardised scaffold
GPQA Diamond93.6%3 of 14SRGraduate-level Google-proof science Q&A
Terminal-Bench74.6%2 of 8SRAgentic Linux terminal task completion
AA Intelligence Index61.4 pts1 of 9AAArtificial Analysis composite score (independent)

Detailed Benchmark Scores

SWE-V SR
88.6%
Real GitHub issue resolution (500 verified issues)
SWE-Pro SR
69.2%
Multi-language, standardised scaffold
GPQA ◇ SR
93.6%
Graduate-level Google-proof science Q&A
Terminal SR
74.6%
Agentic Linux terminal task completion
AA Idx AA
61.4 pts
Artificial Analysis composite score (independent)

Claude Opus 4.8 vs. Frontier Peers

Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.

BenchmarkClaude Opus 4.8GPT-5.5GPT-5.4
SWE-V88.6%88.7%
SWE-Pro69.2%58.6%59.1%
GPQA ◇93.6%93.6%
Terminal74.6%82.7%
AA Idx61.4 pts60.2 pts

About Anthropic

Anthropic Founded 2021 · San Francisco, CA
Website →

Anthropic is an AI safety company founded in 2021 by former OpenAI researchers Dario Amodei and Daniela Amodei, along with a group of colleagues who had led OpenAI's safety and policy teams. The company's stated mission is the responsible development of AI for the long-term benefit of humanity, and its research agenda is defined by that commitment: Constitutional AI (CAI), a technique for training models to be helpful, harmless, and honest through self-critique; mechanistic interpretability, which aims to reverse-engineer what individual neurons and circuits inside large networks actually compute; and scaling science, which tries to predict how model behaviour changes with compute and data. Anthropic publishes this research openly and has become one of the most cited AI safety labs in academic literature. Claude is Anthropic's public model family — named for Claude Shannon, the founder of information theory — and spans four generations (1 through 4.x/Fable 5). Anthropic operates a commercial API and consumer-facing Claude.ai product, and maintains deep cloud partnerships with AWS (Amazon Bedrock) and Google Cloud (Vertex AI).

Use Cases

What to Use Claude Opus 4.8 For

Software Engineering
Scientific & Technical Reasoning
Agentic CLI Automation
Long-Document & Codebase Analysis

Strengths & Limitations

Strengths

  • Strong software engineering — 88% SWE-bench Verified
  • Graduate-level science reasoning — 93% GPQA Diamond
  • Agentic Linux / CLI automation — 74% Terminal-Bench
  • Composite intelligence — 61 pts AA Intelligence Index (independent)
  • Very long documents and large codebases (1M context)

Limitations

  • Closed-weight — cannot be self-hosted or fine-tuned on private data

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...

5 min read
Claude Opus 4.6 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.6 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.6 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.6 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. It is available directly in FlowHunt. Exp...

5 min read
Claude Fable 5 — Benchmarks, Capabilities, and Use Cases
Claude Fable 5 — Benchmarks, Capabilities, and Use Cases

Claude Fable 5 — Benchmarks, Capabilities, and Use Cases

Claude Fable 5 is a proprietary AI model from Anthropic, classified as Super-Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths,...

4 min read