Anthropic logo

Anthropic

Claude Opus 4.6

FrontierAvailable in FlowHunt

Claude Opus 4.6 is a top-tier frontier AI model, developed by Anthropic and released in February 2026. It is a proprietary closed-weight model, available through the Anthropic API. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 6 benchmarks. Standout scores: 69.2% on ARC-AGI-2 (#1 of 3); 53.0% on HLE (#1 of 6); 80.8% on SWE-bench Verified (#5 of 13).

Claude Opus 4.6 was released on February 5, 2026 and gained particular attention for its performance on novel visual reasoning tasks. It holds the highest publicly verified score on ARC-AGI-2 at 69.17% — a benchmark of abstract visual pattern recognition independently verified by the ARC Prize Foundation and designed specifically to prevent memorisation. This result placed Opus 4.6 above the approximately 60% human baseline on a test that previous frontier models had struggled to exceed. It also achieved 91.3% on GPQA Diamond and 53.0% on Humanity’s Last Exam, making it the leading model for hard reasoning tasks in the first half of 2026.

Note: SOTA ARC-AGI-2 69.17% (3P ARC Prize). ASL-3 safety classification.

Released
February 2026
Context
1M tokens
Weights
Closed
Tier
Frontier
Provider
Anthropic

Claude Opus 4.6 Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
SWE-bench Verified80.8%5 of 13SRReal GitHub issue resolution (500 verified issues)
SWE-bench Pro57.3%8 of 9SRMulti-language, standardised scaffold
GPQA Diamond91.3%5 of 14SRGraduate-level Google-proof science Q&A
Terminal-Bench65.4%7 of 8SRAgentic Linux terminal task completion
ARC-AGI-269.2%1 of 33PNovel visual reasoning, no memorisation
HLE53.0%1 of 6SRHumanity's Last Exam (50+ STEM disciplines)

Detailed Benchmark Scores

SWE-V SR
80.8%
Real GitHub issue resolution (500 verified issues)
SWE-Pro SR
57.3%
Multi-language, standardised scaffold
GPQA ◇ SR
91.3%
Graduate-level Google-proof science Q&A
Terminal SR
65.4%
Agentic Linux terminal task completion
ARC-2 3P
69.2%
Novel visual reasoning, no memorisation
HLE SR
53.0%
Humanity's Last Exam (50+ STEM disciplines)

Claude Opus 4.6 vs. Frontier Peers

Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.

BenchmarkClaude Opus 4.6GPT-5.5GPT-5.4
SWE-V80.8%88.7%
SWE-Pro57.3%58.6%59.1%
GPQA ◇91.3%93.6%
Terminal65.4%82.7%
ARC-269.2%
HLE53.0%41.4%

About Anthropic

Anthropic Founded 2021 · San Francisco, CA
Website →

Anthropic is an AI safety company founded in 2021 by former OpenAI researchers Dario Amodei and Daniela Amodei, along with a group of colleagues who had led OpenAI's safety and policy teams. The company's stated mission is the responsible development of AI for the long-term benefit of humanity, and its research agenda is defined by that commitment: Constitutional AI (CAI), a technique for training models to be helpful, harmless, and honest through self-critique; mechanistic interpretability, which aims to reverse-engineer what individual neurons and circuits inside large networks actually compute; and scaling science, which tries to predict how model behaviour changes with compute and data. Anthropic publishes this research openly and has become one of the most cited AI safety labs in academic literature. Claude is Anthropic's public model family — named for Claude Shannon, the founder of information theory — and spans four generations (1 through 4.x/Fable 5). Anthropic operates a commercial API and consumer-facing Claude.ai product, and maintains deep cloud partnerships with AWS (Amazon Bedrock) and Google Cloud (Vertex AI).

Use Cases

What to Use Claude Opus 4.6 For

Software Engineering
Scientific & Technical Reasoning
Agentic CLI Automation
Long-Document & Codebase Analysis

Strengths & Limitations

Strengths

  • Strong software engineering — 80% SWE-bench Verified
  • Graduate-level science reasoning — 91% GPQA Diamond
  • Novel visual reasoning — 69% ARC-AGI-2 (independently verified)
  • Very long documents and large codebases (1M context)

Limitations

  • Closed-weight — cannot be self-hosted or fine-tuned on private data

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

Claude Opus 4.8 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.8 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.8 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.8 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...

5 min read
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...

5 min read
Claude Sonnet 4.6 — Benchmarks, Capabilities, and Use Cases
Claude Sonnet 4.6 — Benchmarks, Capabilities, and Use Cases

Claude Sonnet 4.6 — Benchmarks, Capabilities, and Use Cases

Claude Sonnet 4.6 is a proprietary AI model from Anthropic, classified as Advanced tier with a 1M tokens context window. It is available directly in FlowHunt. E...

4 min read