Anthropic logo

Anthropic

Claude Sonnet 4.6

AdvancedAvailable in FlowHunt

Claude Sonnet 4.6 is an advanced AI model suited for demanding production workloads, developed by Anthropic and released in February 2026. It is a proprietary closed-weight model, available through the Anthropic API. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 4 benchmarks. Standout scores: 89.0% on MATH-500 (#2 of 4); 58.3% on ARC-AGI-2 (#2 of 3); 79.6% on SWE-bench Verified (#8 of 13).

Claude Sonnet 4.6 was released on February 17, 2026 as the mid-tier model in the Claude 4 generation, positioned between the Opus series and the Haiku line in both capability and price. The Sonnet tier has historically offered the best price-to-performance ratio in Anthropic’s lineup — capable enough for the vast majority of coding and reasoning tasks at roughly one-fifth of Opus pricing. Sonnet 4.6 scored 58.3% on ARC-AGI-2 (independently verified by the ARC Prize Foundation), giving it the second-highest score globally on that benchmark at release. It supports the same 1M-token context window as the Opus models.

Note: Best-in-class finance agent (63.3%) and MCP-Atlas (61.3%). Adaptive thinking.

Released
February 2026
Context
1M tokens
Weights
Closed
Tier
Advanced
Provider
Anthropic

Claude Sonnet 4.6 Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
SWE-bench Verified79.6%8 of 13SRReal GitHub issue resolution (500 verified issues)
Terminal-Bench59.1%8 of 8SRAgentic Linux terminal task completion
MATH-50089.0%2 of 4SRHendrycks math (500 problems)
ARC-AGI-258.3%2 of 33PNovel visual reasoning, no memorisation

Detailed Benchmark Scores

SWE-V SR
79.6%
Real GitHub issue resolution (500 verified issues)
Terminal SR
59.1%
Agentic Linux terminal task completion
MATH SR
89.0%
Hendrycks math (500 problems)
ARC-2 3P
58.3%
Novel visual reasoning, no memorisation

Claude Sonnet 4.6 vs. Advanced Peers

Head-to-head benchmark comparison with other Advanced-tier models. Higher is better for all metrics.

BenchmarkClaude Sonnet 4.6GPT-4.1DeepSeek V4-Flash
SWE-V79.6%54.6%79.0%
Terminal59.1%
MATH89.0%
ARC-258.3%

About Anthropic

Anthropic Founded 2021 · San Francisco, CA
Website →

Anthropic is an AI safety company founded in 2021 by former OpenAI researchers Dario Amodei and Daniela Amodei, along with a group of colleagues who had led OpenAI's safety and policy teams. The company's stated mission is the responsible development of AI for the long-term benefit of humanity, and its research agenda is defined by that commitment: Constitutional AI (CAI), a technique for training models to be helpful, harmless, and honest through self-critique; mechanistic interpretability, which aims to reverse-engineer what individual neurons and circuits inside large networks actually compute; and scaling science, which tries to predict how model behaviour changes with compute and data. Anthropic publishes this research openly and has become one of the most cited AI safety labs in academic literature. Claude is Anthropic's public model family — named for Claude Shannon, the founder of information theory — and spans four generations (1 through 4.x/Fable 5). Anthropic operates a commercial API and consumer-facing Claude.ai product, and maintains deep cloud partnerships with AWS (Amazon Bedrock) and Google Cloud (Vertex AI).

Use Cases

What to Use Claude Sonnet 4.6 For

Software Engineering
Agentic CLI Automation
Long-Document & Codebase Analysis

Strengths & Limitations

Strengths

  • Strong software engineering — 79% SWE-bench Verified
  • Novel visual reasoning — 58% ARC-AGI-2 (independently verified)
  • Advanced mathematics — 89% MATH-500
  • Very long documents and large codebases (1M context)

Limitations

  • Closed-weight — cannot be self-hosted or fine-tuned on private data

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

Claude Opus 4.6 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.6 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.6 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.6 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. It is available directly in FlowHunt. Exp...

5 min read
Claude Opus 4.8 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.8 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.8 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.8 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...

5 min read
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...

5 min read