DeepSeek logo

DeepSeek

DeepSeek V4-Pro

FrontierOpen-weightAvailable in FlowHunt

DeepSeek V4-Pro is a top-tier frontier AI model, developed by DeepSeek and released in April 2026. As an open-weight model, its trained weights are publicly available — you can self-host it, fine-tune it on proprietary data, or run it on-premise without API dependencies. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 9 benchmarks. Standout scores: 93.5% on LiveCodeBench (#1 of 5); 87.5% on MMLU-Pro (#2 of 8); 46.0% on ARC-AGI-2 (#3 of 3).

DeepSeek V4-Pro was released on April 24, 2026 as the flagship of DeepSeek’s fourth-generation V series. It achieved 93.5% on LiveCodeBench at launch — the highest competitive programming score of any measured model — and posted the highest self-reported SWE-bench Verified score outside of Claude Fable 5 at 80.6%. The release created significant discussion when CAISI independently measured V4-Pro’s SWE-bench score at 74.0% rather than the 80.6% DeepSeek had reported — a 6.6-point gap that DeepSeek attributed to scaffolding and token-budget differences in evaluation setup. As an open-weight model with a 1M context window and 49B/1.6T MoE architecture, V4-Pro became the leading self-hostable option for teams prioritising coding performance with full data control.

Note: SWE-bench discrepancy: SR 80.6% vs CAISI 74%. LiveCodeBench 93.5% highest reported.

Released
April 2026
Parameters
49B/1.6T MoE
Context
1M tokens
Weights
Open
Tier
Frontier
Provider
DeepSeek

DeepSeek V4-Pro Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
SWE-bench Verified80.6%6 of 13SRReal GitHub issue resolution (500 verified issues)
SWE-bench Pro55.4%9 of 9SRMulti-language, standardised scaffold
GPQA Diamond90.1%6 of 14SRGraduate-level Google-proof science Q&A
MMLU-Pro87.5%2 of 8SRHard knowledge reasoning, 12K questions
LiveCodeBench93.5%1 of 5SRCompetitive programming, continuously updated
Terminal-Bench67.9%4 of 8SRAgentic Linux terminal task completion
ARC-AGI-246.0%3 of 3CNovel visual reasoning, no memorisation
HLE37.7%4 of 6SRHumanity's Last Exam (50+ STEM disciplines)
AA Intelligence Index52.0 pts6 of 9AAArtificial Analysis composite score (independent)

Detailed Benchmark Scores

SWE-V SR
80.6%
Real GitHub issue resolution (500 verified issues)
Independent: 74.0% (CAISI independent)
SWE-Pro SR
55.4%
Multi-language, standardised scaffold
GPQA ◇ SR
90.1%
Graduate-level Google-proof science Q&A
MMLU-P SR
87.5%
Hard knowledge reasoning, 12K questions
LiveCode SR
93.5%
Competitive programming, continuously updated
Terminal SR
67.9%
Agentic Linux terminal task completion
ARC-2 C
46.0%
Novel visual reasoning, no memorisation
HLE SR
37.7%
Humanity's Last Exam (50+ STEM disciplines)
AA Idx AA
52.0 pts
Artificial Analysis composite score (independent)

DeepSeek V4-Pro vs. Frontier Peers

Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.

BenchmarkDeepSeek V4-ProClaude Opus 4.8Claude Opus 4.7
SWE-V80.6%88.6%87.6%
SWE-Pro55.4%69.2%64.3%
GPQA ◇90.1%93.6%94.2%
MMLU-P87.5%
LiveCode93.5%
Terminal67.9%74.6%66.1%
ARC-246.0%
HLE37.7%
AA Idx52.0 pts61.4 pts57.3 pts

About DeepSeek

DeepSeek Founded 2023 · Hangzhou, China
Website →

DeepSeek is a Chinese AI research lab founded in 2023 as a subsidiary of High-Flyer, one of China's largest quantitative hedge funds. The lab became widely known in early 2025 when the release of DeepSeek V3 and R1 triggered the largest single-day drop in NVIDIA's market capitalisation to that date, as investors reassessed the capital intensity of frontier model development in light of DeepSeek's reported lower training costs. DeepSeek is unusual among frontier labs in its commitment to releasing open-weight models alongside its proprietary research, and in operating an active social-media presence that critiques closed-model pricing. Its V-series models (V3 through V4-Pro) are widely used as self-hostable coding and reasoning models in the West, despite ongoing regulatory and security scrutiny of Chinese-origin AI models in several jurisdictions.

Use Cases

What to Use DeepSeek V4-Pro For

Software Engineering
Scientific & Technical Reasoning
Agentic CLI Automation
Self-Hosted & Private Deployments
Long-Document & Codebase Analysis
Competitive Programming

Strengths & Limitations

Strengths

  • Strong software engineering — 80% SWE-bench Verified
  • Graduate-level science reasoning — 90% GPQA Diamond
  • Competitive programming — 93% LiveCodeBench
  • Composite intelligence — 52 pts AA Intelligence Index (independent)
  • Self-hosted & data-private deployments (open-weight)
  • Very long documents and large codebases (1M context)

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

DeepSeek V4-Flash — Benchmarks, Capabilities, and Use Cases
DeepSeek V4-Flash — Benchmarks, Capabilities, and Use Cases

DeepSeek V4-Flash — Benchmarks, Capabilities, and Use Cases

DeepSeek V4-Flash is a open-weight AI model from DeepSeek (13B/284B MoE), classified as Advanced tier with a 1M tokens context window. It is available directly ...

5 min read
Llama 4 Scout — Benchmarks, Capabilities, and Use Cases
Llama 4 Scout — Benchmarks, Capabilities, and Use Cases

Llama 4 Scout — Benchmarks, Capabilities, and Use Cases

Llama 4 Scout is a open-weight AI model from Meta (17B/109B MoE (16 experts)), classified as Established tier with a 10M tokens context window. It is available ...

4 min read
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases
Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 — Benchmarks, Capabilities, and Use Cases

Claude Opus 4.7 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...

5 min read