
DeepSeek V4-Flash — Benchmarks, Capabilities, and Use Cases
DeepSeek V4-Flash is a open-weight AI model from DeepSeek (13B/284B MoE), classified as Advanced tier with a 1M tokens context window. It is available directly ...
DeepSeek V4-Pro is a top-tier frontier AI model, developed by DeepSeek and released in April 2026. As an open-weight model, its trained weights are publicly available — you can self-host it, fine-tune it on proprietary data, or run it on-premise without API dependencies. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.
On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 9 benchmarks. Standout scores: 93.5% on LiveCodeBench (#1 of 5); 87.5% on MMLU-Pro (#2 of 8); 46.0% on ARC-AGI-2 (#3 of 3).
DeepSeek V4-Pro was released on April 24, 2026 as the flagship of DeepSeek’s fourth-generation V series. It achieved 93.5% on LiveCodeBench at launch — the highest competitive programming score of any measured model — and posted the highest self-reported SWE-bench Verified score outside of Claude Fable 5 at 80.6%. The release created significant discussion when CAISI independently measured V4-Pro’s SWE-bench score at 74.0% rather than the 80.6% DeepSeek had reported — a 6.6-point gap that DeepSeek attributed to scaffolding and token-budget differences in evaluation setup. As an open-weight model with a 1M context window and 49B/1.6T MoE architecture, V4-Pro became the leading self-hostable option for teams prioritising coding performance with full data control.
Note: SWE-bench discrepancy: SR 80.6% vs CAISI 74%. LiveCodeBench 93.5% highest reported.
Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.
| Benchmark | Score | Rank | Source | What it measures |
|---|---|---|---|---|
| SWE-bench Verified | 80.6% | 6 of 13 | SR | Real GitHub issue resolution (500 verified issues) |
| SWE-bench Pro | 55.4% | 9 of 9 | SR | Multi-language, standardised scaffold |
| GPQA Diamond | 90.1% | 6 of 14 | SR | Graduate-level Google-proof science Q&A |
| MMLU-Pro | 87.5% | 2 of 8 | SR | Hard knowledge reasoning, 12K questions |
| LiveCodeBench | 93.5% | 1 of 5 | SR | Competitive programming, continuously updated |
| Terminal-Bench | 67.9% | 4 of 8 | SR | Agentic Linux terminal task completion |
| ARC-AGI-2 | 46.0% | 3 of 3 | C | Novel visual reasoning, no memorisation |
| HLE | 37.7% | 4 of 6 | SR | Humanity's Last Exam (50+ STEM disciplines) |
| AA Intelligence Index | 52.0 pts | 6 of 9 | AA | Artificial Analysis composite score (independent) |
Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.
| Benchmark | DeepSeek V4-Pro | Claude Opus 4.8 | Claude Opus 4.7 |
|---|---|---|---|
| SWE-V | 80.6% | 88.6% | 87.6% |
| SWE-Pro | 55.4% | 69.2% | 64.3% |
| GPQA ◇ | 90.1% | 93.6% | 94.2% |
| MMLU-P | 87.5% | — | — |
| LiveCode | 93.5% | — | — |
| Terminal | 67.9% | 74.6% | 66.1% |
| ARC-2 | 46.0% | — | — |
| HLE | 37.7% | — | — |
| AA Idx | 52.0 pts | 61.4 pts | 57.3 pts |
DeepSeek is a Chinese AI research lab founded in 2023 as a subsidiary of High-Flyer, one of China's largest quantitative hedge funds. The lab became widely known in early 2025 when the release of DeepSeek V3 and R1 triggered the largest single-day drop in NVIDIA's market capitalisation to that date, as investors reassessed the capital intensity of frontier model development in light of DeepSeek's reported lower training costs. DeepSeek is unusual among frontier labs in its commitment to releasing open-weight models alongside its proprietary research, and in operating an active social-media presence that critiques closed-model pricing. Its V-series models (V3 through V4-Pro) are widely used as self-hostable coding and reasoning models in the West, despite ongoing regulatory and security scrutiny of Chinese-origin AI models in several jurisdictions.
Use Cases
DeepSeek V4-Pro scores 80% on SWE-bench Verified — ranked #6 of 13 models tracked here. It resolves real GitHub issues, writes production-quality code across multiple languages, and handles multi-step agentic coding workflows — a strong choice for AI-assisted software development teams.
At 90% on GPQA Diamond (#6 of 14 models), DeepSeek V4-Pro exceeds the ~65% human PhD-expert baseline on graduate-level biology, chemistry, and physics questions. Use it for scientific literature synthesis, hypothesis evaluation, medical and legal Q&A, and multi-discipline research tasks where deep domain knowledge matters.
DeepSeek V4-Pro achieves 67% on Terminal-Bench (#4 of 8 models), making it competent for common shell scripting, command-line workflows, and light system administration tasks inside agentic pipelines.
DeepSeek V4-Pro is open-weight (49B/1.6T MoE parameters) — its trained weights are publicly available for download and self-deployment. Run it on your own GPU hardware or private cloud to ensure data never leaves your environment, eliminate per-token API costs at scale, or fine-tune the model on proprietary datasets.
With a 1M-token context window, DeepSeek V4-Pro can ingest entire software repositories, long legal contracts, book-length research reports, or extensive conversation histories in a single prompt — enabling deep document Q&A, cross-file codebase reasoning, and comprehensive summarisation without chunking heuristics or RAG pipelines.
DeepSeek V4-Pro scores 93% on LiveCodeBench (#1 of 5 models), a contamination-resistant benchmark continuously refreshed with new competitive programming problems. At this level it handles advanced data structures, graph algorithms, and interview-level challenges with high reliability.
More to Compare
Compare all 24 models across 11 benchmarks — sortable, sourced, and updated June 2026.

DeepSeek V4-Flash is a open-weight AI model from DeepSeek (13B/284B MoE), classified as Advanced tier with a 1M tokens context window. It is available directly ...

Llama 4 Scout is a open-weight AI model from Meta (17B/109B MoE (16 experts)), classified as Established tier with a 10M tokens context window. It is available ...

Claude Opus 4.7 is a proprietary AI model from Anthropic, classified as Frontier tier with a 1M tokens context window. Explore benchmark scores, strengths, and ...
Cookie Consent
We use cookies to enhance your browsing experience and analyze our traffic. See our privacy policy.