
Nemotron 3 Ultra — Benchmarks, Capabilities, and Use Cases
Nemotron 3 Ultra is a open-weight AI model from NVIDIA (55B/550B MoE), classified as Frontier tier with a 262K tokens context window. It is available directly i...
Nemotron 3 Super is an advanced AI model suited for demanding production workloads, developed by NVIDIA and released in March 2026. As an open-weight model, its trained weights are publicly available — you can self-host it, fine-tune it on proprietary data, or run it on-premise without API dependencies. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.
On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 4 benchmarks. Standout scores: 18.3% on HLE (#6 of 6); 36.0 pts on AA Intelligence Index (#8 of 9); 79.2% on GPQA Diamond (#10 of 14).
Nemotron 3 Super was released on March 11, 2026 as NVIDIA’s mid-tier open-weight model, using 12B active parameters from a 120B total MoE pool. Super was designed for organisations running existing NVIDIA H100 and H200 GPU clusters who wanted strong inference performance without the hardware upgrade to Blackwell. With MMLU-Pro at 80.1% and SWE-bench Verified at 65.8%, it provided frontier-adjacent performance on the GPU generation that most enterprise customers had already deployed. Super fits the recurring pattern of NVIDIA releasing paired model sizes — one for next-generation hardware buyers, one for current installed base.
Note: 2.2x throughput vs GPT-OSS-120B. Reproducible eval suite. 1M context.
Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.
| Benchmark | Score | Rank | Source | What it measures |
|---|---|---|---|---|
| SWE-bench Verified | 60.5% | 11 of 13 | SR | Real GitHub issue resolution (500 verified issues) |
| GPQA Diamond | 79.2% | 10 of 14 | SR | Graduate-level Google-proof science Q&A |
| HLE | 18.3% | 6 of 6 | SR | Humanity's Last Exam (50+ STEM disciplines) |
| AA Intelligence Index | 36.0 pts | 8 of 9 | AA | Artificial Analysis composite score (independent) |
Head-to-head benchmark comparison with other Advanced-tier models. Higher is better for all metrics.
| Benchmark | Nemotron 3 Super | Claude Sonnet 4.6 | GPT-4.1 |
|---|---|---|---|
| SWE-V | 60.5% | 79.6% | 54.6% |
| GPQA ◇ | 79.2% | — | — |
| HLE | 18.3% | — | — |
| AA Idx | 36.0 pts | — | — |
NVIDIA Corporation designs and sells graphics processing units (GPUs) and system-on-chip units that have become the dominant hardware platform for AI model training and inference globally, with an estimated 80-to-95-percent market share in datacentre AI accelerators. Founded in 1993 by Jensen Huang, Chris Malachowsky, and Curtis Priem, NVIDIA has expanded from its original gaming-graphics focus into a full-stack AI computing company. Its Nemotron model family reflects a vertical integration strategy: by releasing open-weight models optimised specifically for NVIDIA hardware, the company aims to demonstrate the value of its Blackwell-generation GPU clusters while shaping the open-weight ecosystem around NVIDIA-optimised software stacks.
Use Cases
Nemotron 3 Super achieves 60% on SWE-bench Verified — ranked #11 of 13 models tracked here. It handles code generation, bug fixing, and code review for typical development tasks, though frontier models score significantly higher on fully autonomous complex engineering.
Nemotron 3 Super scores 79% on GPQA Diamond (#10 of 14 models), around or above the human PhD-expert baseline. It handles demanding scientific reasoning, research literature summarisation, and knowledge-intensive tasks across STEM disciplines.
Nemotron 3 Super is open-weight (12B/120B MoE parameters) — its trained weights are publicly available for download and self-deployment. Run it on your own GPU hardware or private cloud to ensure data never leaves your environment, eliminate per-token API costs at scale, or fine-tune the model on proprietary datasets.
With a 1M-token context window, Nemotron 3 Super can ingest entire software repositories, long legal contracts, book-length research reports, or extensive conversation histories in a single prompt — enabling deep document Q&A, cross-file codebase reasoning, and comprehensive summarisation without chunking heuristics or RAG pipelines.
More to Compare
Compare all 24 models across 11 benchmarks — sortable, sourced, and updated June 2026.

Nemotron 3 Ultra is a open-weight AI model from NVIDIA (55B/550B MoE), classified as Frontier tier with a 262K tokens context window. It is available directly i...

MiniMax M3 is a open-weight AI model from MiniMax (MoE (undisclosed)), classified as Frontier tier with a 1M tokens context window. It is available directly in ...

Mistral Large 3 is a open-weight AI model from Mistral (41B/675B MoE), classified as Advanced tier with a 256K tokens context window. It is available directly i...
Cookie Consent
We use cookies to enhance your browsing experience and analyze our traffic. See our privacy policy.