
Llama 4 Scout — Benchmarks, Capabilities, and Use Cases
Llama 4 Scout is a open-weight AI model from Meta (17B/109B MoE (16 experts)), classified as Established tier with a 10M tokens context window. It is available ...
Llama 4 Maverick is a proven AI model still widely deployed in production systems, developed by Meta and released in April 2025. As an open-weight model, its trained weights are publicly available — you can self-host it, fine-tune it on proprietary data, or run it on-premise without API dependencies. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.
On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 3 benchmarks. Standout scores: 80.5% on MMLU-Pro (#4 of 8); 43.4% on LiveCodeBench (#4 of 5); 69.8% on GPQA Diamond (#11 of 14).
Llama 4 Maverick was released on April 5, 2025 as Meta’s flagship Llama 4 model. Maverick introduced a native multimodal mixture-of-experts architecture — the first Llama model capable of processing images at the reasoning level — using 17B active parameters from a 400B pool with 128 experts. At launch it was the strongest open-weight model on visual understanding tasks, enabling the first generation of vision-capable open-weight agents. Its 1M context window and multimodal capability made it the reference model for open-weight agentic application development through the first half of 2025, before the Frontier-tier closed models extended their lead.
Note: Best-in-class multimodal open-weight at launch. Fits on single H100 host.
Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.
| Benchmark | Score | Rank | Source | What it measures |
|---|---|---|---|---|
| GPQA Diamond | 69.8% | 11 of 14 | SR | Graduate-level Google-proof science Q&A |
| MMLU-Pro | 80.5% | 4 of 8 | SR | Hard knowledge reasoning, 12K questions |
| LiveCodeBench | 43.4% | 4 of 5 | SR | Competitive programming, continuously updated |
Head-to-head benchmark comparison with other Established-tier models. Higher is better for all metrics.
| Benchmark | Llama 4 Maverick | GPT-4o | Llama 4 Scout |
|---|---|---|---|
| GPQA ◇ | 69.8% | 53.6% | — |
| MMLU-P | 80.5% | 72.6% | 52.2% |
| LiveCode | 43.4% | — | 32.8% |
Meta Platforms (formerly Facebook) is one of the world's largest technology companies, headquartered in Menlo Park, California, and best known for Facebook, Instagram, WhatsApp, and Threads. Through Meta AI (formerly FAIR), the company has become the most prolific publisher of open-weight large language models via the Llama family. The Llama 4 series — Maverick, Scout, and Behemoth — were released under permissive community licences and remain the most downloaded open-weight models on platforms like Hugging Face. Meta's AI strategy is distinct among major labs: rather than operating a commercial API, Meta releases models openly for research and commercial use, integrates generative AI features across its own applications, and has publicly advocated for open-weight releases as the most democratic approach to AI development.
Use Cases
Llama 4 Maverick scores 69% on GPQA Diamond (#11 of 14 models), around or above the human PhD-expert baseline. It handles demanding scientific reasoning, research literature summarisation, and knowledge-intensive tasks across STEM disciplines.
Llama 4 Maverick is open-weight (17B/400B MoE (128 experts) parameters) — its trained weights are publicly available for download and self-deployment. Run it on your own GPU hardware or private cloud to ensure data never leaves your environment, eliminate per-token API costs at scale, or fine-tune the model on proprietary datasets.
With a 1M-token context window, Llama 4 Maverick can ingest entire software repositories, long legal contracts, book-length research reports, or extensive conversation histories in a single prompt — enabling deep document Q&A, cross-file codebase reasoning, and comprehensive summarisation without chunking heuristics or RAG pipelines.
More to Compare
Compare all 24 models across 11 benchmarks — sortable, sourced, and updated June 2026.

Llama 4 Scout is a open-weight AI model from Meta (17B/109B MoE (16 experts)), classified as Established tier with a 10M tokens context window. It is available ...

Llama 3.3 70B is a open-weight AI model from Meta (70B dense), classified as Established tier with a 128K tokens context window. It is available directly in Flo...

GPT-4o is a proprietary AI model from OpenAI, classified as Established tier with a 128K tokens context window. It is available directly in FlowHunt. Explore be...
Cookie Consent
We use cookies to enhance your browsing experience and analyze our traffic. See our privacy policy.