Meta logo

Meta

Llama 4 Maverick

EstablishedOpen-weight

Llama 4 Maverick is a proven AI model still widely deployed in production systems, developed by Meta and released in April 2025. As an open-weight model, its trained weights are publicly available — you can self-host it, fine-tune it on proprietary data, or run it on-premise without API dependencies. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 3 benchmarks. Standout scores: 80.5% on MMLU-Pro (#4 of 8); 43.4% on LiveCodeBench (#4 of 5); 69.8% on GPQA Diamond (#11 of 14).

Llama 4 Maverick was released on April 5, 2025 as Meta’s flagship Llama 4 model. Maverick introduced a native multimodal mixture-of-experts architecture — the first Llama model capable of processing images at the reasoning level — using 17B active parameters from a 400B pool with 128 experts. At launch it was the strongest open-weight model on visual understanding tasks, enabling the first generation of vision-capable open-weight agents. Its 1M context window and multimodal capability made it the reference model for open-weight agentic application development through the first half of 2025, before the Frontier-tier closed models extended their lead.

Note: Best-in-class multimodal open-weight at launch. Fits on single H100 host.

Released
April 2025
Parameters
17B/400B MoE (128 experts)
Context
1M tokens
Weights
Open
Tier
Established
Provider
Meta

Llama 4 Maverick Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
GPQA Diamond69.8%11 of 14SRGraduate-level Google-proof science Q&A
MMLU-Pro80.5%4 of 8SRHard knowledge reasoning, 12K questions
LiveCodeBench43.4%4 of 5SRCompetitive programming, continuously updated

Detailed Benchmark Scores

GPQA ◇ SR
69.8%
Graduate-level Google-proof science Q&A
MMLU-P SR
80.5%
Hard knowledge reasoning, 12K questions
LiveCode SR
43.4%
Competitive programming, continuously updated

Llama 4 Maverick vs. Established Peers

Head-to-head benchmark comparison with other Established-tier models. Higher is better for all metrics.

BenchmarkLlama 4 MaverickGPT-4oLlama 4 Scout
GPQA ◇69.8%53.6%
MMLU-P80.5%72.6%52.2%
LiveCode43.4%32.8%

About Meta

Meta Founded 2004 · Menlo Park, CA
Website →

Meta Platforms (formerly Facebook) is one of the world's largest technology companies, headquartered in Menlo Park, California, and best known for Facebook, Instagram, WhatsApp, and Threads. Through Meta AI (formerly FAIR), the company has become the most prolific publisher of open-weight large language models via the Llama family. The Llama 4 series — Maverick, Scout, and Behemoth — were released under permissive community licences and remain the most downloaded open-weight models on platforms like Hugging Face. Meta's AI strategy is distinct among major labs: rather than operating a commercial API, Meta releases models openly for research and commercial use, integrates generative AI features across its own applications, and has publicly advocated for open-weight releases as the most democratic approach to AI development.

Use Cases

What to Use Llama 4 Maverick For

Scientific & Technical Reasoning
Self-Hosted & Private Deployments
Long-Document & Codebase Analysis

Strengths & Limitations

Strengths

  • Advanced scientific Q&A — 69% GPQA Diamond
  • Self-hosted & data-private deployments (open-weight)
  • Very long documents and large codebases (1M context)

Limitations

  • Significantly behind 2026 frontier models on agentic coding and complex reasoning

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

Llama 4 Scout — Benchmarks, Capabilities, and Use Cases
Llama 4 Scout — Benchmarks, Capabilities, and Use Cases

Llama 4 Scout — Benchmarks, Capabilities, and Use Cases

Llama 4 Scout is a open-weight AI model from Meta (17B/109B MoE (16 experts)), classified as Established tier with a 10M tokens context window. It is available ...

4 min read
Llama 3.3 70B — Benchmarks, Capabilities, and Use Cases
Llama 3.3 70B — Benchmarks, Capabilities, and Use Cases

Llama 3.3 70B — Benchmarks, Capabilities, and Use Cases

Llama 3.3 70B is a open-weight AI model from Meta (70B dense), classified as Established tier with a 128K tokens context window. It is available directly in Flo...

4 min read
GPT-4o — Benchmarks, Capabilities, and Use Cases
GPT-4o — Benchmarks, Capabilities, and Use Cases

GPT-4o — Benchmarks, Capabilities, and Use Cases

GPT-4o is a proprietary AI model from OpenAI, classified as Established tier with a 128K tokens context window. It is available directly in FlowHunt. Explore be...

4 min read