Meta logo

Meta

Llama 3.3 70B

EstablishedOpen-weightAvailable in FlowHunt

Llama 3.3 70B is a proven AI model still widely deployed in production systems, developed by Meta and released in December 2024. As an open-weight model, its trained weights are publicly available — you can self-host it, fine-tune it on proprietary data, or run it on-premise without API dependencies. It supports a 128K-token context window, adequate for most documents and multi-turn applications.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 4 benchmarks. Standout scores: 88.4% on HumanEval (#1 of 2); 77.0% on MATH-500 (#3 of 4); 68.9% on MMLU-Pro (#7 of 8).

Llama 3.3 70B was released on December 6, 2024 as one of Meta’s final Llama 3 series models before the architectural transition to Llama 4. As a dense (non-MoE) 70B parameter model, it reflected the Llama 3 generation’s philosophy of straightforward scalable architecture: more parameters, better results, transparent structure. Its 128K context window and strong MMLU-Pro and SWE-bench scores made it the leading self-hosted model for teams with GPU resources adequate for dense 70B inference but not a MoE system. It remains popular in regulated industries — healthcare, finance, legal — where the transparency and auditability of a dense architecture, combined with open-weight availability, matter as much as raw capability.

Note: Near-405B quality at 70B cost. Text-only. IFEval 92.1%.

Released
December 2024
Parameters
70B dense
Context
128K tokens
Weights
Open
Tier
Established
Provider
Meta

Llama 3.3 70B Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
GPQA Diamond50.5%13 of 14SRGraduate-level Google-proof science Q&A
MMLU-Pro68.9%7 of 8SRHard knowledge reasoning, 12K questions
MATH-50077.0%3 of 4SRHendrycks math (500 problems)
HumanEval pass@188.4%1 of 2SRPython code generation (saturated at frontier)

Detailed Benchmark Scores

GPQA ◇ SR
50.5%
Graduate-level Google-proof science Q&A
MMLU-P SR
68.9%
Hard knowledge reasoning, 12K questions
MATH SR
77.0%
Hendrycks math (500 problems)
HumEval SR
88.4%
Python code generation (saturated at frontier)

Llama 3.3 70B vs. Established Peers

Head-to-head benchmark comparison with other Established-tier models. Higher is better for all metrics.

BenchmarkLlama 3.3 70BGPT-4oLlama 4 Maverick
GPQA ◇50.5%53.6%69.8%
MMLU-P68.9%72.6%80.5%
MATH77.0%
HumEval88.4%

About Meta

Meta Founded 2004 · Menlo Park, CA
Website →

Meta Platforms (formerly Facebook) is one of the world's largest technology companies, headquartered in Menlo Park, California, and best known for Facebook, Instagram, WhatsApp, and Threads. Through Meta AI (formerly FAIR), the company has become the most prolific publisher of open-weight large language models via the Llama family. The Llama 4 series — Maverick, Scout, and Behemoth — were released under permissive community licences and remain the most downloaded open-weight models on platforms like Hugging Face. Meta's AI strategy is distinct among major labs: rather than operating a commercial API, Meta releases models openly for research and commercial use, integrates generative AI features across its own applications, and has publicly advocated for open-weight releases as the most democratic approach to AI development.

Use Cases

What to Use Llama 3.3 70B For

Scientific & Technical Reasoning
Self-Hosted & Private Deployments

Strengths & Limitations

Strengths

  • Self-hosted & data-private deployments (open-weight)
  • Long-context tasks (128K context window)

Limitations

  • Below-frontier hard science reasoning (50% GPQA — frontier is 90%+)
  • Significantly behind 2026 frontier models on agentic coding and complex reasoning

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

Llama 4 Scout — Benchmarks, Capabilities, and Use Cases
Llama 4 Scout — Benchmarks, Capabilities, and Use Cases

Llama 4 Scout — Benchmarks, Capabilities, and Use Cases

Llama 4 Scout is a open-weight AI model from Meta (17B/109B MoE (16 experts)), classified as Established tier with a 10M tokens context window. It is available ...

4 min read
Llama 4 Maverick — Benchmarks, Capabilities, and Use Cases
Llama 4 Maverick — Benchmarks, Capabilities, and Use Cases

Llama 4 Maverick — Benchmarks, Capabilities, and Use Cases

Llama 4 Maverick is a open-weight AI model from Meta (17B/400B MoE (128 experts)), classified as Established tier with a 1M tokens context window. Explore bench...

4 min read
Mistral Large 3 — Benchmarks, Capabilities, and Use Cases
Mistral Large 3 — Benchmarks, Capabilities, and Use Cases

Mistral Large 3 — Benchmarks, Capabilities, and Use Cases

Mistral Large 3 is a open-weight AI model from Mistral (41B/675B MoE), classified as Advanced tier with a 256K tokens context window. It is available directly i...

4 min read