Alibaba logo

Alibaba

Qwen 3.7-Max

FrontierAvailable in FlowHunt

Qwen 3.7-Max is a top-tier frontier AI model, developed by Alibaba and released in May 2026. It is a proprietary closed-weight model, available through the Alibaba API. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.

On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 8 benchmarks. Standout scores: 89.6% on MMLU-Pro (#1 of 8); 91.6% on LiveCodeBench (#2 of 5); 41.4% on HLE (#2 of 6).

Qwen 3.7-Max was released on May 20, 2026 as the top model in Alibaba’s Qwen 3 generation, introducing what Alibaba called a ’thinking mode’ — an extended chain-of-thought reasoning capability that can be toggled per query. In thinking mode, the model runs an internal deliberation step before producing its final answer, improving performance substantially on multi-step reasoning tasks. At launch Qwen 3.7-Max scored 92.4% on GPQA Diamond (third globally), 91.6% on LiveCodeBench, and 60.6% on SWE-bench Pro — the highest open-API score on that benchmark. Combined with a 1M context window and competitive API pricing, it established itself as the primary Frontier-tier alternative for cost-conscious enterprise teams.

Note: GPQA Diamond leader (92.4%). SWE-Pro leader (60.6%). AA Index 56.6.

Released
May 2026
Context
1M tokens
Weights
Closed
Tier
Frontier
Provider
Alibaba

Qwen 3.7-Max Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
SWE-bench Verified80.4%7 of 13SRReal GitHub issue resolution (500 verified issues)
SWE-bench Pro60.6%4 of 9SRMulti-language, standardised scaffold
GPQA Diamond92.4%4 of 14SRGraduate-level Google-proof science Q&A
MMLU-Pro89.6%1 of 8SRHard knowledge reasoning, 12K questions
LiveCodeBench91.6%2 of 5SRCompetitive programming, continuously updated
Terminal-Bench69.7%3 of 8SRAgentic Linux terminal task completion
HLE41.4%2 of 6SRHumanity's Last Exam (50+ STEM disciplines)
AA Intelligence Index56.6 pts4 of 9AAArtificial Analysis composite score (independent)

Detailed Benchmark Scores

SWE-V SR
80.4%
Real GitHub issue resolution (500 verified issues)
SWE-Pro SR
60.6%
Multi-language, standardised scaffold
GPQA ◇ SR
92.4%
Graduate-level Google-proof science Q&A
MMLU-P SR
89.6%
Hard knowledge reasoning, 12K questions
LiveCode SR
91.6%
Competitive programming, continuously updated
Terminal SR
69.7%
Agentic Linux terminal task completion
HLE SR
41.4%
Humanity's Last Exam (50+ STEM disciplines)
AA Idx AA
56.6 pts
Artificial Analysis composite score (independent)

Qwen 3.7-Max vs. Frontier Peers

Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.

BenchmarkQwen 3.7-MaxClaude Opus 4.8Claude Opus 4.7
SWE-V80.4%88.6%87.6%
SWE-Pro60.6%69.2%64.3%
GPQA ◇92.4%93.6%94.2%
MMLU-P89.6%
LiveCode91.6%
Terminal69.7%74.6%66.1%
HLE41.4%
AA Idx56.6 pts61.4 pts57.3 pts

About Alibaba

Alibaba Founded 1999 · Hangzhou, China
Website →

Alibaba Group is one of China's largest and most diverse technology conglomerates, founded in 1999 by Jack Ma and headquartered in Hangzhou. Through its DAMO Academy and Tongyi Lab, Alibaba has become a major force in large language model development with the Qwen (Tongyi Qianwen) family. The Qwen 3.7 generation, released in mid-2026, marked Alibaba's transition from domestic deployments to international benchmark competitiveness, with Qwen 3.7-Max scoring 92.4% on GPQA Diamond and 91.6% on LiveCodeBench. Alibaba's AI strategy is defined by open-weight releases as a counterbalance to US API-only models, integration of generative AI across its cloud and e-commerce businesses, and heavy investment in AI infrastructure across its Alibaba Cloud division.

Use Cases

What to Use Qwen 3.7-Max For

Software Engineering
Scientific & Technical Reasoning
Agentic CLI Automation
Long-Document & Codebase Analysis
Competitive Programming

Strengths & Limitations

Strengths

  • Strong software engineering — 80% SWE-bench Verified
  • Graduate-level science reasoning — 92% GPQA Diamond
  • Competitive programming — 91% LiveCodeBench
  • Composite intelligence — 56 pts AA Intelligence Index (independent)
  • Very long documents and large codebases (1M context)

Limitations

  • Closed-weight — cannot be self-hosted or fine-tuned on private data

Browse the Full AI Model Leaderboard

Frequently asked questions

Learn more

Qwen 3.7-Plus — Benchmarks, Capabilities, and Use Cases
Qwen 3.7-Plus — Benchmarks, Capabilities, and Use Cases

Qwen 3.7-Plus — Benchmarks, Capabilities, and Use Cases

Qwen 3.7-Plus is a proprietary AI model from Alibaba, classified as Advanced tier with a 1M tokens context window. It is available directly in FlowHunt. Explore...

3 min read
MiniMax M3 — Benchmarks, Capabilities, and Use Cases
MiniMax M3 — Benchmarks, Capabilities, and Use Cases

MiniMax M3 — Benchmarks, Capabilities, and Use Cases

MiniMax M3 is a open-weight AI model from MiniMax (MoE (undisclosed)), classified as Frontier tier with a 1M tokens context window. It is available directly in ...

4 min read
Nemotron 3 Ultra — Benchmarks, Capabilities, and Use Cases
Nemotron 3 Ultra — Benchmarks, Capabilities, and Use Cases

Nemotron 3 Ultra — Benchmarks, Capabilities, and Use Cases

Nemotron 3 Ultra is a open-weight AI model from NVIDIA (55B/550B MoE), classified as Frontier tier with a 262K tokens context window. It is available directly i...

3 min read