Alibaba logo

Alibaba

Qwen 3.7-Max

FrontierAvailable in FlowHunt

Qwen 3.7-Max 是一款顶级前沿 AI 模型,由阿里巴巴开发,于 2026 年 5 月发布。它是一款专有封闭权重模型,可通过阿里巴巴 API 使用。100 万 Token 的上下文窗口足以容纳整个代码库、长篇技术文档或长时间的智能代理会话而无需截断。

在 FlowHunt AI 排行榜上,它是追踪的 24 个模型之一,在 8 项基准测试上被评估。突出得分:MMLU-Pro 89.6%(8 个模型中排名第 1);LiveCodeBench 91.6%(5 个模型中排名第 2);HLE 41.4%(6 个模型中排名第 2)。

Qwen 3.7-Max 于 2026 年 5 月 20 日发布,是阿里巴巴 Qwen 3 系列中的顶级模型,引入了阿里巴巴称之为"思考模式"的功能——一种可针对每次查询切换的扩展链式推理能力。在思考模式下,模型在生成最终答案之前会进行内部推理步骤,从而显著提升多步推理任务的表现。发布时,Qwen 3.7-Max 在 GPQA Diamond 上得分为 92.4%(全球第三),在 LiveCodeBench 上得分为 91.6%,在 SWE-bench Pro 上得分为 60.6%——这是该基准测试中开放 API 的最高得分。结合 100 万 Token 上下文窗口和具有竞争力的 API 定价,它成为注重成本的企业团队的首要前沿层级替代方案。

注:GPQA Diamond 领先者(92.4%)。SWE-Pro 领先者(60.6%)。AA 指数 56.6。

发布时间
2026 年 5 月
上下文窗口
100 万 Token
权重
封闭
层级
前沿
提供商
阿里巴巴

Qwen 3.7-Max Benchmark Rankings

Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.

BenchmarkScoreRankSourceWhat it measures
SWE-bench Verified80.4%7 of 13SRReal GitHub issue resolution (500 verified issues)
SWE-bench Pro60.6%4 of 9SRMulti-language, standardised scaffold
GPQA Diamond92.4%4 of 14SRGraduate-level Google-proof science Q&A
MMLU-Pro89.6%1 of 8SRHard knowledge reasoning, 12K questions
LiveCodeBench91.6%2 of 5SRCompetitive programming, continuously updated
Terminal-Bench69.7%3 of 8SRAgentic Linux terminal task completion
HLE41.4%2 of 6SRHumanity's Last Exam (50+ STEM disciplines)
AA Intelligence Index56.6 pts4 of 9AAArtificial Analysis composite score (independent)

Detailed Benchmark Scores

SWE-V SR
80.4%
Real GitHub issue resolution (500 verified issues)
SWE-Pro SR
60.6%
Multi-language, standardised scaffold
GPQA ◇ SR
92.4%
Graduate-level Google-proof science Q&A
MMLU-P SR
89.6%
Hard knowledge reasoning, 12K questions
LiveCode SR
91.6%
Competitive programming, continuously updated
Terminal SR
69.7%
Agentic Linux terminal task completion
HLE SR
41.4%
Humanity's Last Exam (50+ STEM disciplines)
AA Idx AA
56.6 pts
Artificial Analysis composite score (independent)

Qwen 3.7-Max vs. Frontier Peers

Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.

BenchmarkQwen 3.7-MaxClaude Opus 4.8Claude Opus 4.7
SWE-V80.4%88.6%87.6%
SWE-Pro60.6%69.2%64.3%
GPQA ◇92.4%93.6%94.2%
MMLU-P89.6%
LiveCode91.6%
Terminal69.7%74.6%66.1%
HLE41.4%
AA Idx56.6 pts61.4 pts57.3 pts

About Alibaba

Alibaba Founded 1999 · Hangzhou, China
Website →

阿里巴巴集团是中国最大、最多元化的科技企业集团之一,由马云于1999年创立,总部位于杭州。通过其达摩院和通义实验室,阿里巴巴凭借Qwen(通义千问)系列成为大语言模型开发的重要力量。2026年中期发布的Qwen 3.7版本标志着阿里巴巴从国内部署转向国际基准竞争,Qwen 3.7-Max在GPQA Diamond上获得92.4%的分数,在LiveCodeBench上获得91.6%的分数。阿里巴巴的AI战略以开放权重发布作为对抗美国仅限API模型的制衡手段,将生成式AI整合到其云和电子商务业务中,并对其阿里云部门的AI基础设施进行大规模投资。

使用场景

Qwen 3.7-Max 适合做什么

软件工程
科学与技术推理
智能代理 CLI 自动化
长文档与代码库分析
竞赛编程

Strengths & Limitations

Strengths

  • 强大的软件工程——SWE-bench Verified 80%
  • 研究生级别的科学推理——GPQA Diamond 92%
  • 竞赛编程——LiveCodeBench 91%
  • 综合智能——AA Intelligence Index 56 分(独立验证)
  • 极长文档与大型代码库(100 万上下文)

Limitations

  • 封闭权重——无法自行托管或在私有数据上微调

浏览完整 AI 模型排行榜

常见问题

了解更多

Qwen 3.7-Plus — 基准测试、能力与用例
Qwen 3.7-Plus — 基准测试、能力与用例

Qwen 3.7-Plus — 基准测试、能力与用例

Qwen 3.7-Plus 是阿里巴巴推出的一款专有AI模型,属于高级(Advanced)级别,拥有100万tokens的上下文窗口,可在 FlowHunt 中直接使用。了解其基准测试得分、优势与推荐用例。...

2 分钟阅读
MiniMax M3 — 基准测试、能力与使用场景
MiniMax M3 — 基准测试、能力与使用场景

MiniMax M3 — 基准测试、能力与使用场景

MiniMax M3 是 MiniMax 推出的开放权重 AI 模型(MoE,未公开参数),属于前沿层级,拥有 100 万 token 的上下文窗口。该模型已在 FlowHunt 中直接可用。探索其基准测试分数、优势及推荐使用场景。...

2 分钟阅读
Qwen Max
Qwen Max

Qwen Max

将 FlowHunt 与 Qwen Max 集成,通过智能 AI 驱动的工作流,实现服务器操作自动化、系统健康监控,并大幅提升生产力,助力可扩展且可靠的基础设施管理。...

1 分钟阅读
AI Qwen Max +3