
Qwen 3.7-Plus — Benchmarks, Capabilities, and Use Cases
Qwen 3.7-Plus is a proprietary AI model from Alibaba, classified as Advanced tier with a 1M tokens context window. It is available directly in FlowHunt. Explore...
Qwen 3.7-Max is a top-tier frontier AI model, developed by Alibaba and released in May 2026. It is a proprietary closed-weight model, available through the Alibaba API. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.
On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 8 benchmarks. Standout scores: 89.6% on MMLU-Pro (#1 of 8); 91.6% on LiveCodeBench (#2 of 5); 41.4% on HLE (#2 of 6).
Qwen 3.7-Max was released on May 20, 2026 as the top model in Alibaba’s Qwen 3 generation, introducing what Alibaba called a ’thinking mode’ — an extended chain-of-thought reasoning capability that can be toggled per query. In thinking mode, the model runs an internal deliberation step before producing its final answer, improving performance substantially on multi-step reasoning tasks. At launch Qwen 3.7-Max scored 92.4% on GPQA Diamond (third globally), 91.6% on LiveCodeBench, and 60.6% on SWE-bench Pro — the highest open-API score on that benchmark. Combined with a 1M context window and competitive API pricing, it established itself as the primary Frontier-tier alternative for cost-conscious enterprise teams.
Note: GPQA Diamond leader (92.4%). SWE-Pro leader (60.6%). AA Index 56.6.
Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.
| Benchmark | Score | Rank | Source | What it measures |
|---|---|---|---|---|
| SWE-bench Verified | 80.4% | 7 of 13 | SR | Real GitHub issue resolution (500 verified issues) |
| SWE-bench Pro | 60.6% | 4 of 9 | SR | Multi-language, standardised scaffold |
| GPQA Diamond | 92.4% | 4 of 14 | SR | Graduate-level Google-proof science Q&A |
| MMLU-Pro | 89.6% | 1 of 8 | SR | Hard knowledge reasoning, 12K questions |
| LiveCodeBench | 91.6% | 2 of 5 | SR | Competitive programming, continuously updated |
| Terminal-Bench | 69.7% | 3 of 8 | SR | Agentic Linux terminal task completion |
| HLE | 41.4% | 2 of 6 | SR | Humanity's Last Exam (50+ STEM disciplines) |
| AA Intelligence Index | 56.6 pts | 4 of 9 | AA | Artificial Analysis composite score (independent) |
Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.
| Benchmark | Qwen 3.7-Max | Claude Opus 4.8 | Claude Opus 4.7 |
|---|---|---|---|
| SWE-V | 80.4% | 88.6% | 87.6% |
| SWE-Pro | 60.6% | 69.2% | 64.3% |
| GPQA ◇ | 92.4% | 93.6% | 94.2% |
| MMLU-P | 89.6% | — | — |
| LiveCode | 91.6% | — | — |
| Terminal | 69.7% | 74.6% | 66.1% |
| HLE | 41.4% | — | — |
| AA Idx | 56.6 pts | 61.4 pts | 57.3 pts |
Alibaba Group is one of China's largest and most diverse technology conglomerates, founded in 1999 by Jack Ma and headquartered in Hangzhou. Through its DAMO Academy and Tongyi Lab, Alibaba has become a major force in large language model development with the Qwen (Tongyi Qianwen) family. The Qwen 3.7 generation, released in mid-2026, marked Alibaba's transition from domestic deployments to international benchmark competitiveness, with Qwen 3.7-Max scoring 92.4% on GPQA Diamond and 91.6% on LiveCodeBench. Alibaba's AI strategy is defined by open-weight releases as a counterbalance to US API-only models, integration of generative AI across its cloud and e-commerce businesses, and heavy investment in AI infrastructure across its Alibaba Cloud division.
Use Cases
Qwen 3.7-Max scores 80% on SWE-bench Verified — ranked #7 of 13 models tracked here. It resolves real GitHub issues, writes production-quality code across multiple languages, and handles multi-step agentic coding workflows — a strong choice for AI-assisted software development teams.
At 92% on GPQA Diamond (#4 of 14 models), Qwen 3.7-Max exceeds the ~65% human PhD-expert baseline on graduate-level biology, chemistry, and physics questions. Use it for scientific literature synthesis, hypothesis evaluation, medical and legal Q&A, and multi-discipline research tasks where deep domain knowledge matters.
Qwen 3.7-Max achieves 69% on Terminal-Bench (#3 of 8 models), making it competent for common shell scripting, command-line workflows, and light system administration tasks inside agentic pipelines.
With a 1M-token context window, Qwen 3.7-Max can ingest entire software repositories, long legal contracts, book-length research reports, or extensive conversation histories in a single prompt — enabling deep document Q&A, cross-file codebase reasoning, and comprehensive summarisation without chunking heuristics or RAG pipelines.
Qwen 3.7-Max scores 91% on LiveCodeBench (#2 of 5 models), a contamination-resistant benchmark continuously refreshed with new competitive programming problems. At this level it handles advanced data structures, graph algorithms, and interview-level challenges with high reliability.
More to Compare
Compare all 24 models across 11 benchmarks — sortable, sourced, and updated June 2026.

Qwen 3.7-Plus is a proprietary AI model from Alibaba, classified as Advanced tier with a 1M tokens context window. It is available directly in FlowHunt. Explore...

MiniMax M3 is a open-weight AI model from MiniMax (MoE (undisclosed)), classified as Frontier tier with a 1M tokens context window. It is available directly in ...

Nemotron 3 Ultra is a open-weight AI model from NVIDIA (55B/550B MoE), classified as Frontier tier with a 262K tokens context window. It is available directly i...
Cookie Consent
We use cookies to enhance your browsing experience and analyze our traffic. See our privacy policy.