
GPT-5 — Benchmarks, Capabilities, and Use Cases
GPT-5 is a proprietary AI model from OpenAI, classified as Frontier tier with a 400K tokens context window. It is available directly in FlowHunt. Explore benchm...
GPT-5.5 is a top-tier frontier AI model, developed by OpenAI and released in April 2026. It is a proprietary closed-weight model, available through the OpenAI API. The 1M-token context window is large enough to hold entire codebases, lengthy technical documents, or extended agentic sessions without truncation.
On the FlowHunt AI Leaderboard it is one of 24 tracked models, evaluated across 6 benchmarks. Standout scores: 82.7% on Terminal-Bench (#1 of 8); 88.7% on SWE-bench Verified (#2 of 13); 93.6% on GPQA Diamond (#2 of 14).
GPT-5.5 was released by OpenAI on April 23, 2026 as an iterative refinement of GPT-5. At launch it matched Claude Opus 4.6 on GPQA Diamond at 93.6% and recorded the highest Terminal-Bench score of any model measured — 82.7% — reflecting particular strength in autonomous system administration, shell scripting, and DevOps automation. OpenAI positioned GPT-5.5 as the primary general-purpose model in its API, and it became the default model powering ChatGPT for most users. It also recorded 68.4% on SWE-bench Pro and 88.7% on SWE-bench Verified, placing it among the top three coding models globally.
Note: Highest SWE-bench Verified before Fable 5. Best OpenAI Terminal-Bench score.
Rank among all 24 models tracked on this leaderboard. SR = self-reported · 3P / AA / C = independently verified.
| Benchmark | Score | Rank | Source | What it measures |
|---|---|---|---|---|
| SWE-bench Verified | 88.7% | 2 of 13 | SR | Real GitHub issue resolution (500 verified issues) |
| SWE-bench Pro | 58.6% | 7 of 9 | SR | Multi-language, standardised scaffold |
| GPQA Diamond | 93.6% | 2 of 14 | SR | Graduate-level Google-proof science Q&A |
| Terminal-Bench | 82.7% | 1 of 8 | SR | Agentic Linux terminal task completion |
| HLE | 41.4% | 3 of 6 | SR | Humanity's Last Exam (50+ STEM disciplines) |
| AA Intelligence Index | 60.2 pts | 2 of 9 | AA | Artificial Analysis composite score (independent) |
Head-to-head benchmark comparison with other Frontier-tier models. Higher is better for all metrics.
| Benchmark | GPT-5.5 | Claude Opus 4.8 | Claude Opus 4.7 |
|---|---|---|---|
| SWE-V | 88.7% | 88.6% | 87.6% |
| SWE-Pro | 58.6% | 69.2% | 64.3% |
| GPQA ◇ | 93.6% | 93.6% | 94.2% |
| Terminal | 82.7% | 74.6% | 66.1% |
| HLE | 41.4% | — | — |
| AA Idx | 60.2 pts | 61.4 pts | 57.3 pts |
OpenAI was founded in December 2015 as a nonprofit AI research laboratory by Sam Altman, Elon Musk, Greg Brockman, Ilya Sutskever, John Schulman, and Wojciech Zaremba, with the mission of developing artificial general intelligence (AGI) that benefits all of humanity. In 2019 it created a capped-profit subsidiary, OpenAI LP, to raise institutional capital and scale model training. The company is known for the GPT series of language models (GPT-1 through GPT-5.5), the DALL-E image generation series, Whisper speech recognition, and Sora video generation. OpenAI operates the ChatGPT consumer application, a developer API, and enterprise services through Microsoft Azure. As of 2026 it maintains a close strategic partnership with Microsoft, which is the company's largest investor and cloud provider.
Use Cases
GPT-5.5 scores 88% on SWE-bench Verified — ranked #2 of 13 models tracked here. It resolves real GitHub issues, writes production-quality code across multiple languages, and handles multi-step agentic coding workflows — a strong choice for AI-assisted software development teams.
At 93% on GPQA Diamond (#2 of 14 models), GPT-5.5 exceeds the ~65% human PhD-expert baseline on graduate-level biology, chemistry, and physics questions. Use it for scientific literature synthesis, hypothesis evaluation, medical and legal Q&A, and multi-discipline research tasks where deep domain knowledge matters.
Scoring 82% on Terminal-Bench (#1 of 8 models), GPT-5.5 is well-suited for agentic DevOps and infrastructure automation. Terminal-Bench measures autonomous completion of Linux shell tasks — package management, file system operations, process control, and multi-step system administration — without human guidance mid-task.
With a 1M-token context window, GPT-5.5 can ingest entire software repositories, long legal contracts, book-length research reports, or extensive conversation histories in a single prompt — enabling deep document Q&A, cross-file codebase reasoning, and comprehensive summarisation without chunking heuristics or RAG pipelines.
More to Compare
Compare all 24 models across 11 benchmarks — sortable, sourced, and updated June 2026.

GPT-5 is a proprietary AI model from OpenAI, classified as Frontier tier with a 400K tokens context window. It is available directly in FlowHunt. Explore benchm...

GPT-5.4 is a proprietary AI model from OpenAI, classified as Frontier tier with a 1M tokens context window. It is available directly in FlowHunt. Explore benchm...

GPT-4.1 is a proprietary AI model from OpenAI, classified as Advanced tier with a 1M tokens context window. It is available directly in FlowHunt. Explore benchm...
Cookie Consent
We use cookies to enhance your browsing experience and analyze our traffic. See our privacy policy.