Compare models
Higher is better on every test, and the best published score in each column is highlighted. A "β" means no score has been published yet. It never means the model scored zero.
What do these tests measure?
- SWE-bench Verified
- Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.
- SWE-bench Pro
- The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.
- GPQA Diamond
- PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.
- Terminal-Bench 2.1
- Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.
- Humanity's Last Exam
- PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.
- GDPval-AA v2
- Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.
- Artificial Analysis Intelligence Index
- A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.
Claude Opus 5
Released 7/24/2026
Top benchmarks:
Claude Fable 5
Released 6/9/2026
Top benchmarks:
Claude Opus 4.8
Released 5/28/2026
Top benchmarks:
Claude Sonnet 5
Released 6/30/2026
Top benchmarks:
Claude Haiku 4.5
Released 10/15/2025
Top benchmarks:
GPT-5.6 Sol
Released 7/9/2026
Top benchmarks:
GPT-5.6 Terra
Released 7/9/2026
Top benchmarks:
GPT-5.6 Luna
Released 7/9/2026
Top benchmarks:
Gemini 3.1 Pro
Released 2/19/2026
Top benchmarks:
Gemini 3.6 Flash
Released 7/21/2026
Top benchmarks:
Gemini 3.5 Flash-Lite
Released 7/21/2026
Top benchmarks:
Gemini 3.5 Flash
Released 5/19/2026
Top benchmarks:
Grok 4.5
Released 7/16/2026
Top benchmarks:
Grok 4.1 Fast
Released 11/19/2025
Top benchmarks:
Muse Spark 1.1
Released 7/9/2026
Top benchmarks:
Kimi K3
Released 7/16/2026
Top benchmarks:
Inkling
Released 7/15/2026
Top benchmarks:
Mistral Medium 3.5
Released 4/28/2026
Top benchmarks:
GLM-5.2
Released 6/13/2026
Top benchmarks:
DeepSeek V4 Pro
Released 4/24/2026
Top benchmarks:
Qwen 3.6
Released 4/16/2026
Top benchmarks:
Llama 4 Maverick
Released 4/5/2025
Top benchmarks:
Llama 4 Scout
Released 4/5/2025
Top benchmarks:
*Open-source models are free to download and run yourself; hosted-API pricing varies by vendor. Scores are provider-published evals where available, otherwise independent leaderboard runs, collected 2026-07-30. The dot next to each score shows who measured it: an independent run, the provider's own number, the provider's number where an independent run differs (hover for details). βreasoningβ means the model thinks step by step before answering; βwebβ means the provider's assistant can search the live internet.