Skip to main content

Compare models

Higher is better on every test, and the best published score in each column is highlighted. A "β€”" means no score has been published yet. It never means the model scored zero.

What do these tests measure?
SWE-bench Verified
Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.
SWE-bench Pro
The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.
GPQA Diamond
PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.
Terminal-Bench 2.1
Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.
Humanity's Last Exam
PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.
GDPval-AA v2
Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.
Artificial Analysis Intelligence Index
A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.
Showing 23 of 23 models
Showing 23 of 23 models

Claude Opus 5

Released 7/24/2026

Anthropicflagship
Context:1M
Price:$5
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1861.0 pts
swe bench verified97.0%
gpqa diamond93.4%

Claude Fable 5

Released 6/9/2026

Anthropicflagship
Context:1M
Price:$10
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1759.6 pts
swe bench verified95.0%
gpqa diamond92.6%

Claude Opus 4.8

Released 5/28/2026

Anthropicflagship
Context:1M
Price:$5
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1600.1 pts
gpqa diamond93.6%
swe bench verified88.6%

Claude Sonnet 5

Released 6/30/2026

Anthropicbalanced
Context:1M
Price:$2
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1603.0 pts
gpqa diamond91.1%
swe bench verified82.1%

Claude Haiku 4.5

Released 10/15/2025

Context:200K
Price:$1
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa911.0 pts
swe bench verified73.3%

GPT-5.6 Sol

Released 7/9/2026

OpenAIflagship
Context:1.05M
Price:$5
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1747.8 pts
swe bench verified96.2%
gpqa diamond94.6%

GPT-5.6 Terra

Released 7/9/2026

OpenAIbalanced
Context:1.05M
Price:$2.50
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1593.0 pts
gpqa diamond92.9%
terminal bench87.4%

GPT-5.6 Luna

Released 7/9/2026

OpenAIfast
Context:1.05M
Price:$1
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1591.8 pts
swe bench verified93.0%
gpqa diamond92.3%

Gemini 3.1 Pro

Released 2/19/2026

Googleflagship
Context:1M
Price:$2
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa962.3 pts
gpqa diamond94.3%
swe bench verified80.6%

Gemini 3.6 Flash

Released 7/21/2026

Googlebalanced
Context:1.05M
Price:$1.50
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1423.0 pts
terminal bench78.0%
swe bench pro58.7%

Gemini 3.5 Flash-Lite

Released 7/21/2026

Googlefast
Context:1.05M
Price:$0.30
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1139.0 pts
swe bench pro54.2%
terminal bench54.0%

Gemini 3.5 Flash

Released 5/19/2026

Googlefast
Context:1M
Price:$1.50
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1348.8 pts
gpqa diamond92.2%
swe bench verified78.0%

Grok 4.5

Released 7/16/2026

xAIflagship
Context:500K
Price:$2
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1527.0 pts
gpqa diamond93.1%
swe bench verified86.6%

Grok 4.1 Fast

Released 11/19/2025

xAIfast
Context:2M
Price:$0.20
🧠 reasoning🌐 web

Top benchmarks:

gpqa diamond85.3%
swe bench pro70.0%
hle17.6%

Muse Spark 1.1

Released 7/9/2026

Metaflagship
Context:1M
Price:$1.25
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1375.0 pts
gpqa diamond88.4%
swe bench verified82.0%

Kimi K3

Released 7/16/2026

Moonshot AIflagship
Context:1M
Price:$3
🧠 reasoning🌐 webπŸ‘οΈ vision

Top benchmarks:

gdpval aa1686.0 pts
gpqa diamond93.5%
swe bench verified93.4%

Inkling

Released 7/15/2026

Context:1M
Price:β€”
🧠 reasoningπŸ‘οΈ visionopen

Top benchmarks:

gdpval aa1237.0 pts
gpqa diamond87.2%
swe bench verified77.6%

Mistral Medium 3.5

Released 4/28/2026

Mistral AIbalanced
Context:256K
Price:β€”
🧠 reasoning🌐 webπŸ‘οΈ visionopen

Top benchmarks:

aa intelligence index30.0 pts
terminal bench15.6%

GLM-5.2

Released 6/13/2026

Z.ai (GLM)flagship
Context:1M
Price:β€”
🧠 reasoningopen

Top benchmarks:

gdpval aa1510.0 pts
gpqa diamond91.2%
terminal bench81.0%

DeepSeek V4 Pro

Released 4/24/2026

DeepSeekflagship
Context:β€”
Price:β€”
🧠 reasoningopen

Top benchmarks:

gdpval aa1306.0 pts
gpqa diamond90.1%
swe bench verified80.6%

Qwen 3.6

Released 4/16/2026

Context:β€”
Price:β€”
🧠 reasoningπŸ‘οΈ visionopen

Top benchmarks:

gdpval aa1139.0 pts
gpqa diamond86.0%
swe bench verified73.4%

Llama 4 Maverick

Released 4/5/2025

Metaflagship
Context:1M
Price:β€”
πŸ‘οΈ visionopen

Top benchmarks:

gpqa diamond69.8%
gdpval aa7.0 pts

Llama 4 Scout

Released 4/5/2025

Metabalanced
Context:10M
Price:β€”
πŸ‘οΈ visionopen

Top benchmarks:

gdpval aa111.0 pts
gpqa diamond57.2%

*Open-source models are free to download and run yourself; hosted-API pricing varies by vendor. Scores are provider-published evals where available, otherwise independent leaderboard runs, collected 2026-07-30. The dot next to each score shows who measured it: an independent run, the provider's own number, the provider's number where an independent run differs (hover for details). β€œreasoning” means the model thinks step by step before answering; β€œweb” means the provider's assistant can search the live internet.