Skip to main content

Claude Opus 5

Anthropic

๐Ÿง  Reasoning๐ŸŒ Web Search

Anthropic's flagship for agentic coding and long business workflows. It comes close to Fable 5 at half the price.

Released July 24, 2026

Pricing

Input tokens$5.00/M
Output tokens$25.00/M

Capacity

Context window1.0M tokens
Max output128K tokens

Capabilities

โœ“Reasoning & Planning
โœ“Web Search
โœ—Open Source

Best Scores

GDPval-AA v21861.0 pts
SWE-bench Verified97.0%
GPQA Diamond93.4%
Terminal-Bench 2.184.6%
Artificial Analysis Intelligence Index61.0 pts

Best For

๐Ÿ’ปcoding
๐Ÿ“Šanalysis
๐Ÿ”ฌresearch

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1861.0 pts

Independently measured by Artificial Analysis.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

97.0%

Independently measured by vals.ai.

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

84.6%

Independently measured by vals.ai.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

93.4%

Independently measured by vals.ai.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

52.6%

Independently measured by Artificial Analysis.

Reasoning

Artificial Analysis Intelligence Index

A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.

61.0 pts

Independently measured by Artificial Analysis.

Why Choose Claude Opus 5?

Opus 5 is the strongest coding model tracked here, and it tops the SWE-Bench Verified leaderboard. It costs half of what Fable 5 costs, so it is the better default for everyday engineering work. The 1M token context lets you hand it a whole codebase at once.

How It Compares

vs Claude Fable 5

Opus 5 scores higher on coding benchmarks at half the price. Fable 5 still leads on the hardest reasoning tasks.

vs GPT-5.6 Sol

Opus 5 leads on SWE-Bench Verified. Sol edges ahead on Terminal-Bench.

Release History

๐Ÿ†• NewJuly 24, 2026

Claude Opus 5 released

Anthropic ships a flagship built for agentic coding and business workflows at $5/$25 per million tokens, which is half the price of Fable 5. It has a 1M token context window, and it now leads the SWE-Bench Verified leaderboard at 97%.

You Might Also Like