Skip to main content

Claude Opus 4.8

Anthropic

๐Ÿง  Reasoning๐ŸŒ Web Search

A coding workhorse. Near the top of the toughest coding benchmarks at half the price of Fable 5.

Released May 28, 2026

Pricing

Input tokens$5.00/M
Output tokens$25.00/M

Capacity

Context window1.0M tokens
Max output128K tokens

Capabilities

โœ“Reasoning & Planning
โœ“Web Search
โœ—Open Source

Best Scores

GDPval-AA v21600.1 pts
GPQA Diamond93.6%
SWE-bench Verified88.6%
Terminal-Bench 2.174.6%
SWE-bench Pro69.2%

Best For

๐Ÿ’ปcoding
๐Ÿ›debugging
๐Ÿ“Šanalysis

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1600.1 pts

Provider-reported; no independent run recorded yet.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

93.6%

Provider-reported; no independent run recorded yet.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

45.7%

Independently measured by Artificial Analysis.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

88.6%

Provider-reported; no independent run recorded yet.

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

74.6%

Provider-reported; an independent run by Vals AI (Terminus 2) lands at 71.91%.

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

69.2%

Provider-reported; no independent run recorded yet.

Why Choose Claude Opus 4.8?

Opus 4.8 is the sweet spot for professional developers. It scores near the top on coding benchmarks at half the cost of Fable 5, and it maintains the full 1M token context for working with large projects. The best balance of power and price.

How It Compares

vs Claude Sonnet 5

Opus is stronger on complex code; Sonnet is faster and cheaper for simpler tasks.

vs GPT-5.6 Sol

Opus leads on SWE-Bench coding; Sol has a slight edge on reasoning-heavy problems.

Release History

๐Ÿ’ฐ PriceJune 5, 2026

Claude Opus 4.8 price reduced

Input price reduced from $15 to $5 per million tokens. Output price halved from $60 to $25.

You Might Also Like