Skip to main content

Claude Sonnet 5

Anthropic

๐Ÿง  Reasoning๐ŸŒ Web Search

The most agentic Sonnet yet, released June 30. Can make plans and run tools autonomously. Best value for daily work, now at intro pricing.

Released June 30, 2026

Pricing

Input tokens$2.00/M
Output tokens$10.00/M

Capacity

Context window1.0M tokens
Max output128K tokens

Capabilities

โœ“Reasoning & Planning
โœ“Web Search
โœ—Open Source

Best Scores

GDPval-AA v21603.0 pts
GPQA Diamond91.1%
SWE-bench Verified82.1%
Terminal-Bench 2.180.4%
SWE-bench Pro63.2%

Best For

โœ๏ธwriting
๐Ÿ“Šanalysis
๐Ÿ’ปcoding

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1603.0 pts

Independently measured by Artificial Analysis.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

91.1%

Independently measured by Artificial Analysis.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

57.4%

Independently measured by BenchLM.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

82.1%

Provider-reported; no independent run recorded yet.

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

80.4%

Provider-reported; an independent run by Vals AI (Terminus 2) lands at 74.53%.

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

63.2%

Provider-reported; no independent run recorded yet.

Reasoning

Artificial Analysis Intelligence Index

A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.

53.0 pts

Independently measured by Artificial Analysis.

Why Choose Claude Sonnet 5?

Sonnet 5 is the best all-rounder for teams on a budget. It can plan and execute multi-step workflows autonomously, reason through complex problems, and maintain 1M-token context. At intro pricing ($2/$10), it's unbeatable value.

How It Compares

vs Claude Haiku 4.5

Sonnet is significantly more capable; Haiku is 5x cheaper and better for simple tasks.

vs GPT-5.6 Terra

Sonnet now costs less at intro pricing; Terra is strong but Sonnet edges it on autonomy.

Release History

๐Ÿ’ฐ PriceJuly 18, 2026

Claude Sonnet 5 intro pricing extended

Intro pricing extended through August 31, 2026: $2/$10 per million tokens (input/output).

๐Ÿ†• NewJune 30, 2026

Claude Sonnet 5 released

New Sonnet model with autonomous planning, multi-step tool use, and 1M token context. Available at intro pricing.

You Might Also Like