Skip to main content

Claude Fable 5

Anthropic

๐Ÿง  Reasoning๐ŸŒ Web Search

Anthropic's most capable model. Built for the hardest reasoning and long autonomous work, at a premium price.

Released June 9, 2026

Pricing

Input tokens$10.00/M
Output tokens$50.00/M

Capacity

Context window1.0M tokens
Max output128K tokens

Capabilities

โœ“Reasoning & Planning
โœ“Web Search
โœ—Open Source

Best Scores

GDPval-AA v21759.6 pts
SWE-bench Verified95.0%
GPQA Diamond92.6%
Terminal-Bench 2.180.5%
SWE-bench Pro80.0%

Best For

๐Ÿ’ปcoding
๐Ÿ”ฌresearch
๐Ÿ›debugging

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1759.6 pts

Provider-reported; no independent run recorded yet.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

95.0%

Provider-reported; no independent run recorded yet.

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

80.5%

Independently measured by vals.ai.

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

80.0%

Provider-reported; no independent run recorded yet.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

92.6%

Provider-reported; no independent run recorded yet.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

53.3%

Independently measured by Artificial Analysis.

Reasoning

Artificial Analysis Intelligence Index

A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.

59.9 pts

Provider-reported; no independent run recorded yet.

Why Choose Claude Fable 5?

Fable 5 is the strongest pure reasoning model on the market. If your task requires solving novel problems, multi-step planning, or deep scientific reasoning, Fable 5 consistently delivers. The 1M token context lets you work with entire codebases or research papers in a single request.

How It Compares

vs GPT-5.6 Sol

Fable 5 excels at complex multi-step reasoning and debugging. Sol is faster for immediate responses.

vs Gemini 3.1 Pro

Fable 5 edges out on reasoning-heavy tasks; Gemini wins on long-context retrieval and speed.

Release History

๐Ÿ†• NewJune 9, 2026

Claude Fable 5 released

Anthropic launches its most capable reasoning model with 1M token context and state-of-the-art performance on SWE-Bench. Access was paused June 12 under a US export directive and restored July 1.

You Might Also Like