Skip to main content

DeepSeek V4 Pro

DeepSeek

๐Ÿง  Reasoning

An open-source powerhouse for code and math. Free to self-host under an MIT license.

Released April 24, 2026

Pricing

Input tokensN/A/M
Output tokensN/A/M

Capacity

Context windowNot published

Capabilities

โœ“Reasoning & Planning
โœ—Web Search
โœ“Open Source

Best Scores

GDPval-AA v21306.0 pts
GPQA Diamond90.1%
SWE-bench Verified80.6%
Terminal-Bench 2.167.9%
SWE-bench Pro55.4%

Best For

๐Ÿ’ปcoding
๐Ÿ”ฌresearch
๐Ÿ›debugging

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1306.0 pts

Independently measured by Artificial Analysis.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

90.1%

Provider-reported and independently reproduced by NIST CAISI.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

37.7%

Provider-reported; no independent run recorded yet.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

80.6%

Provider-reported; no independent run recorded yet.

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

67.9%

Provider-reported; an independent run by Vals AI (Terminus 2) lands at 50.19%.

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

55.4%

Provider-reported; no independent run recorded yet.

Reasoning

Artificial Analysis Intelligence Index

A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.

44.0 pts

Independently measured by Artificial Analysis.

Why Choose DeepSeek V4 Pro?

DeepSeek V4 Pro excels at coding and mathematical reasoning in open-source form. It scores in the 80s on SWE-Bench coding tasks and competes with closed-source flagships. Perfect for teams who need frontier-level capability without vendor lock-in.

How It Compares

vs GLM-5.2

DeepSeek is slightly cheaper to compute; GLM-5.2 is slightly more capable.

vs Claude Opus 4.8

Opus is closed-source and costs $5-25 per million tokens; DeepSeek is free.

You Might Also Like