Skip to main content
๐Ÿง  Reasoning๐ŸŒ Web Search

Moonshot AI's brand-new 2.8-trillion-parameter flagship. Frontier scores on agentic and reasoning work, with open weights promised within weeks.

Released July 16, 2026

Pricing

Input tokens$3.00/M
Output tokens$15.00/M

Capacity

Context window1.0M tokens

Capabilities

โœ“Reasoning & Planning
โœ“Web Search
โœ—Open Source

Best Scores

GDPval-AA v21686.0 pts
GPQA Diamond93.5%
SWE-bench Verified93.4%
Terminal-Bench 2.188.3%
Artificial Analysis Intelligence Index57.0 pts

Best For

๐Ÿ’ปcoding
๐Ÿ”ฌresearch
๐Ÿ“Šanalysis

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1686.0 pts

Independently measured by Artificial Analysis.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

93.5%

Provider-reported; no independent run recorded yet.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

56.0%

Provider-reported; no independent run recorded yet.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

93.4%

Independently measured by Vals AI (mini-swe-agent).

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

88.3%

Provider-reported; an independent run by Vals AI (Terminus 2) lands at 80.9%.

Reasoning

Artificial Analysis Intelligence Index

A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.

57.0 pts

Independently measured by Artificial Analysis.

Why Choose Kimi K3?

Kimi K3 posts some of the strongest published agentic numbers of any model, closed or open, and Moonshot has promised to release the weights, a rare combination of frontier capability now and self-hosting later. The 1M context and native vision make it a strong pick for long-horizon tool-use agents.

How It Compares

vs GPT-5.6 Sol

K3 is cheaper and matches Sol's published terminal scores; Sol has independent verification and a mature ecosystem.

vs GLM-5.2

K3 posts higher published agentic scores; GLM-5.2 is already downloadable and MIT-licensed today.

Release History

๐Ÿ†• NewJuly 16, 2026

Kimi K3 released

Moonshot AI launches its 2.8T-parameter flagship with 1M token context, native vision, and always-on reasoning. Open weights promised for late July.

You Might Also Like