Skip to main content
๐Ÿง  Reasoning

Thinking Machines' first model: a 975B-parameter open-weights multimodal flagship under Apache 2.0. The leading US open model on real-world coding.

Released July 15, 2026

Pricing

Input tokensN/A/M
Output tokensN/A/M

Capacity

Context window1.0M tokens

Capabilities

โœ“Reasoning & Planning
โœ—Web Search
โœ“Open Source

Best Scores

GDPval-AA v21237.0 pts
GPQA Diamond87.2%
SWE-bench Verified77.6%
Terminal-Bench 2.163.8%
SWE-bench Pro54.3%

Best For

๐Ÿ’ปcoding
๐Ÿ”ฌresearch
๐Ÿ“Šanalysis

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1237.0 pts

Independently measured by Artificial Analysis.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

87.2%

Provider-reported; no independent run recorded yet.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

46.0%

Provider-reported; no independent run recorded yet.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

77.6%

Provider-reported; no independent run recorded yet.

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

63.8%

Provider-reported; an independent run by Vals AI (Terminus 2) lands at 47.57%.

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

54.3%

Provider-reported; no independent run recorded yet.

Why Choose Inkling?

Inkling is the strongest US-built open-weights model, and the first flagship from Mira Murati's Thinking Machines Lab. It handles text, images, and audio natively, and the Apache 2.0 license is as business-friendly as it gets. Weights are on Hugging Face, with hosted APIs available from several providers.

How It Compares

vs GLM-5.2

Inkling leads on SWE-bench coding and is natively multimodal; GLM-5.2 is stronger on science reasoning and terminal work.

vs DeepSeek V4 Pro

Inkling adds vision and audio; DeepSeek edges it on pure coding benchmarks.

Release History

๐Ÿ†• NewJuly 15, 2026

Inkling released by Thinking Machines

Mira Murati's Thinking Machines Lab ships its first model: a 975B open-weights multimodal flagship under Apache 2.0, available on Hugging Face.

You Might Also Like