Skip to main content

GLM-5.2

Z.ai (GLM)

๐Ÿง  Reasoning

The current #1 open-source model. Frontier-level science reasoning you can download and run yourself, MIT-licensed.

Released June 13, 2026

Pricing

Input tokensN/A/M
Output tokensN/A/M

Capacity

Context window1.0M tokens

Capabilities

โœ“Reasoning & Planning
โœ—Web Search
โœ“Open Source

Best Scores

GDPval-AA v21510.0 pts
GPQA Diamond91.2%
Terminal-Bench 2.181.0%
SWE-bench Pro62.1%
Humanity's Last Exam40.5%

Best For

๐Ÿ”ฌresearch
๐Ÿ’ปcoding
๐Ÿ“Šanalysis

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1510.0 pts

Independently measured by Artificial Analysis.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

91.2%

Provider-reported; no independent run recorded yet.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

40.5%

Provider-reported; no independent run recorded yet.

Software Engineering

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

81.0%

Provider-reported; an independent run by Vals AI (Terminus 2) lands at 67.79%.

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

62.1%

Provider-reported; no independent run recorded yet.

Why Choose GLM-5.2?

GLM-5.2 is the strongest open-source model available. It beats most closed-source competitors on GPQA reasoning and has respectable coding scores. MIT-licensed means you can use it anywhere, from your laptop to production servers.

How It Compares

vs DeepSeek V4 Pro

GLM-5.2 has better reasoning; DeepSeek is slightly cheaper to self-host.

vs Claude Opus 4.8

Opus is more capable but costs $5-25 per million tokens; GLM is free to run.

You Might Also Like