GPT-5.6 family released
OpenAI ships three tiers at once: Sol (flagship, $5/$30 per million tokens), Terra (balanced, $2.50/$15), and Luna (budget, $1/$6), all with a 1.05M token context window.
OpenAI's brand-new flagship. State of the art on autonomous terminal work, strong all-rounder.
Released July 9, 2026
GDPval-AA v2
Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.
Provider-reported; no independent run recorded yet.
SWE-bench Verified
Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.
Independently measured by Vals AI (mini-swe-agent).
Terminal-Bench 2.1
Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.
Provider-reported; an independent run by Vals AI (Terminus 2) lands at 85.77%.
SWE-bench Pro
The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.
Provider-reported; no independent run recorded yet.
GPQA Diamond
PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.
Provider-reported; no independent run recorded yet.
Humanity's Last Exam
PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.
Independently measured by Artificial Analysis.
Artificial Analysis Intelligence Index
A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.
Independently measured by Artificial Analysis.
Sol leads the market on autonomous terminal and DevOps tasks, executing shell commands and complex workflows better than any competitor. If you're automating infrastructure or running long-running agents, Sol is the benchmark.
Sol is faster and stronger on terminal tasks; Fable edges it on pure reasoning.
Sol excels at agentic work; Gemini is more aggressive on pricing.
OpenAI ships three tiers at once: Sol (flagship, $5/$30 per million tokens), Terra (balanced, $2.50/$15), and Luna (budget, $1/$6), all with a 1.05M token context window.