Gemini 3.1 Pro
Google's flagship. A top-tier reasoner with strong long-context skills at an aggressive price.
Released February 19, 2026
Pricing
Capacity
Capabilities
Best Scores
Best For
Benchmark Scores
Specialized Skills
GDPval-AA v2
Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.
Provider-reported; no independent run recorded yet.
Knowledge
GPQA Diamond
PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.
Independently measured by Artificial Analysis.
Humanity's Last Exam
PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.
Provider-reported; no independent run recorded yet.
Software Engineering
SWE-bench Verified
Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.
Provider-reported; no independent run recorded yet.
Terminal-Bench 2.1
Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.
Provider-reported; an independent run by Vals AI (Terminus 2) lands at 70.79%.
SWE-bench Pro
The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.
Independently measured by Anthropic's comparison table.
Reasoning
Artificial Analysis Intelligence Index
A frequently refreshed overall score made from nine modern tests: real work, tool use, terminal tasks, science, hard questions, and long-context reasoning. It is a scorecard rather than a percent correct.
Provider-reported; no independent run recorded yet.
Why Choose Gemini 3.1 Pro?
Gemini 3.1 Pro is Google's aggressive flagship. It competes on GPQA reasoning with the best models while underpricing competitors. The 1M context and web search make it excellent for research tasks.
How It Compares
vs Claude Fable 5
Gemini is cheaper and faster; Fable has a slight reasoning edge.
vs GPT-5.6 Sol
Gemini is more aggressive on pricing; Sol leads on terminal/agentic work.