Skip to main content

Muse Spark 1.1

Meta

๐Ÿง  Reasoning๐ŸŒ Web Search

Meta's new flagship and its first paid, closed-weights model after the open Llama era. Built for agent work at an aggressive price.

Released July 9, 2026

Pricing

Input tokens$1.25/M
Output tokens$4.25/M

Capacity

Context window1.0M tokens
Max output256K tokens

Capabilities

โœ“Reasoning & Planning
โœ“Web Search
โœ—Open Source

Best Scores

GDPval-AA v21375.0 pts
GPQA Diamond88.4%
SWE-bench Verified82.0%
Terminal-Bench 2.169.3%
Humanity's Last Exam62.1%

Best For

๐Ÿ’ปcoding
๐Ÿ“Šanalysis
โœ๏ธwriting

Benchmark Scores

Specialized Skills

GDPval-AA v2

Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.

1375.0 pts

Independently measured by Artificial Analysis.

Knowledge

GPQA Diamond

PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.

88.4%

Independently measured by Artificial Analysis.

Humanity's Last Exam

PhD-level questions across many subjects. Tests deep reasoning on the hardest questions humans can ask.

62.1%

Independently measured by BenchLM.

Software Engineering

SWE-bench Verified

Hands the AI bugs from actual software projects and counts how many it fixes. Like a coding job interview, but with real work.

82.0%

Independently measured by BenchmarkList (mini-SWE-agent).

Terminal-Bench 2.1

Puts the AI in front of a computer terminal and asks it to finish multi-step tasks on its own. Measures how good an "AI agent" it is.

69.3%

Independently measured by Vals AI (Terminus 2).

SWE-bench Pro

The harder version of the coding test. Bigger codebases, trickier bugs. Scores drop for everyone, so the gaps between models become clearer.

61.5%

Provider-reported; no independent run recorded yet.

Why Choose Muse Spark 1.1?

Muse Spark 1.1 is Meta's first foray into paid models, and it's aggressively priced at $1.25/$4.25. At that price, it's one of the cheapest frontier flagships. The 256K max output is exceptional for generating large documents.

How It Compares

vs Claude Sonnet 5

At intro pricing, Sonnet is cheaper; Spark is strong but less proven.

vs GPT-5.6 Terra

Spark is cheaper; Terra has broader adoption.

Release History

๐Ÿ†• NewJuly 9, 2026

Muse Spark 1.1 released

Meta's first paid, closed-weights model after the open Llama era: 1M token context, an exceptional 256K max output, and aggressive $1.25/$4.25 pricing aimed at agent workloads.

You Might Also Like