Llama 4 Scout
The long-context champion. A 10-million-token window, enough to read hundreds of books at once.
Released April 5, 2025
Pricing
Capacity
Capabilities
Best Scores
Best For
Benchmark Scores
Specialized Skills
GDPval-AA v2
Real paid work from 44 different jobs, such as law, nursing, and software. Judges compare two answers side by side without knowing which model wrote them, and the winner gains rating points. A typical human expert scores 1000, so a higher number means the work was picked over a human more often.
Independently measured by Artificial Analysis.
Knowledge
GPQA Diamond
PhD-level science questions written so you cannot just Google the answer. Tests whether the model can reason about hard science.
Provider-reported; no independent run recorded yet.
Why Choose Llama 4 Scout?
Scout has the largest context window of any model: 10 million tokens. That's enough to load an entire codebase, hundreds of documents, or several books at once. Perfect for long-context retrieval and bulk processing tasks.
How It Compares
vs Llama 4 Maverick
Scout has 10x more context but is less capable on reasoning; Maverick is more general.
vs Grok 4.1 Fast
Scout is slightly longer (10M vs 2M) and more capable; Grok is cheaper.