See it on a graph
Start with one of the views below, then look for models in the sweet spot, usually high on performance, low on price. Tap or hover a point to see which model it is.
Which models give the most coding skill per dollar?
Zoomed in to the data: SWE-bench Verified (%) starts at 70. The baseline is not zero.
Can’t tap a point? Choose a model
Choose your own axes
Not shown (no published data on both axes): GPT-5.6 Terra, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Grok 4.1 Fast, Inkling, Mistral Medium 3.5, GLM-5.2, DeepSeek V4 Pro, Qwen 3.6, Llama 4 Maverick, Llama 4 Scout.