Pedro Navarrete

Model notes

Back to blog

Deciding which AI model to use.

A benchmark score helps, but I also want to know what each task costs and how long it takes.

AI providers offer several model families, often with low, medium, high, and max effort settings. The names hint at what is newer or larger, but I choose with two questions: how well does it perform, and what does one completed task cost? Token prices miss models that burn more tokens or need several attempts.

The first chart compares the Artificial Analysis Intelligence Index with cost or time per task. The second uses DeepSWE for long-horizon software engineering work.

DeepSWE publishes token usage, agent duration, agent steps, and peak context for each run. That lets me keep its results fixed, reprice them with current token costs, and compare task duration.

How to read the chart

Up is better: the Artificial Analysis Intelligence Index in the first chart, and average pass@1 across included DeepSWE rollouts in the second. Left is cheaper, or faster in the time view. The best trade-offs sit near the top-left.

A model family shares one icon and color, and a line joins its effort settings. Darker families are closer to the Pareto frontier. The 50% to 95% score filter is synced across all three views, but each benchmark calculates its own cutoff. Selecting a provider does not change that baseline.

The marginal return rows compare adjacent effort settings in each DeepSWE model family: score change divided by estimated cost change. This is a finite difference, not a continuous derivative. A larger number means more benchmark performance per additional dollar. Only steps ending above the selected cutoff remain.

All providers · ≥50% · Frontier off

Legend · 15 model families
5.6 Sol 5.6 Terra 5.6 Luna GPT-5.5 5.4 5.4 mini 5.4 nano GPT-5 / 5.1 / mini Fable Opus Sonnet Gemini V4 Flash Grok Grok Build

Models in the same family share an icon, shade, and connecting line. Families closer to the Pareto frontier appear darker in light mode and lighter in dark mode. When two families are equally close, the one with higher benchmark score gets the stronger shade. Hover, tap, or focus a point to inspect the model.

What DeepSWE costs at today’s token prices

DeepSWE tests long-horizon software engineering agents on 113 tasks from real repositories. I keep its v1.1 scores and token usage fixed, then apply prices from August 20, 2026. Because input counts include cache hits, I subtract those hits, price them at the cache-read rate, price the rest at the regular input rate, and add output cost.

The agent-time view averages DeepSWE’s measured duration across included rollouts. It measures the coding agent’s work, not raw output speed, and excludes verifier runtime.

Muse Spark 1.2 has two points because Meta offers a separate Contributor tier. Both use identical results and token counts. The cheaper Contributor tier requires opting in to let Meta use prompts and completions for model training, so the privacy trade-off is different.

All providers · ≥50% · Frontier off

Legend · 15 model families
5.6 Sol 5.6 Terra 5.6 Luna GPT-5.5 5.4 Fable Opus Sonnet Gemini V4 Flash Grok Muse Spark Kimi Qwen GLM

Marginal return by effort step

Score gained per additional dollar between adjacent available effort settings. Higher is better; a negative return means the higher effort setting scored lower in this benchmark sample.

Synced cutoff · 50% of top score · 36.8 minimum

GPT-5.6 Luna Best: medium → high 294 pts / $1
Effort stepScore gainAdded costMarginal return
medium → high+33.0+$0.11294 pts / $1
high → xhigh+12.6+$0.1488.2 pts / $1
xhigh → max+10.3+$0.2443.3 pts / $1
Gemini 3.7 Flash Best: low → medium 60.8 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+11.7+$0.1960.8 pts / $1
medium → high-0.2+$0.15-1.46 pts / $1
GPT-5.6 Terra Best: medium → high 42.3 pts / $1
Effort stepScore gainAdded costMarginal return
medium → high+18.6+$0.4442.3 pts / $1
high → xhigh+6.4+$0.798.09 pts / $1
xhigh → max+9.4+$2.054.61 pts / $1
GPT-5.6 Sol Best: low → medium 20.0 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+15.7+$0.7920.0 pts / $1
medium → high+8.3+$1.605.20 pts / $1
high → xhigh+1.3+$1.201.11 pts / $1
xhigh → max+1.9+$3.290.59 pts / $1
GPT-5.5 Best: low → medium 17.4 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+27.0+$1.5517.4 pts / $1
medium → high+10.4+$2.354.43 pts / $1
high → xhigh+2.7+$2.011.32 pts / $1
Grok 4.6 Best: low → medium 10.7 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+25.8+$2.4110.7 pts / $1
medium → high-2.3+$0.94-2.45 pts / $1
high → xhigh+1.6+$1.111.39 pts / $1
Claude Sonnet 5 Best: low → medium 7.47 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+9.3+$1.247.47 pts / $1
medium → high+8.5+$2.203.84 pts / $1
high → xhigh+1.4+$2.940.49 pts / $1
xhigh → max+4.2+$9.590.44 pts / $1
Claude Opus 4.8 Best: low → medium 7.01 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+7.9+$1.127.01 pts / $1
medium → high+3.1+$0.823.78 pts / $1
high → xhigh+2.6+$3.650.71 pts / $1
xhigh → max+4.6+$5.130.90 pts / $1
GLM-5.2 Best: high → max 6.91 pts / $1
Effort stepScore gainAdded costMarginal return
high → max+7.5+$1.086.91 pts / $1
Claude Opus 5 Best: low → medium 6.79 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+10.8+$1.596.79 pts / $1
medium → high+3.9+$2.731.44 pts / $1
high → xhigh+0.3+$2.940.11 pts / $1
xhigh → max+0.5+$2.710.18 pts / $1
Claude Fable 5 Best: low → medium 2.58 pts / $1
Effort stepScore gainAdded costMarginal return
low → medium+5.8+$2.242.58 pts / $1
medium → high+3.2+$2.991.08 pts / $1
high → xhigh+1.3+$4.080.32 pts / $1
xhigh → max-0.2+$7.95-0.02 pts / $1

Models in the same family share an icon, shade, and connecting line. Families closer to the Pareto frontier appear darker in light mode and lighter in dark mode. When two families are equally close, the one with higher benchmark score gets the stronger shade. Hover, tap, or focus a point to inspect the model.

What the Pareto frontier means

A model is on the Pareto frontier if no other model beats it on both score and cost. The switch fades everything else and joins the frontier with a dotted line. I use that line as a shortlist: start cheap, then move up if the model struggles.

GPT-5.6 Luna’s price cut moved every effort setting farther left without changing its score. Some GPT-5.5 settings score higher, but Luna balances score and cost better. Intelligence breaks ties between families equally close to the frontier.

Claude Opus 5 reaches the top at higher effort settings, but cost and time climb quickly. High or xhigh is a better starting point than max when I do not need the last point of benchmark performance.

In the repriced DeepSWE data, DeepSeek V4 Pro at max effort reaches a 62.8 score for about $0.24 per task. GPT-5.6 Luna at max reaches 67.2 for about $0.54. Among the new higher-cost results, Grok 4.6 at medium reaches 67.5 for about $3.45, and GLM-5.3 at max reaches 69.0 for about $3.99. Claude Opus 5 at max still leads at 73.6, but its estimated cost is about $11.54 per task.

My current default

Right now I use GPT-5.6 Sol at medium effort by default and GPT-5.6 Luna for subagent work. Both sit on the Pareto frontier.

Benchmarks have limits

These numbers narrow the field, but they cannot show how a model will handle my code, instructions, or tools. I test a few candidates on real project work and start with the smallest one I can trust.