Model notes
Back to blogDeciding which AI model to use.
A benchmark score helps, but I also want to know what each task costs and how long it takes.
AI providers offer several model families, often with low, medium, high, and max effort settings. The names hint at what is newer or larger, but I choose with two questions: how well does it perform, and what does one completed task cost? Token prices miss models that burn more tokens or need several attempts.
The first chart compares the Artificial Analysis Intelligence Index with cost or time per task. The second uses DeepSWE for long-horizon software engineering work.
DeepSWE publishes token usage, agent duration, agent steps, and peak context for each run. That lets me keep its results fixed, reprice them with current token costs, and compare task duration.
How to read the chart
Up is better: the Artificial Analysis Intelligence Index in the first chart, and average pass@1 across included DeepSWE rollouts in the second. Left is cheaper, or faster in the time view. The best trade-offs sit near the top-left.
A model family shares one icon and color, and a line joins its effort settings. Darker families are closer to the Pareto frontier. The 50% to 95% score filter is synced across all three views, but each benchmark calculates its own cutoff. Selecting a provider does not change that baseline.
The marginal return rows compare adjacent effort settings in each DeepSWE model family: score change divided by estimated cost change. This is a finite difference, not a continuous derivative. A larger number means more benchmark performance per additional dollar. Only steps ending above the selected cutoff remain.
All providers · ≥50% · Frontier off
Legend · 15 model families
Models in the same family share an icon, shade, and connecting line. Families closer to the Pareto frontier appear darker in light mode and lighter in dark mode. When two families are equally close, the one with higher benchmark score gets the stronger shade. Hover, tap, or focus a point to inspect the model.
What DeepSWE costs at today’s token prices
DeepSWE tests long-horizon software engineering agents on 113 tasks from real repositories. I keep its v1.1 scores and token usage fixed, then apply prices from August 20, 2026. Because input counts include cache hits, I subtract those hits, price them at the cache-read rate, price the rest at the regular input rate, and add output cost.
The agent-time view averages DeepSWE’s measured duration across included rollouts. It measures the coding agent’s work, not raw output speed, and excludes verifier runtime.
Muse Spark 1.2 has two points because Meta offers a separate Contributor tier. Both use identical results and token counts. The cheaper Contributor tier requires opting in to let Meta use prompts and completions for model training, so the privacy trade-off is different.
All providers · ≥50% · Frontier off
Legend · 15 model families
Marginal return by effort step
Score gained per additional dollar between adjacent available effort settings. Higher is better; a negative return means the higher effort setting scored lower in this benchmark sample.
Synced cutoff · 50% of top score · 36.8 minimum
GPT-5.6 Luna Best: medium → high 294 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| medium → high | +33.0 | +$0.11 | 294 pts / $1 |
| high → xhigh | +12.6 | +$0.14 | 88.2 pts / $1 |
| xhigh → max | +10.3 | +$0.24 | 43.3 pts / $1 |
Gemini 3.7 Flash Best: low → medium 60.8 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +11.7 | +$0.19 | 60.8 pts / $1 |
| medium → high | -0.2 | +$0.15 | -1.46 pts / $1 |
GPT-5.6 Terra Best: medium → high 42.3 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| medium → high | +18.6 | +$0.44 | 42.3 pts / $1 |
| high → xhigh | +6.4 | +$0.79 | 8.09 pts / $1 |
| xhigh → max | +9.4 | +$2.05 | 4.61 pts / $1 |
GPT-5.6 Sol Best: low → medium 20.0 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +15.7 | +$0.79 | 20.0 pts / $1 |
| medium → high | +8.3 | +$1.60 | 5.20 pts / $1 |
| high → xhigh | +1.3 | +$1.20 | 1.11 pts / $1 |
| xhigh → max | +1.9 | +$3.29 | 0.59 pts / $1 |
GPT-5.5 Best: low → medium 17.4 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +27.0 | +$1.55 | 17.4 pts / $1 |
| medium → high | +10.4 | +$2.35 | 4.43 pts / $1 |
| high → xhigh | +2.7 | +$2.01 | 1.32 pts / $1 |
Grok 4.6 Best: low → medium 10.7 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +25.8 | +$2.41 | 10.7 pts / $1 |
| medium → high | -2.3 | +$0.94 | -2.45 pts / $1 |
| high → xhigh | +1.6 | +$1.11 | 1.39 pts / $1 |
Claude Sonnet 5 Best: low → medium 7.47 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +9.3 | +$1.24 | 7.47 pts / $1 |
| medium → high | +8.5 | +$2.20 | 3.84 pts / $1 |
| high → xhigh | +1.4 | +$2.94 | 0.49 pts / $1 |
| xhigh → max | +4.2 | +$9.59 | 0.44 pts / $1 |
Claude Opus 4.8 Best: low → medium 7.01 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +7.9 | +$1.12 | 7.01 pts / $1 |
| medium → high | +3.1 | +$0.82 | 3.78 pts / $1 |
| high → xhigh | +2.6 | +$3.65 | 0.71 pts / $1 |
| xhigh → max | +4.6 | +$5.13 | 0.90 pts / $1 |
GLM-5.2 Best: high → max 6.91 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| high → max | +7.5 | +$1.08 | 6.91 pts / $1 |
Claude Opus 5 Best: low → medium 6.79 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +10.8 | +$1.59 | 6.79 pts / $1 |
| medium → high | +3.9 | +$2.73 | 1.44 pts / $1 |
| high → xhigh | +0.3 | +$2.94 | 0.11 pts / $1 |
| xhigh → max | +0.5 | +$2.71 | 0.18 pts / $1 |
Claude Fable 5 Best: low → medium 2.58 pts / $1
| Effort step | Score gain | Added cost | Marginal return |
|---|---|---|---|
| low → medium | +5.8 | +$2.24 | 2.58 pts / $1 |
| medium → high | +3.2 | +$2.99 | 1.08 pts / $1 |
| high → xhigh | +1.3 | +$4.08 | 0.32 pts / $1 |
| xhigh → max | -0.2 | +$7.95 | -0.02 pts / $1 |
Models in the same family share an icon, shade, and connecting line. Families closer to the Pareto frontier appear darker in light mode and lighter in dark mode. When two families are equally close, the one with higher benchmark score gets the stronger shade. Hover, tap, or focus a point to inspect the model.
What the Pareto frontier means
A model is on the Pareto frontier if no other model beats it on both score and cost. The switch fades everything else and joins the frontier with a dotted line. I use that line as a shortlist: start cheap, then move up if the model struggles.
GPT-5.6 Luna’s price cut moved every effort setting farther left without changing its score. Some GPT-5.5 settings score higher, but Luna balances score and cost better. Intelligence breaks ties between families equally close to the frontier.
Claude Opus 5 reaches the top at higher effort settings, but cost and time climb quickly. High or xhigh is a better starting point than max when I do not need the last point of benchmark performance.
In the repriced DeepSWE data, DeepSeek V4 Pro at max effort reaches a 62.8 score for about $0.24 per task. GPT-5.6 Luna at max reaches 67.2 for about $0.54. Among the new higher-cost results, Grok 4.6 at medium reaches 67.5 for about $3.45, and GLM-5.3 at max reaches 69.0 for about $3.99. Claude Opus 5 at max still leads at 73.6, but its estimated cost is about $11.54 per task.
My current default
Right now I use GPT-5.6 Sol at medium effort by default and GPT-5.6 Luna for subagent work. Both sit on the Pareto frontier.
Benchmarks have limits
These numbers narrow the field, but they cannot show how a model will handle my code, instructions, or tools. I test a few candidates on real project work and start with the smallest one I can trust.