Back to the full leaderboard

LLM Effort Levels Compared

Last update: September 30th, 2026

GPT-6.1-Sol: Low vs Medium vs High

Tested with Codex CLI.

Effort Total points (max 60) Avg cost
per prompt
Avg time
per prompt
vs previous level Video
Low
55.78
$0.10 03:30 —
Medium Sweet spot
57.3
$0.14 03:53 +1.52 pts · +46% cost · +00:23 time
High
58.1
$0.22 07:26 +0.80 pts · +59% cost · +03:33 time

Opus 5.5: Low vs Medium vs High

Tested with Claude Code.

Effort Total points (max 60) Avg cost
per prompt
Avg time
per prompt
vs previous level Video
Low
53.69
$0.37 01:13 —
Medium Sweet spot
57.37
$0.56 02:04 +3.68 pts · +53% cost · +00:51 time
High
57.83
$0.79 03:10 +0.46 pts · +42% cost · +01:06 time

GPT-6-Astra: Medium vs High

Tested with Codex CLI.

Effort Total points (max 60) Avg cost
per prompt
Avg time
per prompt
vs previous level Video
Medium Sweet spot
56.03
$0.69 03:08 —
High
56.77
$0.96 04:34 +0.74 pts · +39% cost · +01:26 time

Sonnet 5.5: Medium vs High

Tested with Claude Code.

Effort Total points (max 60) Avg cost
per prompt
Avg time
per prompt
vs previous level Video
Medium
53.46
$0.20 00:58 —
High Sweet spot
56.18
$0.34 01:53 +2.72 pts · +67% cost · +00:55 time

GPT-6-Luna: Medium vs High vs Xhigh vs Max

Tested with Codex CLI.

Effort Total points (max 60) Avg cost
per prompt
Avg time
per prompt
vs previous level Video
Medium
40.2
$0.01 02:13 —
High
47.73
$0.01 03:51 +7.53 pts · +15% cost · +01:38 time
Xhigh
49.94
$0.03 10:08 +2.21 pts · +122% cost · +06:17 time
Max Sweet spot
51.99
$0.03 11:03 +2.05 pts · +12% cost · +00:55 time

Deepseek-V4.1-Flash: High vs Max

Tested with OpenCode.

Effort Total points (max 60) Avg cost
per prompt
Avg time
per prompt
vs previous level Video
High Sweet spot
48.52
$0.03 03:17 —
Max
48.62
$0.05 05:15 +0.10 pts · +61% cost · +01:58 time

How to read this

  • The sweet spot is the lowest effort level that scores within 1 point of the model's best result.
  • Score bars start at 30 points, so the differences between effort levels are easier to see.
  • Each score measures a model-and-harness configuration, not the model in isolation.
  • Prices are calculated with API costs. Models tested on a subscription plan show N/A.
  • Full methodology and scoring formulas are explained in this article.

See every tested model side by side on the full LLM Coding Leaderboard.

Povilas Korop

Get Weekly AI Coding News

Sent every Wednesday. No spam, ever. Unsubscribe anytime.