LLM Effort Levels Compared
Last update: September 30th, 2026
GPT-6.1-Sol: Low vs Medium vs High
Tested with Codex CLI.
Opus 5.5: Low vs Medium vs High
Tested with Claude Code.
GPT-6-Astra: Medium vs High
Tested with Codex CLI.
Sonnet 5.5: Medium vs High
Tested with Claude Code.
GPT-6-Luna: Medium vs High vs Xhigh vs Max
Tested with Codex CLI.
Deepseek-V4.1-Flash: High vs Max
Tested with OpenCode.
How to read this
- The sweet spot is the lowest effort level that scores within 1 point of the model's best result.
- Score bars start at 30 points, so the differences between effort levels are easier to see.
- Each score measures a model-and-harness configuration, not the model in isolation.
- Prices are calculated with API costs. Models tested on a subscription plan show N/A.
- Full methodology and scoring formulas are explained in this article.
See every tested model side by side on the full LLM Coding Leaderboard.