GPT-6-Sol Coding Benchmark
How GPT-6-Sol ranks on the LLM Coding Leaderboard, shown next to the three models directly above and below it, with the top three models included for reference. GPT-6-Sol was tested at 2 effort levels, each ranked separately.
- Last evaluated: 1 day ago (September 23, 2026), with Codex CLI.
- Last evaluated: 1 day ago (September 23, 2026), with Codex CLI.
| # | Model | Total points (max 60) |
Avg cost per prompt |
Avg time per prompt |
Video | Tested with | Points per project | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
Laravel Code Quality
(max 20) |
React-TS Code Quality
(max 20) |
CSV Import (PHP) | Offline Sync (PHP) | Bank Feed (Dart / Flutter) | Shipping Quotes (Go) | |||||||
| 1 | Opus 5.5 (High) | 57.83 | $0.79 | 03:10 | Claude Code | 18.5 | 19.33 | 5 | 5 | 5 | 5 | |
| 2 | Opus 5.5 (Medium) | 57.37 | $0.56 | 02:04 | Claude Code | 18.95 | 19.67 | 5 | 4.75 | 5 | 4 | |
| 3 | GPT-6-Astra (Medium) | 56.03 | $0.69 | 03:08 | Codex CLI | 18.2 | 17.83 | 5 | 5 | 5 | 5 | |
| ⋯ 1 place not shown ⋯ | ||||||||||||
| 5 | Opus 5 (High) | 53.72 | $1.65 | 06:33 | Claude Code | 18.55 | 17.67 | 5 | 3.5 | 4 | 5 | |
| 6 | Fable 5.1 (Medium) | 53.63 | $1.53 | 03:06 | Claude Code | 18.55 | 16.83 | 4.5 | 3.75 | 5 | 5 | |
| 7 | Opus 5 (Medium) | 52.57 | $1.10 | 03:48 | Claude Code | 17.65 | 17.17 | 5 | 3.25 | 4.5 | 5 | |
| 8 | GPT-6-Sol (High) | 52.42 | $0.31 | 05:18 | Codex CLI | 17.67 | 17.5 | 3.5 | 4.25 | 4.5 | 5 | |
| 9 | Grok 4.7 (High) | 52.28 | $1.85 | 23:27 | Cursor | 17.7 | 18.33 | 4 | 3.75 | 5 | 3.5 | |
| 10 | GPT-5.6-Sol (Medium) | 52.12 | $1.01 | 05:14 | Codex CLI | 17.45 | 16.67 | 5 | 4 | 4 | 5 | |
| 11 | GPT-6-Luna (Max) | 51.99 | $0.03 | 11:03 | Codex CLI | 18.12 | 17.17 | 5 | 4 | 3.2 | 4.5 | |
| ⋯ 1 place not shown ⋯ | ||||||||||||
| 13 | Kimi K3 (High) | 50.85 | $0.63 | 12:49 | OpenCode | 18.4 | 17.5 | 4.5 | 3.25 | 3.2 | 4 | |
| 14 | GPT-5.6-Terra (High) | 50.72 | $0.36 | 05:13 | Codex CLI | 17.3 | 16.67 | 5 | 4.75 | 2.5 | 4.5 | |
| 15 | Grok 4.6 (High) | 50.15 | $0.77 | 09:19 | Cursor | 17.4 | 18 | 4 | 3.75 | 4.5 | 2.5 | |
| 16 | GPT-6-Sol (Medium) | 50.07 | $0.21 | 03:26 | Codex CLI | 17.17 | 16 | 2.9 | 5 | 5 | 4 | |
| 17 | GPT-6-Luna (Xhigh) | 49.94 | $0.03 | 10:08 | Codex CLI | 17.02 | 17.17 | 4.5 | 3.75 | 3 | 4.5 | |
| 18 | GLM-5.3 (High) | 49.92 | $0.19 | 04:58 | OpenCode | 17.55 | 17.67 | 3.5 | 3.2 | 4.5 | 3.5 | |
| 19 | SWE-2 (Max) | 49.78 | N/A | 11:59 | Devin | 17.5 | 17.83 | 4 | 2.45 | 4 | 4 | |
How to read this
- Each score measures a model-and-harness configuration, not the model in isolation.
- Prices are calculated with API costs. Models tested on a subscription plan show N/A.
- Full methodology and scoring formulas are explained in this article.
See every tested model side by side on the full LLM Coding Leaderboard.