GPT-5.6-Luna Coding Benchmark
How GPT-5.6-Luna ranks on the LLM Coding Leaderboard, shown next to the three models directly above and below it. GPT-5.6-Luna was tested at 5 effort levels, each ranked separately.
- Max effort — place #3, 17.5 points (max 20). Last evaluated on August 10, 2026, with Codex CLI.
- Xhigh effort — place #7, 15.75 points (max 20). Last evaluated on August 16, 2026, with Codex CLI.
- High effort — place #12, 14 points (max 20). Last evaluated on August 10, 2026, with Codex CLI.
- Medium effort — place #16, 12.5 points (max 20). Last evaluated on August 12, 2026, with Codex CLI.
- Low effort — place #29, 5.1 points (max 20). Last evaluated on August 11, 2026, with Codex CLI.
| # | Model | Total points (max 20) |
Avg cost per prompt |
Avg time per prompt |
Video | Tested with | Points per project (max 5) | |||
|---|---|---|---|---|---|---|---|---|---|---|
| CSV Import (PHP) | Offline Sync (PHP) | Bank Feed (Dart/Flutter) | Shipping Quotes (Go) | |||||||
| 1 | GPT-5.6-Sol (Medium) | 18 | $1.01 | 05:14 | Codex CLI | 5 | 4 | 4 | 5 | |
| 2 | Opus 5 (Medium) | 17.75 | $1.10 | 03:48 | Claude Code | 5 | 3.25 | 4.5 | 5 | |
| 3 | GPT-5.6-Luna (Max) | 17.5 | $0.10 | 14:18 | Codex CLI | 5 | 4.5 | 3 | 5 | |
| 4 | Opus 5 (High) | 17.5 | $1.65 | 06:33 | Claude Code | 5 | 3.5 | 4 | 5 | |
| 5 | Tencent Hy3 (High) | 16.45 | $0.05 | 06:35 | OpenCode | 3.2 | 3.75 | 4.5 | 5 | |
| 6 | GPT-5.6-Terra (Medium) | 16.45 | $0.20 | 02:44 | Codex CLI | 4.2 | 3.75 | 4 | 4.5 | |
| 7 | GPT-5.6-Luna (Xhigh) | 15.75 | $0.06 | 09:07 | Codex CLI | 4 | 4.75 | 2 | 5 | |
| 8 | Opus 4.8 (Medium) | 15.45 | $1.08 | 04:13 | Claude Code | 4.5 | 3.75 | 4 | 3.2 | |
| 9 | Kimi K3 | 14.95 | $0.63 | 12:49 | OpenCode | 4.5 | 3.25 | 3.2 | 4 | |
| 10 | Grok 4.6 (High) | 14.75 | $0.77 | 09:19 | Cursor | 4 | 3.75 | 4.5 | 2.5 | |
| 11 | GLM-5.3 (High) | 14.7 | $0.19 | 04:58 | OpenCode | 3.5 | 3.2 | 4.5 | 3.5 | |
| 12 | GPT-5.6-Luna (High) | 14 | $0.04 | 06:18 | Codex CLI | 4 | 4.5 | 1.5 | 4 | |
| 13 | Kimi K2.7 Code | 12.9 | $0.42 | 08:21 | OpenCode | 3 | 4 | 3.5 | 2.4 | |
| 14 | Grok 4.5 | 12.8 | $0.24 | 03:30 | OpenCode | 3.2 | 2.7 | 3.5 | 3.4 | |
| 15 | Deepseek-V4-Flash (Max) | 12.6 | $0.02 | 05:04 | OpenCode | 3.2 | 3.5 | 3.5 | 2.4 | |
| 16 | GPT-5.6-Luna (Medium) | 12.5 | $0.02 | 03:00 | Codex CLI | 4 | 4 | 1.5 | 3 | |
| 17 | Gemini-3.7-Flash (High) | 12.45 | N/A | 02:01 | Antigravity | 3.5 | 3.75 | 2.7 | 2.5 | |
| 18 | Qwen 3.8 27B (Xhigh) | 12.15 | $0.45 | 23:22 | OpenCode | 4.2 | 2.75 | 3 | 2.2 | |
| 19 | Deepseek-V4-Pro-0813 (High) | 12.05 | $0.04 | 07:15 | OpenCode | 2.4 | 3.25 | 4 | 2.4 | |
| ⋯ 6 places not shown ⋯ | ||||||||||
| 26 | Deepseek-V4-Flash (High) | 8.2 | $0.02 | 06:31 | OpenCode | 1.9 | 2.2 | 2.5 | 1.6 | |
| 27 | Minimax M3 | 7.95 | $0.21 | 06:32 | OpenCode | 3.5 | 0.95 | 1.9 | 1.6 | |
| 28 | Composer 2.5 | 5.6 | $0.10 | 02:41 | Cursor | 2.2 | 1.4 | 0.9 | 1.1 | |
| 29 | GPT-5.6-Luna (Low) | 5.1 | $0.01 | 01:36 | Codex CLI | 3 | 0.7 | 0.2 | 1.2 | |
Tutorials about GPT-5.6-Luna
PREMIUM
Article
· Aug 17, 2026
Save Money: GPT-5.6-Luna with Sol Review (Experiment)
Video
· Aug 1, 2026
I Re-Tested GPT-5.6-Luna and Deepseek-v4-Flash (New Benchmarks)
How to read this
- Each score measures a model-and-harness configuration, not the model in isolation.
- Prices are calculated with API costs. Models tested on a subscription plan show N/A.
- Full methodology and scoring formulas are explained in this article.
See every tested model side by side on the full LLM Coding Leaderboard.