LLM Coding Leaderboard
Last update: October 3rd, 2026
Full methodology and scoring formulas explained in this article.
Compare: Effort levels (Low vs Medium vs High vs Max) →
| # | Model |
Total points (max 70) |
Avg cost per prompt |
Avg time per prompt |
Video | Tested with | Points per project | |||
|---|---|---|---|---|---|---|---|---|---|---|
|
Laravel Code Quality
(max 20) |
React-TS Code Quality
(max 20) |
Bug Finding Test
(max 10) |
Edge-case Testing (4 projects, max 20) |
|||||||
| 1 | GPT-6.1-Sol (High) | 68.1 | $0.22 | 07:26 | Codex CLI | 19.1 | 19 | 10 | 20 | |
| 2 | Opus 5.5 (High) | 67.41 | $0.79 | 03:10 | Claude Code | 18.5 | 19.33 | 9.58 | 20 | |
| 3 | GPT-6.1-Sol (Medium) | 67.09 | $0.14 | 03:53 | Codex CLI | 18.63 | 18.67 | 9.79 | 20 | |
| 4 | GPT-6-Astra (High) | 66.35 | $0.96 | 04:34 | Codex CLI | 18.6 | 18.17 | 9.58 | 20 | |
| 5 | GPT-6-Astra (Medium) | 66.03 | $0.69 | 03:08 | Codex CLI | 18.2 | 17.83 | 10 | 20 | |
| 6 | Opus 5.5 (Medium) | 65.5 | $0.56 | 02:04 | Claude Code | 18.95 | 19.67 | 8.13 | 18.75 | |
| 7 | GPT-6.1-Sol (Low) | 65.36 | $0.10 | 03:30 | Codex CLI | 18.03 | 18 | 9.58 | 19.75 | |
| 8 | Sonnet 5.5 (High) | 64.31 | $0.34 | 01:53 | Claude Code | 18.68 | 18.5 | 8.13 | 19 | |
| 9 | Fable 5.1 (Medium) | 61.03 | $1.53 | 03:06 | Claude Code | 18.55 | 16.83 | 7.4 | 18.25 | |
| 10 | Opus 5.5 (Low) | 60.57 | $0.37 | 01:13 | Claude Code | 17.52 | 17.17 | 6.88 | 19 | |
| 11 | Sonnet 5.5 (Medium) | 60.54 | $0.20 | 00:58 | Claude Code | 17.88 | 18.33 | 7.08 | 17.25 | |
| 12 | Grok 4.7 (High) | 60.09 | $1.85 | 23:27 | Cursor | 17.7 | 18.33 | 7.81 | 16.25 | |
| 13 | GPT-6-Sol (High) | 59.92 | $0.31 | 05:18 | Codex CLI | 17.67 | 17.5 | 7.5 | 17.25 | |
| 14 | GPT-6-Luna (Max) | 59.07 | $0.03 | 11:03 | Codex CLI | 18.12 | 17.17 | 7.08 | 16.7 | |
| 15 | Space Bunny (Max) | 57.19 | N/A | 19:19 | OpenCode | 17.61 | 16.5 | 4.58 | 18.5 | |
| 16 | SWE-2 (Max) | 57.07 | N/A | 11:59 | Devin | 17.5 | 17.83 | 7.29 | 14.45 | |
| 17 | Kimi K3 (High) | 56.89 | $0.63 | 12:49 | OpenCode | 18.4 | 17.5 | 6.04 | 14.95 | |
| 18 | Grok 4.6 (High) | 56.71 | $0.77 | 09:19 | Cursor | 17.4 | 18 | 6.56 | 14.75 | |
| 19 | GPT-6-Luna (Xhigh) | 56.4 | $0.03 | 10:08 | Codex CLI | 17.02 | 17.17 | 6.46 | 15.75 | |
| 20 | GPT-6-Sol (Medium) | 56.32 | $0.21 | 03:26 | Codex CLI | 17.17 | 16 | 6.25 | 16.9 | |
| 21 | GLM-5.3 (High) | 54.82 | $0.19 | 04:58 | OpenCode | 17.55 | 17.67 | 4.9 | 14.7 | |
| 22 | Deepseek-V4.1-Flash (High) | 54.35 | $0.03 | 03:17 | OpenCode | 16.65 | 18.67 | 5.83 | 13.2 | |
| 23 | MiMo-V2.6 Pro (High) | 53.81 | $0.09 | 15:04 | OpenCode | 17.41 | 18 | 6.25 | 12.15 | |
| 24 | Qwen 3.8 Max (0902) (High) | 53.39 | $0.40 | 11:49 | OpenCode | 17.75 | 18 | 4.69 | 12.95 | |
| 25 | Deepseek-V4.1-Flash (Max) | 53.31 | $0.05 | 05:15 | OpenCode | 17.7 | 18.17 | 4.69 | 12.75 | |
| 26 | MiMo-V2.6 Flash (High) | 53.29 | $0.04 | 17:56 | OpenCode | 17.83 | 17.5 | 7.81 | 10.15 | |
| 27 | GPT-6-Luna (High) | 53.04 | $0.01 | 03:51 | Codex CLI | 16.9 | 15.83 | 5.31 | 15 | |
| 28 | Muse Spark 1.3 (Max) | 53.01 | $0.01 | 03:42 | OpenCode | 15.6 | 18.33 | 4.58 | 14.5 | |
| 29 | Qwen 3.8 Flash (Max) | 52.64 | $0.04 | 08:38 | OpenCode | 16.35 | 17.67 | 5.42 | 13.2 | |
| 30 | Deepseek-V4-Pro-0813 (Max) | 50.35 | $0.05 | 09:06 | OpenCode | 17.3 | 16.5 | 4.9 | 11.65 | |
| 31 | Tencent Hy3 (High) | 49.87 | $0.05 | 06:35 | OpenCode | 16.8 | 14.33 | 2.29 | 16.45 | |
| 32 | Deepseek-V4-Pro-0813 (High) | 48.34 | $0.04 | 07:15 | OpenCode | 17.25 | 16.33 | 2.71 | 12.05 | |
| 33 | Qwen 3.8 27B (Xhigh) | 47.76 | $0.45 | 23:22 | OpenCode | 17.23 | 15.67 | 2.71 | 12.15 | |
| 34 | GLM-5.3-Flash (Max) | 47.51 | $0.02 | 08:54 | OpenCode | 17.2 | 17.17 | 4.69 | 8.45 | |
| 35 | Gemini-3.8-Flash (High) | 47.09 | N/A | 02:56 | Antigravity | 14.8 | 16.67 | 5.42 | 10.2 | |
| 36 | LongCat-2.5-Preview | 46.65 | N/A | 16:50 | OpenCode | 16.98 | 16.67 | 2.4 | 10.6 | |
| 37 | GPT-6-Luna (Medium) | 43.74 | $0.01 | 02:13 | Codex CLI | 16.33 | 14.17 | 3.54 | 9.7 | |
| 38 | Gemini-3.1-Pro | 42.36 | $0.52 | 03:04 | OpenCode | 12.55 | 15.33 | 4.38 | 10.1 | |
| 39 | Composer 2.5 | 41.39 | $0.10 | 02:41 | Cursor | 16 | 15.83 | 3.96 | 5.6 | |
| 40 | Minimax M3 (High) | 40.35 | $0.21 | 06:32 | OpenCode | 15.3 | 15 | 2.1 | 7.95 | |
Support Leaderboard Updates
Running dozens of prompts on multiple projects with different LLMs is costly.
Help keep it updated: get Premium membership.
Recent Premium Tutorials
Oct 6, 2026
I Tried /retro from Matt Pocock Skills v1.3
Sep 27, 2026
I Tested Opus 5.5 vs GPT-6-Astra as Code REVIEWERS
Explanations
- Full methodology and scoring formulas explained in this article
- Prices are calculated with API costs
- When a model is tested via a subscription plan rather than the API, there's no per-prompt price to measure, so the avg cost column shows N/A for comparison
Updates
- October 3rd: Added new project worth 10 points for finding bugs
- September 30th: Added GPT-6.1-Sol
- September 29th: Added Sonnet 5.5
- September 28th: Added LongCat-2.5-Preview
- September 27th: Added MiMo v2.6 Pro/Flash models
- September 25th: Added Space Bunny stealth model
- September 24th: Added GPT-6-Sol and GPT-6-Luna
- September 23rd: Added Opus 5.5
- September 22nd: Added Grok 4.7
- September 15th: Added SWE-2 model by Devin
- September 12th: Added new project React/Typescript code quality (20 points max)
- September 11th: Added Deepseek v4.1 Flash model
- September 9th: Added Qwen 3.8 Max (0902)
- September 8th: Added Fable 5.1 and re-calculate scoring on Code Quality project
- September 6th: Added GPT-6-Astra
- September 3rd: Added Gemini-3.8-Flash
- August 31st: New project evaluating CODE QUALITY (max 20 points)
- August 29th: Added Qwen3.8-Flash model
- August 27th: Added GLM-5.3-Flash (ex Ox Alpha) model
- August 20th: Added Qwen 3.8 27B
- August 18th: Added Gemini-3.7-Flash tested in Antigravity
- August 15th: Added GLM-5.3 model
- August 14th: Added DeepSeek-v4-Pro-0813 and Grok 4.6 models
- August 12th: Added new evaluation project no.4 with Go language
- August 7th: Added Muse Spark 1.2 and GPT-5.6-Terra High levels
- August 4th: Added Qwen 3.8 Max and NEW 3rd Project Flutter/Dart
- August 2nd: Added Max levels of GPT-5.6-Luna and Deepseek-v4-Flash
- August 1st: Re-tested GPT-5.6-Luna/Terra and Deepseek-v4-Flash with new prices
- July 27th: added Opus 5 High and GPT-5.6-Sol Medium
- July 27th: new v2 version of the leaderboard - with only two (but more complex) projects.
- July 26th: cleanup - removed older models: GPT-5.5, Kimi K2.6, Gemini 3.5 Flash, Sonnet 4.6
- July 22nd: added Gemini 3.6 Flash and Gemini 3.5 Flash Lite
- July 18th: added Kimi K3
- July 11th: added GPT-5.6 Luna and Terra
- July 9th: added Grok 4.5
- July 8th: added Tencent Hy3
- July 1st: added Sonnet 5
- June 24th: added GPT-5.4-Mini and Gemini-3.5-Flash
- June 23rd: added 5th benchmark project, removed Opus 4.7 and GLM-5.1, added GPT-5.4
- June 17th: added GLM-5.2
- June 13th: added Kimi-K2.7-Code
- June 5th: added Deepseek v4 Flash and removed Minimax M2.7
- June 4th: added Qwen 3.7 Plus
- June 2nd: added Qwen 3.7 Max and Tested with column
- June 2nd: added Minimax M3
- May 30th: added Avg Cost for all LLMs (based on API pricing)
- May 29th: added Claude Opus 4.8
- May 24th: added 4th benchmark project - for React/TypeScript
- May 20th: added Composer 2.5