Back to Homepage

LLM Coding Leaderboard

Last update: October 3rd, 2026

Full methodology and scoring formulas explained in this article.

Compare: Effort levels (Low vs Medium vs High vs Max) →

# Model Total points
(max 70)
Avg cost
per prompt
Avg time
per prompt
Video Tested with Points per project
Laravel Code Quality
(max 20)
React-TS Code Quality
(max 20)
Bug Finding Test
(max 10)
Edge-case Testing
(4 projects, max 20)
1 GPT-6.1-Sol (High) 68.1 $0.22 07:26 Codex CLI 19.1 19 10 20
2 Opus 5.5 (High) 67.41 $0.79 03:10 Claude Code 18.5 19.33 9.58 20
3 GPT-6.1-Sol (Medium) 67.09 $0.14 03:53 Codex CLI 18.63 18.67 9.79 20
4 GPT-6-Astra (High) 66.35 $0.96 04:34 Codex CLI 18.6 18.17 9.58 20
5 GPT-6-Astra (Medium) 66.03 $0.69 03:08 Codex CLI 18.2 17.83 10 20
6 Opus 5.5 (Medium) 65.5 $0.56 02:04 Claude Code 18.95 19.67 8.13 18.75
7 GPT-6.1-Sol (Low) 65.36 $0.10 03:30 Codex CLI 18.03 18 9.58 19.75
8 Sonnet 5.5 (High) 64.31 $0.34 01:53 Claude Code 18.68 18.5 8.13 19
9 Fable 5.1 (Medium) 61.03 $1.53 03:06 Claude Code 18.55 16.83 7.4 18.25
10 Opus 5.5 (Low) 60.57 $0.37 01:13 Claude Code 17.52 17.17 6.88 19
11 Sonnet 5.5 (Medium) 60.54 $0.20 00:58 Claude Code 17.88 18.33 7.08 17.25
12 Grok 4.7 (High) 60.09 $1.85 23:27 Cursor 17.7 18.33 7.81 16.25
13 GPT-6-Sol (High) 59.92 $0.31 05:18 Codex CLI 17.67 17.5 7.5 17.25
14 GPT-6-Luna (Max) 59.07 $0.03 11:03 Codex CLI 18.12 17.17 7.08 16.7
15 Space Bunny (Max) 57.19 N/A 19:19 OpenCode 17.61 16.5 4.58 18.5
16 SWE-2 (Max) 57.07 N/A 11:59 Devin 17.5 17.83 7.29 14.45
17 Kimi K3 (High) 56.89 $0.63 12:49 OpenCode 18.4 17.5 6.04 14.95
18 Grok 4.6 (High) 56.71 $0.77 09:19 Cursor 17.4 18 6.56 14.75
19 GPT-6-Luna (Xhigh) 56.4 $0.03 10:08 Codex CLI 17.02 17.17 6.46 15.75
20 GPT-6-Sol (Medium) 56.32 $0.21 03:26 Codex CLI 17.17 16 6.25 16.9
21 GLM-5.3 (High) 54.82 $0.19 04:58 OpenCode 17.55 17.67 4.9 14.7
22 Deepseek-V4.1-Flash (High) 54.35 $0.03 03:17 OpenCode 16.65 18.67 5.83 13.2
23 MiMo-V2.6 Pro (High) 53.81 $0.09 15:04 OpenCode 17.41 18 6.25 12.15
24 Qwen 3.8 Max (0902) (High) 53.39 $0.40 11:49 OpenCode 17.75 18 4.69 12.95
25 Deepseek-V4.1-Flash (Max) 53.31 $0.05 05:15 OpenCode 17.7 18.17 4.69 12.75
26 MiMo-V2.6 Flash (High) 53.29 $0.04 17:56 OpenCode 17.83 17.5 7.81 10.15
27 GPT-6-Luna (High) 53.04 $0.01 03:51 Codex CLI 16.9 15.83 5.31 15
28 Muse Spark 1.3 (Max) 53.01 $0.01 03:42 OpenCode 15.6 18.33 4.58 14.5
29 Qwen 3.8 Flash (Max) 52.64 $0.04 08:38 OpenCode 16.35 17.67 5.42 13.2
30 Deepseek-V4-Pro-0813 (Max) 50.35 $0.05 09:06 OpenCode 17.3 16.5 4.9 11.65
31 Tencent Hy3 (High) 49.87 $0.05 06:35 OpenCode 16.8 14.33 2.29 16.45
32 Deepseek-V4-Pro-0813 (High) 48.34 $0.04 07:15 OpenCode 17.25 16.33 2.71 12.05
33 Qwen 3.8 27B (Xhigh) 47.76 $0.45 23:22 OpenCode 17.23 15.67 2.71 12.15
34 GLM-5.3-Flash (Max) 47.51 $0.02 08:54 OpenCode 17.2 17.17 4.69 8.45
35 Gemini-3.8-Flash (High) 47.09 N/A 02:56 Antigravity 14.8 16.67 5.42 10.2
36 LongCat-2.5-Preview 46.65 N/A 16:50 OpenCode 16.98 16.67 2.4 10.6
37 GPT-6-Luna (Medium) 43.74 $0.01 02:13 Codex CLI 16.33 14.17 3.54 9.7
38 Gemini-3.1-Pro 42.36 $0.52 03:04 OpenCode 12.55 15.33 4.38 10.1
39 Composer 2.5 41.39 $0.10 02:41 Cursor 16 15.83 3.96 5.6
40 Minimax M3 (High) 40.35 $0.21 06:32 OpenCode 15.3 15 2.1 7.95
Older LLMs hidden from table

These models are hidden to keep the leaderboard focused on current LLMs. Untick any of them to add it back to the table.

Support Leaderboard Updates

Running dozens of prompts on multiple projects with different LLMs is costly.
Help keep it updated: get Premium membership.


Recent Premium Tutorials


Explanations

  • Full methodology and scoring formulas explained in this article
  • Prices are calculated with API costs
  • When a model is tested via a subscription plan rather than the API, there's no per-prompt price to measure, so the avg cost column shows N/A for comparison

Updates

  • October 3rd: Added new project worth 10 points for finding bugs
  • September 30th: Added GPT-6.1-Sol
  • September 29th: Added Sonnet 5.5
  • September 28th: Added LongCat-2.5-Preview
  • September 27th: Added MiMo v2.6 Pro/Flash models
  • September 25th: Added Space Bunny stealth model
  • September 24th: Added GPT-6-Sol and GPT-6-Luna
  • September 23rd: Added Opus 5.5
  • September 22nd: Added Grok 4.7
  • September 15th: Added SWE-2 model by Devin
  • September 12th: Added new project React/Typescript code quality (20 points max)
  • September 11th: Added Deepseek v4.1 Flash model
  • September 9th: Added Qwen 3.8 Max (0902)
  • September 8th: Added Fable 5.1 and re-calculate scoring on Code Quality project
  • September 6th: Added GPT-6-Astra
  • September 3rd: Added Gemini-3.8-Flash
  • August 31st: New project evaluating CODE QUALITY (max 20 points)
  • August 29th: Added Qwen3.8-Flash model
  • August 27th: Added GLM-5.3-Flash (ex Ox Alpha) model
  • August 20th: Added Qwen 3.8 27B
  • August 18th: Added Gemini-3.7-Flash tested in Antigravity
  • August 15th: Added GLM-5.3 model
  • August 14th: Added DeepSeek-v4-Pro-0813 and Grok 4.6 models
  • August 12th: Added new evaluation project no.4 with Go language
  • August 7th: Added Muse Spark 1.2 and GPT-5.6-Terra High levels
  • August 4th: Added Qwen 3.8 Max and NEW 3rd Project Flutter/Dart
  • August 2nd: Added Max levels of GPT-5.6-Luna and Deepseek-v4-Flash
  • August 1st: Re-tested GPT-5.6-Luna/Terra and Deepseek-v4-Flash with new prices
  • July 27th: added Opus 5 High and GPT-5.6-Sol Medium
  • July 27th: new v2 version of the leaderboard - with only two (but more complex) projects.
  • July 26th: cleanup - removed older models: GPT-5.5, Kimi K2.6, Gemini 3.5 Flash, Sonnet 4.6
  • July 22nd: added Gemini 3.6 Flash and Gemini 3.5 Flash Lite
  • July 18th: added Kimi K3
  • July 11th: added GPT-5.6 Luna and Terra
  • July 9th: added Grok 4.5
  • July 8th: added Tencent Hy3
  • July 1st: added Sonnet 5
  • June 24th: added GPT-5.4-Mini and Gemini-3.5-Flash
  • June 23rd: added 5th benchmark project, removed Opus 4.7 and GLM-5.1, added GPT-5.4
  • June 17th: added GLM-5.2
  • June 13th: added Kimi-K2.7-Code
  • June 5th: added Deepseek v4 Flash and removed Minimax M2.7
  • June 4th: added Qwen 3.7 Plus
  • June 2nd: added Qwen 3.7 Max and Tested with column
  • June 2nd: added Minimax M3
  • May 30th: added Avg Cost for all LLMs (based on API pricing)
  • May 29th: added Claude Opus 4.8
  • May 24th: added 4th benchmark project - for React/TypeScript
  • May 20th: added Composer 2.5
Povilas Korop

Get Weekly AI Coding News

Sent every Wednesday. No spam, ever. Unsubscribe anytime.