Back to Homepage

LLM Coding Leaderboard

Last update: August 20th, 2026

Results are based on 4 different PHP/Flutter/Go projects.
Full methodology and scoring formulas explained in this article.

# Model Total points
(max 20)
Avg cost
per prompt
Avg time
per prompt
Video Tested with Points per project (max 5)
CSV Import (PHP) Offline Sync (PHP) Bank Feed (Dart/Flutter) Shipping Quotes (Go)
1 GPT-5.6-Sol (Medium) 18 $1.01 05:14 Codex CLI 5 4 4 5
2 Opus 5 (Medium) 17.75 $1.10 03:48 Claude Code 5 3.25 4.5 5
3 GPT-5.6-Luna (Max) 17.5 $0.10 14:18 Codex CLI 5 4.5 3 5
4 Opus 5 (High) 17.5 $1.65 06:33 Claude Code 5 3.5 4 5
5 Tencent Hy3 (High) 16.45 $0.05 06:35 OpenCode 3.2 3.75 4.5 5
6 GPT-5.6-Terra (Medium) 16.45 $0.20 02:44 Codex CLI 4.2 3.75 4 4.5
7 GPT-5.6-Luna (Xhigh) 15.75 $0.06 09:07 Codex CLI 4 4.75 2 5
8 Opus 4.8 (Medium) 15.45 $1.08 04:13 Claude Code 4.5 3.75 4 3.2
9 Kimi K3 14.95 $0.63 12:49 OpenCode 4.5 3.25 3.2 4
10 Grok 4.6 (High) 14.75 $0.77 09:19 Cursor 4 3.75 4.5 2.5
11 GLM-5.3 (High) 14.7 $0.19 04:58 OpenCode 3.5 3.2 4.5 3.5
12 GPT-5.6-Luna (High) 14 $0.04 06:18 Codex CLI 4 4.5 1.5 4
13 Kimi K2.7 Code 12.9 $0.42 08:21 OpenCode 3 4 3.5 2.4
14 Grok 4.5 12.8 $0.24 03:30 OpenCode 3.2 2.7 3.5 3.4
15 Deepseek-V4-Flash (Max) 12.6 $0.02 05:04 OpenCode 3.2 3.5 3.5 2.4
16 GPT-5.6-Luna (Medium) 12.5 $0.02 03:00 Codex CLI 4 4 1.5 3
17 Gemini-3.7-Flash (High) 12.45 N/A 02:01 Antigravity 3.5 3.75 2.7 2.5
18 Qwen 3.8 27B (Xhigh) 12.15 $0.45 23:22 OpenCode 4.2 2.75 3 2.2
19 Deepseek-V4-Pro-0813 (High) 12.05 $0.04 07:15 OpenCode 2.4 3.25 4 2.4
20 Sonnet 5 (Medium) 12 $0.79 03:29 Claude Code 4.5 1.9 2.4 3.2
21 Deepseek-V4-Pro-0813 (Max) 11.65 $0.05 09:06 OpenCode 1.9 3.25 4 2.5
22 Qwen 3.8 Max (High) 11.6 $0.45 10:21 OpenCode 2.2 2.7 4 2.7
23 Gemini-3.6-Flash (High) 11.45 $0.75 02:58 OpenCode 2.5 3.75 2.7 2.5
24 GLM-5.2 10.3 $0.43 07:59 OpenCode 2.7 3.5 2.7 1.4
25 Gemini-3.1-Pro 10.1 $0.52 03:04 OpenCode 1.9 2.5 3.5 2.2
26 Deepseek-V4-Flash (High) 8.2 $0.02 06:31 OpenCode 1.9 2.2 2.5 1.6
27 Minimax M3 7.95 $0.21 06:32 OpenCode 3.5 0.95 1.9 1.6
28 Composer 2.5 5.6 $0.10 02:41 Cursor 2.2 1.4 0.9 1.1
29 GPT-5.6-Luna (Low) 5.1 $0.01 01:36 Codex CLI 3 0.7 0.2 1.2

Support Leaderboard Updates

Running dozens of prompts on multiple projects with different LLMs is costly.
So, if you want to support my mission and help keep the leaderboard updated with new models/variants, subscribe to Premium membership of AI Coding Daily.


Recent Premium Tutorials


Explanations

  • Prices are calculated with API costs
  • GPT/Opus were tested only on Medium effort: it was enough to get top spots, High effort wasn't needed
  • When a model is tested via a subscription plan rather than the API, there's no per-prompt price to measure, so the avg cost column shows N/A for comparison

Updates

  • August 20th: Added Qwen 3.8 27B
  • August 18th: Added Gemini-3.7-Flash tested in Antigravity
  • August 15th: Added GLM-5.3 model
  • August 14th: Added DeepSeek-v4-Pro-0813 and Grok 4.6 models
  • August 12th: Added new evaluation project no.4 with Go language
  • August 7th: Added Muse Spark 1.2 and GPT-5.6-Terra High levels
  • August 4th: Added Qwen 3.8 Max and NEW 3rd Project Flutter/Dart
  • August 2nd: Added Max levels of GPT-5.6-Luna and Deepseek-v4-Flash
  • August 1st: Re-tested GPT-5.6-Luna/Terra and Deepseek-v4-Flash with new prices
  • July 27th: added Opus 5 High and GPT-5.6-Sol Medium
  • July 27th: new v2 version of the leaderboard - with only two (but more complex) projects.
  • July 26th: cleanup - removed older models: GPT-5.5, Kimi K2.6, Gemini 3.5 Flash, Sonnet 4.6
  • July 22nd: added Gemini 3.6 Flash and Gemini 3.5 Flash Lite
  • July 18th: added Kimi K3
  • July 11th: added GPT-5.6 Luna and Terra
  • July 9th: added Grok 4.5
  • July 8th: added Tencent Hy3
  • July 1st: added Sonnet 5
  • June 24th: added GPT-5.4-Mini and Gemini-3.5-Flash
  • June 23rd: added 5th benchmark project, removed Opus 4.7 and GLM-5.1, added GPT-5.4
  • June 17th: added GLM-5.2
  • June 13th: added Kimi-K2.7-Code
  • June 5th: added Deepseek v4 Flash and removed Minimax M2.7
  • June 4th: added Qwen 3.7 Plus
  • June 2nd: added Qwen 3.7 Max and Tested with column
  • June 2nd: added Minimax M3
  • May 30th: added Avg Cost for all LLMs (based on API pricing)
  • May 29th: added Claude Opus 4.8
  • May 24th: added 4th benchmark project - for React/TypeScript
  • May 20th: added Composer 2.5

Explanation / Methodology

Each prompt was launched 5 times on the same project, making it total of 25 points max.

For all evaluation tests passing, LLM got 1 point for the task. If at least one test failed, LLM got 0 points for that task.

So this above is the summary table.

I will continue testing models constantly - will come up with new tasks for evaluation, and will update when new LLMs are released.

Povilas Korop

Get Weekly AI Coding News

Sent every Wednesday. No spam, ever. Unsubscribe anytime.