Kimi K3 Coding Benchmark
How Kimi K3 ranks on the LLM Coding Leaderboard, shown next to the three models directly above and below it, with the top three models included for reference.
- Last evaluated: 1 week ago (September 29, 2026), with OpenCode.
| # | Model | Total points (max 80) |
Avg cost per prompt |
Avg time per prompt |
Video | Tested with | Points per project | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
Laravel Code Quality
(max 20) |
React-TS Code Quality
(max 20) |
Bug Finding Test
(max 20) |
CSV Import (PHP) | Offline Sync (PHP) | Bank Feed (Dart / Flutter) | Shipping Quotes (Go) | |||||||
| 1 | GPT-6.1-Sol (High) | 68.1 | $0.22 | 07:26 | Codex CLI | 19.1 | 19 | 10 | 5 | 5 | 5 | 5 | |
| 2 | Opus 5.5 (High) | 67.41 | $0.79 | 03:10 | Claude Code | 18.5 | 19.33 | 9.58 | 5 | 5 | 5 | 5 | |
| 3 | GPT-6.1-Sol (Medium) | 67.09 | $0.14 | 03:53 | Codex CLI | 18.63 | 18.67 | 9.79 | 5 | 5 | 5 | 5 | |
| ⋯ 10 places not shown ⋯ | |||||||||||||
| 14 | GPT-6-Luna (Max) | 59.07 | $0.03 | 11:03 | Codex CLI | 18.12 | 17.17 | 7.08 | 5 | 4 | 3.2 | 4.5 | |
| 15 | Space Bunny (Max) | 57.19 | N/A | 19:19 | OpenCode | 17.61 | 16.5 | 4.58 | 5 | 4.5 | 4 | 5 | |
| 16 | SWE-2 (Max) | 57.07 | N/A | 11:59 | Devin | 17.5 | 17.83 | 7.29 | 4 | 2.45 | 4 | 4 | |
| 17 | Kimi K3 (High) | 56.89 | $0.63 | 12:49 | OpenCode | 18.4 | 17.5 | 6.04 | 4.5 | 3.25 | 3.2 | 4 | |
| 18 | Grok 4.6 (High) | 56.71 | $0.77 | 09:19 | Cursor | 17.4 | 18 | 6.56 | 4 | 3.75 | 4.5 | 2.5 | |
| 19 | GPT-6-Luna (Xhigh) | 56.4 | $0.03 | 10:08 | Codex CLI | 17.02 | 17.17 | 6.46 | 4.5 | 3.75 | 3 | 4.5 | |
| 20 | GPT-6-Sol (Medium) | 56.32 | $0.21 | 03:26 | Codex CLI | 17.17 | 16 | 6.25 | 2.9 | 5 | 5 | 4 | |
Tutorials about Kimi K3
Video
· Aug 13, 2026
I Tried Kimi Code with Kimi K3: Deep-Dive Review
PREMIUM
Article
· Jul 22, 2026
I Tested Kimi K3 vs Fable / Sol / Opus on 3 Coding Projects
How to read this
- Each score measures a model-and-harness configuration, not the model in isolation.
- Prices are calculated with API costs. Models tested on a subscription plan show N/A.
- Full methodology and scoring formulas are explained in this article.
See every tested model side by side on the full LLM Coding Leaderboard.