Grok 4.5 Coding Benchmark
How Grok 4.5 ranks on the LLM Coding Leaderboard, shown next to the three models directly above and below it.
- Grok 4.5 — place #14, 12.8 points (max 20). Last evaluated on August 11, 2026, with OpenCode.
| # | Model | Total points (max 20) |
Avg cost per prompt |
Avg time per prompt |
Video | Tested with | Points per project (max 5) | |||
|---|---|---|---|---|---|---|---|---|---|---|
| CSV Import (PHP) | Offline Sync (PHP) | Bank Feed (Dart/Flutter) | Shipping Quotes (Go) | |||||||
| 11 | GLM-5.3 (High) | 14.7 | $0.19 | 04:58 | OpenCode | 3.5 | 3.2 | 4.5 | 3.5 | |
| 12 | GPT-5.6-Luna (High) | 14 | $0.04 | 06:18 | Codex CLI | 4 | 4.5 | 1.5 | 4 | |
| 13 | Kimi K2.7 Code | 12.9 | $0.42 | 08:21 | OpenCode | 3 | 4 | 3.5 | 2.4 | |
| 14 | Grok 4.5 | 12.8 | $0.24 | 03:30 | OpenCode | 3.2 | 2.7 | 3.5 | 3.4 | |
| 15 | Deepseek-V4-Flash (Max) | 12.6 | $0.02 | 05:04 | OpenCode | 3.2 | 3.5 | 3.5 | 2.4 | |
| 16 | GPT-5.6-Luna (Medium) | 12.5 | $0.02 | 03:00 | Codex CLI | 4 | 4 | 1.5 | 3 | |
| 17 | Gemini-3.7-Flash (High) | 12.45 | N/A | 02:01 | Antigravity | 3.5 | 3.75 | 2.7 | 2.5 | |
Tutorials about Grok 4.5
Article
· Aug 10, 2026
I Tried Using Grok 4.5 in OpenCode vs Cursor
Video
· Jul 9, 2026
I Tested NEW Grok 4.5 for Coding. Wow. Just Wow.
How to read this
- Each score measures a model-and-harness configuration, not the model in isolation.
- Prices are calculated with API costs. Models tested on a subscription plan show N/A.
- Full methodology and scoring formulas are explained in this article.
See every tested model side by side on the full LLM Coding Leaderboard.