Back to the full leaderboard

Omen Alpha Coding Benchmark

How Omen Alpha ranks on the LLM Coding Leaderboard, shown next to the three models directly above and below it, with the top three models included for reference.

  • Last evaluated: 14 hours ago (September 4, 2026), with OpenCode.
# Model Total points
(max 20)
Avg cost
per prompt
Avg time
per prompt
Video Tested with Points per project (max 5)
CSV Import (PHP) Offline Sync (PHP) Bank Feed (Dart/Flutter) Shipping Quotes (Go)
1 GPT-5.6-Sol (High) 34.72 $1.18 10:15 Codex CLI 5 4.5 5 5
2 Opus 5 (High) 33 $1.65 06:33 Claude Code 5 3.5 4 5
3 Opus 5 (Medium) 32.97 $1.10 03:48 Claude Code 5 3.25 4.5 5
⋯ 8 places not shown ⋯
12 GPT-5.6-Luna (High) 24.86 $0.04 06:18 Codex CLI 4 4.5 1.5 4
13 Grok 4.6 (High) 23.65 $0.77 09:19 Cursor 4 3.75 4.5 2.5
14 Deepseek-V4-Flash (Max) 23.18 $0.02 05:04 OpenCode 3.2 3.5 3.5 2.4
15 Omen Alpha (High) 23.14 $0.03 01:51 OpenCode 4 3.5 2.7 3
16 Qwen 3.8 Flash (Max) 22.94 $0.04 08:38 OpenCode 3.5 3.5 2.2 4
17 Kimi K2.7 Code 22.78 $0.42 08:21 OpenCode 3 4 3.5 2.4
18 Tencent Hy3 (High) 22.11 $0.05 06:35 OpenCode 3.2 3.75 4.5 5

How to read this

  • Each score measures a model-and-harness configuration, not the model in isolation.
  • Prices are calculated with API costs. Models tested on a subscription plan show N/A.
  • Full methodology and scoring formulas are explained in this article.

See every tested model side by side on the full LLM Coding Leaderboard.

Povilas Korop

Get Weekly AI Coding News

Sent every Wednesday. No spam, ever. Unsubscribe anytime.