Back to Articles
Tutorials

LLM Coding Leaderboard: My Methodology and Scoring Formulas

August 21, 2026
16 min read

After looking at my LLM benchmark table, many people ask how exactly I test models and what the evaluation criteria are. To avoid repeating it in every video, I published this article explaining my methodology.

This benchmark measures two things: whether a coding model produces maintainable, production-quality code, and how reliably it can finish production-shaped tasks in unfamiliar, partially implemented projects.


The whole benchmark on one page

The benchmark is 40 points total: Code Quality is worth 20 points (one build, judged on a 100-point rubric, then scaled to 20 — "Is this code worth keeping?") and Behavioral reliability is worth 20 points (4 projects x 5 fresh attempts = 20 independent attempts — "Does it work, every time?")

Comparison of the two halves. Code Quality runs one Laravel build (Community Equipment Library), 1 attempt, graded by a blinded judge model against a fixed 100-point rubric across 8 dimensions and 49 criteria with every deduction cited file:line, max 20 points. Behavioral reliability runs CSV Import, Offline Sync, Bank Feed and Shipping Quotes, 5 attempts per project (20 total), graded by automated deterministic tests on the number of mapped evaluator checks failed, max 20 points (1 per attempt).

Both halves weigh exactly the same. Code Quality is listed first on purpose: passing a task today matters less if the result is difficult or unsafe to extend tomorrow.

The one-line summary: the leaderboard rewards engineering quality and consistent, repeatable task completion — not a model's best result after retries or coaching.


The two formulas, up front

1. Code Quality — the raw 100-point score is scaled straight down to 20.

Code Quality points = Q / 5

What that looks like across the range:

Raw score to leaderboard points, Q divided by 5: raw 100 gives 20.00 (perfect), raw 90 gives 18.00, raw 85 gives 17.00 (the Strong band starts here), raw 80 gives 16.00, raw 70 gives 14.00, raw 60 gives 12.00, raw 50 gives 10.00 and raw 40 gives 8.00.

Every raw point carries the same weight — five rubric points to one leaderboard point. All of the separation between models is created inside the rubric, not by the conversion.

2. Behavioral reliability — each attempt is scored by how many checks it fails.

Points per attempt by failing checks. Most projects: 0 failing checks scores 1.0, 1 scores 0.5, 2 scores 0.2, 3 or more scores 0. Offline Sync: 0 scores 1.0, 1 scores 0.75, 2 scores 0.5, 3 scores 0.2, 4 or more scores 0.

Five attempts per project, so each project lands anywhere from 0 to 5, fractions included. Four projects → 20 points.

Total score equals Q divided by 5 (0 to 20 points) plus CSV Import (0 to 5) plus Offline Sync (0 to 5) plus Bank Feed (0 to 5) plus Shipping Quotes (0 to 5).


Part 1 — Code Quality: Community Equipment Library (20 points)

The model receives the shared prompt and Phase 1 prompt for a private Community Equipment Library built with Laravel. It must implement invited-member access, administrator equipment management, an equipment catalog with search and pagination, maintenance states, role-based authorization, and immediate enforcement when an account is deactivated.

The project is PHP/Laravel, but the rubric is language-agnostic in spirit: clear architecture, appropriate boundaries, meaningful tests, safe data modeling, access control, maintainability, useful error handling. Framework conventions are judged only as "does the model use its chosen stack correctly".

📄 The full scorecard — every criterion, every point, and real recorded judgments — is published here: Equipment Library Scorecard: how a judge model spends 100 points on a Laravel codebase

How one judgment happens

Four steps. 1: the model gets the spec plus a fixed Laravel starter and writes the catalog, roles, search, deactivation and its own tests. 2: the harness runs and records the gates — Pint, PHPStan, the test suite, composer test, plus hidden behavioural probes. 3: the judge is blinded and sees only the model-authored diff and that evidence packet: no model name, no other submission, no ranking. 4: the judge scores one criterion at a time, with a deduction ledger so a flaw is charged exactly once and a file:line citation for every point lost.

Anything byte-identical to the starter template is neutral — neither credited nor penalized. Scoring is absolute: each repository is judged against the rubric alone, never against the other submissions.

The judge is fixed so results stay comparable: gpt-5.6-sol via the Codex CLI at medium reasoning effort, one judgment per submission, roughly $1 per judgment.

Where the 100 points live

Point distribution across the eight rubric dimensions: D1 Verification gate 15, D2 Test quality 18 (the biggest slice), D3 Design and boundaries 14, D4 Schema and model integrity 12, D5 Laravel/PHP idiom 12, D6 Duplication and dead code 11, D7 Access control 10, D8 UX and error surfacing 8.

The eight dimensions and the question each really asks. D1 Verification gate, 15 points: purely mechanical — Pint, PHPStan, the test run, composer test; no judgment, no arguing. D2 Model-authored test quality, 18 points: not how many tests, but which contractual behaviours they actually pin down. D3 Design and boundaries, 14 points: thin controllers, one validation boundary, one home for roles and statuses, middleware seams. D4 Schema and model integrity, 12 points: additive reversible migrations, indexes a real query uses, allow-listed writes, useful factories. D5 Laravel/PHP idiom, 12 points: types, correct framework APIs, named routes, translated strings, presentation-only Blade. D6 Duplication and dead code, 11 points: starts at 11 and only goes down — copy-pasted forms, unused methods, misleading names. D7 Access control and activity, 10 points: deactivation must bite on the next request, on every protected route, settings included. D8 UX and error surfacing, 8 points: validation errors, flash messages, empty states, retained search, a delete confirmation that works.

D1 + D2 alone are 33 points — effectively "does it ship, and did the model prove it?"

The five rules that keep the judging honest

1. One flaw, one deduction. Before scoring, the judge builds a ledger: each root cause gets an ID, exactly one owning criterion, the exact point loss, and file:line evidence. Every other affected criterion records "covered by that ID; no additional deduction" and keeps its credit. A scattered status label is charged once — not four times across design, idiom, Blade and duplication.

2. Names are never trusted. A test called "pagination keeps the search query" earns nothing unless it inspects a page link or requests page 2 with the query. A UI control earns credit only if its state scope and submitted fields actually work.

3. Volume is never rewarded. Assertion counts are descriptive, never a score. Padding a test with extra assertions gains nothing; what counts is which contractual behaviour it pins down.

4. Evidence beats suspicion. A suspected flaw with no code trace or reproduction goes into an uncertainties list, not into a deduction. Missing a preferred pattern is not proof of broken behaviour, and valid framework alternatives are accepted rather than punished.

5. UI findings must state their evidence level — source trace, rendered HTML, or a real browser interaction. An Alpine @click attribute is not proof that a delete confirmation works.

What the deductions actually look like

Across 31 scored submissions, totals ranged from 62.75 to 92.75. In one representative set of five, all five passed the hidden behavioural probes — the app worked in every case. The 30-point spread was entirely about the code:

Six representative deductions. Minus 9.0 on D1.2 plus D1.4: ten PHPStan errors — the scale is a cliff, not a slope (0 errors gives 6 points, 1-2 gives 4, 3-9 gives 2, 10+ gives 0), and a red composer test took the other 3. Minus 7.0 on D1.3 plus D1.4: the submission shipped red because of its own test — random factory data, then an assumption that Equipment::first() lands on page one. Minus 3.0 on D4.1: role and is_active inserted into the original users migration, so a fresh database passes but an existing deployment never gets those columns. Minus 2.0 on D4.2: five indexes, one useful query — over-indexing is scored as harshly as under-indexing. Minus 1.5 on D7.1: activity refreshed on the equipment routes, but the dashboard and settings routes sat outside the middleware, so a deactivated member kept browsing their profile. Minus 0.75 on D2.5: a search fixture matching the query in both name and category, so a green test proved nothing about which branch fired.

Grade bands for the raw total:

Grade bands: 85 or above is Strong, 70 to 84.99 is Solid, 55 to 69.99 is Mixed, 40 to 54.99 is Weak, below 40 is Poor.

From 100 raw points to 20 leaderboard points

Let Q be the raw score:

Code Quality points = Q / 5

No curve, no clamping, no percentage conversion: the raw rubric total is divided by five. 100 raw is 20 leaderboard points, and each raw point is worth exactly 0.2.

Applied to those recorded totals:

The conversion applied to recorded totals: raw 92.75 (Strong) gives 18.55 of 20 leaderboard points; raw 85.25 (Strong) gives 17.05 of 20; raw 76.50 (Solid) gives 15.30 of 20; raw 65.375 (Mixed) gives 13.075 of 20; raw 62.75 (Mixed) gives 12.55 of 20.

Two submissions that both work — 92.75 and 62.75 raw — still land 6 leaderboard points apart. That gap comes entirely from the rubric.


Part 2 — Behavioral reliability: four projects × five attempts (20 points)

The four projects. 1: CSV Import, untrusted CSV input, validation, idempotency and batching, Laravel/PHP, 29 mapped checks. 2: Offline Sync, multi-device state, conflicts, retries and tombstones, Laravel/PHP, 41 mapped checks. 3: Bank Feed, async UI state, pagination, parsing and performance, Flutter/Dart, 48 mapped checks. 4: Shipping Quotes, concurrency, cancellation, caching and partial failure, Go, 19 mapped checks.

What the model starts with

Every project has a committed initial template: a small, runnable project with the existing structure, dependencies, schema or public interfaces, a deliberately incomplete or first-pass implementation, and a small public smoke test where appropriate.

  • CSV Import — a Laravel contacts app with an existing /api/contacts/import route, Contact model, table, and first-pass importer.
  • Offline Sync — a Laravel notes app with Note, SyncChange and SyncMutation persistence, plus a deliberately naive sync service.
  • Bank Feed — a Flutter app where TransactionFeed has its fixed constructor and data-source interfaces, but the widget itself is a stub.
  • Shipping Quotes — a Go package with the public carrier/request/result contract, helpers, and a sequential first-pass QuoteService.

The model works inside the template. It does not build an application from scratch, and it may not change the public contract where the prompt says that contract is fixed.

The loop, per attempt

The per-attempt loop: fresh workspace from template, then install pinned dependencies and initialize the app, then hand the agent the prompt and workspace (the agent never sees the evaluator tests), then the agent inspects, edits and verifies within its session limits, then the harness copies in the hidden evaluator tests and runs them, then every pass and fail is recorded and the workspace is discarded.

The five attempts are independent. No code, notes, or feedback carries over. This measures repeatable first-pass performance, not the best result after a round of coaching.

Prompts and evaluator tests stay separate

The model sees the prompt and the workspace. It never sees the hidden evaluator test files or the scoring map. That prevents optimizing against test implementation details.

The tests are behavioral: they exercise the public route, widget or package API and inspect observable results — stored records, JSON responses, rendered widgets, returned quotes, errors, ordering, resource behavior. Different internal designs pass equally, as long as they satisfy the contract.

Laravel:  php artisan test .eval-tests --filter="(EvalTest|ArchTest)"
Flutter:  flutter test .eval-tests --reporter=json --no-pub
Go:       go test -json -race -count=1 ./evaluator_tests

The runner writes a machine-readable report and records the error message for every failed case. A timeout or a missing report counts as an unsuccessful attempt.

Checks are split into floor checks (basic capability) and signal checks (the discriminating ones), which is how I diagnose why an attempt failed.

Scoring an attempt

Scoring one attempt. With 0 failing checks an attempt scores 1.0 in most projects and 1.0 in Offline Sync; 1 failing check scores 0.5 and 0.75; 2 failing checks score 0.2 and 0.5; 3 failing checks score 0 and 0.2; 4 or more failing checks score 0 in both.

Offline Sync uses the more forgiving scale because it has the largest check count and the most ways to fail partially. An attempt that produces no usable change, or fails verification entirely, is scored as failing every relevant check → 0.

A project's score is the sum of its five attempts, so 3.7/5 means the attempts were on average close to clean — not that two of five failed outright. The detailed passed/total, the raw failing-check count, and the floor/signal split are all retained; they are exactly what drives the point value, not a coarser pass/fail layered on top.


Worked example: one CSV Import attempt

The model imports valid rows but rejects the entire mixed file instead of salvaging the valid row:

$ php artisan test .eval-tests --filter="(EvalTest|ArchTest)"

FAILED ... reports skipped rows and row-level errors for a mixed import
Failed asserting that 0 is greater than or equal to 2.

Tests: 29, Passed: 28, Failed: 1

One failing check → 0.5 of 1 point. Not clean, but not zeroed out either. The 28/29 detail is kept for diagnosis; the failing-check count drives the score.


The four projects in detail

1. CSV Import — untrusted input

The task: harden an existing contact CSV importer without changing its public API. The fixed contract is POST /api/contacts/import, upload under the file multipart field, columns name,email,phone. Email is the natural key. The importer must be idempotent, salvage valid rows from mixed input, return a structured JSON reconciliation summary, handle malformed or hostile uploads cleanly, and survive production-sized files.

What it really tests: whether the model thinks beyond explode(',', ...) and one happy-path demo file. Quoted commas and newlines, UTF-8 BOMs, CRLF files, invalid bytes, invalid rows, duplicate emails, large files, injection-shaped content.

One evaluator, concretely — a mixed file with one valid row, one invalid email, one missing email:

name,email,phone
Alice,alice@example.com,111
Bad,not-an-email,222
NoEmail,,333

It expects a success response, ≥1 imported, ≥2 skipped, and row-level errors. All-or-nothing validation fails it:

FAILED ... reports skipped rows and row-level errors for a mixed import
Failed asserting that 0 is greater than or equal to 2.

Coverage — 29 checks: 1 floor (import a clean file) · 27 signal (structured responses, duplicate prevention, parsing and encoding, row validation, partial acceptance, batching, large inputs, bad uploads, security, DB invariants) · 1 optional memory-bound check on a ~20,000-row file.

2. Offline Sync — stateful, and the hardest to fake

The task: harden POST /api/sync for clients editing notes on multiple devices while offline. The API takes a device_id, a cursor, and ordered create/update/delete mutations, and must keep returning results, changes and next_cursor.

The rules: the server is authoritative. Mutation IDs give idempotency within a device, versions detect stale edits, conflicts must not overwrite newer state, deletes must produce tombstones, mixed batches must salvage valid siblings, pull-only requests must return an accurate ordered change feed, and malformed input must produce clean reconciliation responses rather than 500s.

One evaluator, concretely: submit the same create mutation twice, then pull from another device. The note must exist once, the retry must report applied or duplicate, and the observer must see one canonical change. An implementation that applies retries twice fails:

FAILED ... retrying a create mutation does not apply it twice
Expected collection to have count 1. Got 2.

Other tests replay conflicts, reuse mutation IDs with different contents, apply stale updates and deletes, verify tombstones, check cursor boundaries, and simulate a failed change-log write to prove canonical state is never left half-committed.

Coverage — 41 checks: 3 floor (create, update, delete correctly) · 38 signal (retries, batch idempotency, device scoping, conflicts, tombstones, cursor delivery, patch semantics, malformed and mixed input, ordering, transactions, production-sized batches).

A solution can look perfect for one request and still collapse when requests are retried, reordered, replayed from another device, or combined into one batch.

3. Bank Feed — async UI that has to survive real data

The task: implement the Flutter TransactionFeed widget without changing its constructor, data-source interface, or required widget keys. The source returns pages of untrusted raw maps: duplicated, out of order, malformed, wrong currency, or slow.

What it must do: show valid records only, count skipped ones, parse integer amounts, normalize timestamps with the supplied display offset, sort newest first, group by calendar day, compute the signed total, show loading/empty/error states, retry failed pages, paginate without duplicate requests, stay responsive with years of history and very long descriptions, and update correctly when rebuilt with changed inputs.

One evaluator, concretely: five unusable records plus one valid one. Only txn-ok may render, and feed-skipped must read 5 skipped. A widget that renders malformed records fails:

Expected: [txn-ok]
Actual:   [txn-no-identity, txn-42, txn-bad-amount, ...]

Another good one: a later page fails, but already-loaded rows and the current total must stay visible while an inline feed-page-error and retry control appear.

Coverage — 48 checks: 27 floor (rendered feed, parsing, ordering, grouping, totals, initial errors, retries, pagination, layout safety) · 21 signal (malformed records, deduplication, cross-page ordering, skipped counts, duplicate requests, lazy list construction, partial failures, in-place widget updates).

Tests run on a phone-sized viewport with large pages, long text, emoji, right-to-left text and large amounts — a logically correct widget that cannot render reliably does not get full credit.

4. Shipping Quotes — concurrency under hostile dependencies

The task: harden the Go QuoteService without changing the exported contract in contract.go. It aggregates quotes from multiple third-party carriers, all treated as untrusted: a carrier may be slow, panic, fail, return malformed data, ignore cancellation, or return duplicate services.

What it must do: normalize and validate the request; run carrier calls concurrently with a configurable bound; honor per-carrier and overall timeouts; preserve valid sibling quotes; reconcile failures deterministically; return partial success when possible. A successful non-empty result may be cached, equivalent requests must share cache entries, concurrent identical misses must coalesce, and callers sharing a flight must cancel independently.

One evaluator, concretely: 12 carriers with MaxConcurrency: 3. The observed peak of active calls must be ≤ 3 while still proving calls overlap. Sequential or unbounded implementations fail:

bad concurrency peak 12, err=<nil>

Timeout tests require a fast carrier's quote to survive a stuck sibling; cache tests verify returned slices cannot mutate the cached result or another caller's result.

Coverage — 19 test groups: 1 floor (carrier fan-out, usable quote) · 18 signal (validation, normalization, quote reconciliation, deterministic ordering, panic and nil-carrier isolation, ownership, concurrency, timeouts, cancellation, cache lifecycle, cache keys, in-flight coalescing, production-sized carrier sets).


What the final score means

High score: the model produces well-designed code and repeatedly delivers complete solutions that survive both the happy path and adversarial checks.

Low score: not necessarily "can't write code" — more often a maintainability, testing, security or reliability gap that a demo would never reveal.

The 50/50 weighting keeps either side from hiding the other:

Elegant code cannot compensate for features that do not work. Passing tests cannot compensate for a weak, unsafe or unmaintainable implementation.

One successful behavioral attempt may be luck; five is a dependable agent. And on the code-quality half, five submissions all passed the hidden behavioural probes and still landed between 62.75 and 92.75 raw — between 12.55 and 18.55 leaderboard points.

If you want to see exactly where every one of those points went, criterion by criterion and file:line by file:line: 👉 Equipment Library Scorecard

Share this article

Povilas Korop

Get Weekly AI Coding News

Sent every Wednesday. No spam, ever. Unsubscribe anytime.