After looking at my LLM benchmark table, many people ask how exactly I test models and what the evaluation criteria are. To avoid repeating it in every video, I published this article explaining my methodology.
This benchmark measures two things: whether a coding model produces maintainable, production-quality code, and how reliably it can finish production-shaped tasks in unfamiliar, partially implemented projects.
The whole benchmark on one page


Both halves weigh exactly the same. Code Quality is listed first on purpose: passing a task today matters less if the result is difficult or unsafe to extend tomorrow.
The one-line summary: the leaderboard rewards engineering quality and consistent, repeatable task completion — not a model's best result after retries or coaching.
The two formulas, up front
1. Code Quality — the raw 100-point score is scaled straight down to 20.
Code Quality points = Q / 5
What that looks like across the range:

Every raw point carries the same weight — five rubric points to one leaderboard point. All of the separation between models is created inside the rubric, not by the conversion.
2. Behavioral reliability — each attempt is scored by how many checks it fails.

Five attempts per project, so each project lands anywhere from 0 to 5, fractions included. Four projects → 20 points.

Part 1 — Code Quality: Community Equipment Library (20 points)
The model receives the shared prompt and Phase 1 prompt for a private Community Equipment Library built with Laravel. It must implement invited-member access, administrator equipment management, an equipment catalog with search and pagination, maintenance states, role-based authorization, and immediate enforcement when an account is deactivated.
The project is PHP/Laravel, but the rubric is language-agnostic in spirit: clear architecture, appropriate boundaries, meaningful tests, safe data modeling, access control, maintainability, useful error handling. Framework conventions are judged only as "does the model use its chosen stack correctly".
📄 The full scorecard — every criterion, every point, and real recorded judgments — is published here: Equipment Library Scorecard: how a judge model spends 100 points on a Laravel codebase
How one judgment happens

Anything byte-identical to the starter template is neutral — neither credited nor penalized. Scoring is absolute: each repository is judged against the rubric alone, never against the other submissions.
The judge is fixed so results stay comparable: gpt-5.6-sol via the Codex CLI at medium reasoning effort, one judgment per submission, roughly $1 per judgment.
Where the 100 points live


D1 + D2 alone are 33 points — effectively "does it ship, and did the model prove it?"
The five rules that keep the judging honest
1. One flaw, one deduction. Before scoring, the judge builds a ledger: each root cause gets an ID, exactly one owning criterion, the exact point loss, and
file:lineevidence. Every other affected criterion records "covered by that ID; no additional deduction" and keeps its credit. A scattered status label is charged once — not four times across design, idiom, Blade and duplication.
2. Names are never trusted. A test called "pagination keeps the search query" earns nothing unless it inspects a page link or requests page 2 with the query. A UI control earns credit only if its state scope and submitted fields actually work.
3. Volume is never rewarded. Assertion counts are descriptive, never a score. Padding a test with extra assertions gains nothing; what counts is which contractual behaviour it pins down.
4. Evidence beats suspicion. A suspected flaw with no code trace or reproduction goes into an
uncertaintieslist, not into a deduction. Missing a preferred pattern is not proof of broken behaviour, and valid framework alternatives are accepted rather than punished.
5. UI findings must state their evidence level — source trace, rendered HTML, or a real browser interaction. An Alpine
@clickattribute is not proof that a delete confirmation works.
What the deductions actually look like
Across 31 scored submissions, totals ranged from 62.75 to 92.75. In one representative set of five, all five passed the hidden behavioural probes — the app worked in every case. The 30-point spread was entirely about the code:

Grade bands for the raw total:

From 100 raw points to 20 leaderboard points
Let Q be the raw score:
Code Quality points = Q / 5
No curve, no clamping, no percentage conversion: the raw rubric total is divided by five. 100 raw is 20 leaderboard points, and each raw point is worth exactly 0.2.
Applied to those recorded totals:

Two submissions that both work — 92.75 and 62.75 raw — still land 6 leaderboard points apart. That gap comes entirely from the rubric.
Part 2 — Behavioral reliability: four projects × five attempts (20 points)

What the model starts with
Every project has a committed initial template: a small, runnable project with the existing structure, dependencies, schema or public interfaces, a deliberately incomplete or first-pass implementation, and a small public smoke test where appropriate.
- CSV Import — a Laravel contacts app with an existing
/api/contacts/importroute,Contactmodel, table, and first-pass importer. - Offline Sync — a Laravel notes app with
Note,SyncChangeandSyncMutationpersistence, plus a deliberately naive sync service. - Bank Feed — a Flutter app where
TransactionFeedhas its fixed constructor and data-source interfaces, but the widget itself is a stub. - Shipping Quotes — a Go package with the public carrier/request/result contract, helpers, and a sequential first-pass
QuoteService.
The model works inside the template. It does not build an application from scratch, and it may not change the public contract where the prompt says that contract is fixed.
The loop, per attempt

The five attempts are independent. No code, notes, or feedback carries over. This measures repeatable first-pass performance, not the best result after a round of coaching.
Prompts and evaluator tests stay separate
The model sees the prompt and the workspace. It never sees the hidden evaluator test files or the scoring map. That prevents optimizing against test implementation details.
The tests are behavioral: they exercise the public route, widget or package API and inspect observable results — stored records, JSON responses, rendered widgets, returned quotes, errors, ordering, resource behavior. Different internal designs pass equally, as long as they satisfy the contract.
Laravel: php artisan test .eval-tests --filter="(EvalTest|ArchTest)"
Flutter: flutter test .eval-tests --reporter=json --no-pub
Go: go test -json -race -count=1 ./evaluator_tests
The runner writes a machine-readable report and records the error message for every failed case. A timeout or a missing report counts as an unsuccessful attempt.
Checks are split into floor checks (basic capability) and signal checks (the discriminating ones), which is how I diagnose why an attempt failed.
Scoring an attempt

Offline Sync uses the more forgiving scale because it has the largest check count and the most ways to fail partially. An attempt that produces no usable change, or fails verification entirely, is scored as failing every relevant check → 0.
A project's score is the sum of its five attempts, so 3.7/5 means the attempts were on average close to clean — not that two of five failed outright. The detailed passed/total, the raw failing-check count, and the floor/signal split are all retained; they are exactly what drives the point value, not a coarser pass/fail layered on top.
Worked example: one CSV Import attempt
The model imports valid rows but rejects the entire mixed file instead of salvaging the valid row:
$ php artisan test .eval-tests --filter="(EvalTest|ArchTest)"
FAILED ... reports skipped rows and row-level errors for a mixed import
Failed asserting that 0 is greater than or equal to 2.
Tests: 29, Passed: 28, Failed: 1
One failing check → 0.5 of 1 point. Not clean, but not zeroed out either. The 28/29 detail is kept for diagnosis; the failing-check count drives the score.
The four projects in detail
1. CSV Import — untrusted input
The task: harden an existing contact CSV importer without changing its public API. The fixed contract is POST /api/contacts/import, upload under the file multipart field, columns name,email,phone. Email is the natural key. The importer must be idempotent, salvage valid rows from mixed input, return a structured JSON reconciliation summary, handle malformed or hostile uploads cleanly, and survive production-sized files.
What it really tests: whether the model thinks beyond explode(',', ...) and one happy-path demo file. Quoted commas and newlines, UTF-8 BOMs, CRLF files, invalid bytes, invalid rows, duplicate emails, large files, injection-shaped content.
One evaluator, concretely — a mixed file with one valid row, one invalid email, one missing email:
name,email,phone
Alice,alice@example.com,111
Bad,not-an-email,222
NoEmail,,333
It expects a success response, ≥1 imported, ≥2 skipped, and row-level errors. All-or-nothing validation fails it:
FAILED ... reports skipped rows and row-level errors for a mixed import
Failed asserting that 0 is greater than or equal to 2.
Coverage — 29 checks: 1 floor (import a clean file) · 27 signal (structured responses, duplicate prevention, parsing and encoding, row validation, partial acceptance, batching, large inputs, bad uploads, security, DB invariants) · 1 optional memory-bound check on a ~20,000-row file.
2. Offline Sync — stateful, and the hardest to fake
The task: harden POST /api/sync for clients editing notes on multiple devices while offline. The API takes a device_id, a cursor, and ordered create/update/delete mutations, and must keep returning results, changes and next_cursor.
The rules: the server is authoritative. Mutation IDs give idempotency within a device, versions detect stale edits, conflicts must not overwrite newer state, deletes must produce tombstones, mixed batches must salvage valid siblings, pull-only requests must return an accurate ordered change feed, and malformed input must produce clean reconciliation responses rather than 500s.
One evaluator, concretely: submit the same create mutation twice, then pull from another device. The note must exist once, the retry must report applied or duplicate, and the observer must see one canonical change. An implementation that applies retries twice fails:
FAILED ... retrying a create mutation does not apply it twice
Expected collection to have count 1. Got 2.
Other tests replay conflicts, reuse mutation IDs with different contents, apply stale updates and deletes, verify tombstones, check cursor boundaries, and simulate a failed change-log write to prove canonical state is never left half-committed.
Coverage — 41 checks: 3 floor (create, update, delete correctly) · 38 signal (retries, batch idempotency, device scoping, conflicts, tombstones, cursor delivery, patch semantics, malformed and mixed input, ordering, transactions, production-sized batches).
A solution can look perfect for one request and still collapse when requests are retried, reordered, replayed from another device, or combined into one batch.
3. Bank Feed — async UI that has to survive real data
The task: implement the Flutter TransactionFeed widget without changing its constructor, data-source interface, or required widget keys. The source returns pages of untrusted raw maps: duplicated, out of order, malformed, wrong currency, or slow.
What it must do: show valid records only, count skipped ones, parse integer amounts, normalize timestamps with the supplied display offset, sort newest first, group by calendar day, compute the signed total, show loading/empty/error states, retry failed pages, paginate without duplicate requests, stay responsive with years of history and very long descriptions, and update correctly when rebuilt with changed inputs.
One evaluator, concretely: five unusable records plus one valid one. Only txn-ok may render, and feed-skipped must read 5 skipped. A widget that renders malformed records fails:
Expected: [txn-ok]
Actual: [txn-no-identity, txn-42, txn-bad-amount, ...]
Another good one: a later page fails, but already-loaded rows and the current total must stay visible while an inline feed-page-error and retry control appear.
Coverage — 48 checks: 27 floor (rendered feed, parsing, ordering, grouping, totals, initial errors, retries, pagination, layout safety) · 21 signal (malformed records, deduplication, cross-page ordering, skipped counts, duplicate requests, lazy list construction, partial failures, in-place widget updates).
Tests run on a phone-sized viewport with large pages, long text, emoji, right-to-left text and large amounts — a logically correct widget that cannot render reliably does not get full credit.
4. Shipping Quotes — concurrency under hostile dependencies
The task: harden the Go QuoteService without changing the exported contract in contract.go. It aggregates quotes from multiple third-party carriers, all treated as untrusted: a carrier may be slow, panic, fail, return malformed data, ignore cancellation, or return duplicate services.
What it must do: normalize and validate the request; run carrier calls concurrently with a configurable bound; honor per-carrier and overall timeouts; preserve valid sibling quotes; reconcile failures deterministically; return partial success when possible. A successful non-empty result may be cached, equivalent requests must share cache entries, concurrent identical misses must coalesce, and callers sharing a flight must cancel independently.
One evaluator, concretely: 12 carriers with MaxConcurrency: 3. The observed peak of active calls must be ≤ 3 while still proving calls overlap. Sequential or unbounded implementations fail:
bad concurrency peak 12, err=<nil>
Timeout tests require a fast carrier's quote to survive a stuck sibling; cache tests verify returned slices cannot mutate the cached result or another caller's result.
Coverage — 19 test groups: 1 floor (carrier fan-out, usable quote) · 18 signal (validation, normalization, quote reconciliation, deterministic ordering, panic and nil-carrier isolation, ownership, concurrency, timeouts, cancellation, cache lifecycle, cache keys, in-flight coalescing, production-sized carrier sets).
What the final score means
High score: the model produces well-designed code and repeatedly delivers complete solutions that survive both the happy path and adversarial checks.
Low score: not necessarily "can't write code" — more often a maintainability, testing, security or reliability gap that a demo would never reveal.
The 50/50 weighting keeps either side from hiding the other:
Elegant code cannot compensate for features that do not work. Passing tests cannot compensate for a weak, unsafe or unmaintainable implementation.
One successful behavioral attempt may be luck; five is a dependable agent. And on the code-quality half, five submissions all passed the hidden behavioural probes and still landed between 62.75 and 92.75 raw — between 12.55 and 18.55 leaderboard points.
If you want to see exactly where every one of those points went, criterion by criterion and file:line by file:line:
👉 Equipment Library Scorecard