model-bench is an internal Python CLI that screens OpenRouter LLM models against a curated pack of coding and config tasks drawn from this repo's real commits, to identify budget-tier models that clear the quality bar for offloadable work.
Two run modes
- single-shot (
mode: single-shot, the default): the model emits a whole file or answer in one turn; a deterministic verifier or an LLM judge grades it. - agentic (
mode: agentic): the model is dropped into a real repo snapshot with file tools (list_dir/read_file/write_file/done) and edits the code itself over several turns. This is the primary contract: native tool-calling carries file content in API-serialized JSON, so the output-format noise that dominated single-shot is gone, and it measures agentic reliability and token/turn efficiency alongside raw capability.
SWE-bench-style real-monolith tasks
An agentic task graded by the repo's own tests works like SWE-bench:
snapshot:pins the parent of a real fix commit and lists the monolith paths to materialize intofixture/(bench snapshotusesgit archive, so fixtures are real repo state and re-generate deterministically).- The model explores that fixture and makes the change the prompt describes.
- The
pytestverifier drops the gold test from the fix commit onto the workdir (a hidden grader the model never sees) and runs it on the monolith venv. On the buggy snapshot it fails; a correct edit makes it pass (fail-to-pass).
For example, hikes-walkhighlands-dom-01 and hikes-walkhighlands-duration-01 are both
agentic tasks against the hikes doability model (DOM scraping and duration-aware doability,
respectively). The pack currently has 17 agentic and 3 single-shot tasks in total; tasks/
is the source of truth for the full, current list.
Setup
Two interpreters are involved:
The harness runs on your bare
python3and needspyyaml,pydantic,httpx(pluspytestfor the unit tests).The verifier venv runs real monolith code (fixture + gold tests) and must have the monolith runtime deps. Recreate it from the pinned list:
python3 -m venv ~/.cache/model-bench-venv ~/.cache/model-bench-venv/bin/pip install -r requirements-venv.txtThe
pytestverifier resolves this venv from$MODEL_BENCH_VENV(default~/.cache/model-bench-venv).
Providers: candidates vs the Claude ceiling
Each model in models.yaml has a provider:
openrouter(default,role: candidate) rents the model per-token and records real cost / turns / tokens. These are the models you would actually deploy, so their cost is the point.OPENROUTER_API_KEYmust be set to run any of them; calls are billed (cents per task).claude-code(role: anchor, the Claude models) runs through the localclaudeCLI under the Max subscription. It is a capability ceiling, not a cost-ranked competitor: free, so cost is 0, and it uses Claude Code's own agent harness (not the bench tool loop), so its turns/tokens are not comparable to candidates. The judge also runs this way. Anchors need theclaudeCLI on PATH and no OpenRouter key.
An anchors-only run (--model claude) needs no OPENROUTER_API_KEY at all.
In-cluster llama.cpp (self-hosted qwen)
bench run --base-url <url>/v1 points the OpenRouter client at any OpenAI-compatible
endpoint instead, skipping the API key and OpenRouter pricing (cost records as 0). Two
routes reach the in-cluster llama.cpp:
- on-cluster:
kubectl -n inference port-forward svc/inference 18080:8080and--base-url http://127.0.0.1:18080/v1; - off-cluster:
--base-url https://private.jomcgi.dev/llm/v1plus--header "CF-Access-Client-Id: ..." --header "CF-Access-Client-Secret: ...". Cloudflare Access authenticates on those two headers only and ignoresAuthorization.
The qwen/qwen3.8-27b entry in models.yaml is the self-hosted row: it is
status: experimental so a bare bench run never sends that slug to OpenRouter, and its
api_model is the alias llama.cpp actually serves. The comments on that entry are
canonical for how it is reached.
Result cells (billed output — kept out of the worktree)
Each run writes one JSON cell per (task, model) under a durable per-user cache dir,
~/.cache/model-bench/results by default (override with MODEL_BENCH_RESULTS or
--results). They are deliberately NOT inside the git worktree: results/ is gitignored,
so a git worktree remove would delete them and force a full paid re-run. The cache is
keyed on prompt + fixture + verifier + model + budget, so re-running skips unchanged
cells; only the committed reports/leaderboard.md and the page's leaderboard.json are
version-controlled. bench report --json-out writes the JSON; the committed copy the
public page reads lives at
projects/monolith/frontend/src/lib/public/llm-leaderboard/leaderboard.json.
Commands
python3 -m bench snapshot # materialize all task fixtures
python3 -m bench run # run every active (task, model) cell
python3 -m bench run --task worldcup-swing-settled-01 --model qwen3-coder-30b # one cell, cheap
python3 -m bench report # regenerate reports/leaderboard.md
python3 -m bench list # models and their status/role
python3 -m bench drop <id> --reason "..." # retire a model in models.yaml
python3 -m bench prune # delete result cells for retired models
The leaderboard uses a gate model. Each task carries a difficulty tier
(easy / standard / hard): easy + standard form the qualification floor. A model
may miss at most one floor task and stay viable (FLOOR_MISS_TOLERANCE in bench/cli.py,
so a single flaky miss does not exclude it); missing more disqualifies it. The hard tasks
differentiate the qualified. Among the qualified, ranking is hard-task pass, then cost.
It reads on two lenses. The self-host lens (hard-task pass, median tokens/turns, tool-use reliability) is model-intrinsic and carries over to local hardware. The cloud lens (median wall-time, cost, cost-per-solve) is the real time and money to rent the model via OpenRouter, versus the Claude anchor rows. Remote wall-time reflects a typical cloud request, not local GPU throughput.
Probing the live qwen lane
probe/ is a separate CLI that runs the same tasks through the monolith's agent-session
API on the in-cluster qwen and pi lane, recording wall time, diff, verifier result and a
SigNoz span breakdown. It measures the lane, not the model, and is the harness for the
#5051 efficiency loop. See probe/README.md.