alden

Evals

Prompt and model changes are measured against recorded public PRs before they ship.

Running

pnpm eval                                  # prints what a run would cost; calls nothing
pnpm eval --yes                            # runs every recorded PR through the model review
pnpm eval --yes --model opus --effort medium
pnpm eval --yes --only django              # one recorded PR
pnpm eval --yes --clones ~/src             # with the code graph, from clones in ~/src/<name>
pnpm eval --no-llm                         # the checks and code graph alone: free, no key needed
pnpm eval --no-llm --only reg              # just the regression cases

With --yes, the eval uses your configured provider and spends against your caps and ledger like any review. Results are written to eval/results/<time>-<model>.json, which is gitignored.

--clones <dir> gives recorded PRs the code graph when their repo is cloned at <dir>/<name> (for example ~/src/django), as a real review from a clone would.

What it measures

Without any labels:

Regressions: did the briefing point at what broke?

The phase 1 test is whether a briefing flags the code that later broke. The reg-* fixtures are PRs in all four languages that a later PR had to fix (Next.js for TypeScript, Django for Python, the GitHub CLI and Prometheus for Go, Laravel for PHP). Each one's expectation names the files the fix changed, and the eval reports, overall and by language:

pnpm regressions:find owner/repo finds more candidates: merged PRs that mention a regression, traced back through #123, PR links or commit SHAs to the PR that caused it, keeping pairs where the fix changed the original's code. Pick the clear ones, add them to scripts/record-fixtures.ts, record them with pnpm fixtures:record --only <name>, and add their expectations.

The first run, checks only without clones (30 Sep 2026), pointed at 2 of 29 broken files; with Laravel's clone, the graph added one more. The deterministic checks look for tests, migrations, config and secrets, not at the logic that broke, so this is the baseline the model has to improve on.

With labels, from eval/expectations.json:

When Alden misses something that mattered in real use, add it to expectations.json so the next change is measured against it.

The recorded PRs

pnpm fixtures:record records GitHub responses (PR details, descriptions and diffs) for a pinned list of public PRs into packages/core/src/github/__fixtures__. The PR numbers are pinned so results stay comparable over time. The same fixtures drive the unit tests.