vanguarstew

Live evidence

Every perf:* label, measured — never claimed.

Each pull request touching agent/ is scored against both a public repo set and a private, undisclosed one the author has never seen — the label follows the worse of the two. Published automatically by the maintainer bot, straight after every real score. No cherry-picking: every scored PR appears here, including the ones that got closed.

Maintainer foresight (M7)

Did it predict what the maintainers actually did next?

The objective, independently-checkable half of the score, on the held-out target the agent is never tuned against — no trust in our judge required. From the most recently scored real pull request.

Optimization journey

One bar per pull request.

Bar height is the agent's composite score at that moment — the worse of its public and private target results, same rule the perf:* label itself uses. The first bar is the fixed anchor release; every bar after it is a real, published PR score. The bar only moves when something real happened — no line connects two points that aren't directly comparable.

Full record

Every merged or regression-blocked PR.

The bot's label, and the measured delta on each target, for every agent/ PR that either merged (any outcome, including no measurable change) or was closed because it measurably regressed the agent. A PR closed for any other reason doesn't appear here — this table is evidence, not a scoreboard of every review the bot has ever made.