Live evidence
Every perf:* label, measured — never claimed.
Each pull request touching agent/ is scored against
both a public repo set and a private, undisclosed one the author has never seen —
the label follows the worse of the two. Published automatically by the maintainer
bot, straight after every real score. No cherry-picking: every scored PR appears here,
including the ones that got closed.
Maintainer foresight (M7)
Did it predict what the maintainers actually did next?
The objective, independently-checkable half of the score, on the held-out target the agent is never tuned against — no trust in our judge required. From the most recently scored real pull request.
Optimization journey
One bar per pull request.
Bar height is the agent's composite score at that moment — the
worse of its public and private target results, same rule the perf:*
label itself uses. The first bar is the fixed anchor release; every bar after it is a
real, published PR score. The bar only moves when something real happened — no line
connects two points that aren't directly comparable.
Full record
Every merged or regression-blocked PR.
The bot's label, and the measured delta on each target, for every
agent/ PR that either merged (any outcome, including no measurable
change) or was closed because it measurably regressed the agent. A PR closed for
any other reason doesn't appear here — this table is evidence, not a scoreboard of
every review the bot has ever made.