Every number, and how it was measured.
Results are only worth the method behind them, so the method comes first. Each figure below links back to the file in the repository that records it.
How the ablation was run.
Adaptive routing versus always-expensive.
| Task set | Arm | Success | Calls | Tokens | Cost | Wall |
|---|---|---|---|---|---|---|
| 5 fixture bugs | always-expensive | 5/5 | 17 | 38,680 | $0.0528 | 575s |
| 5 fixture bugs | adaptive | 5/5 | 31 | 69,615 | $0.0237 | 300s |
| 16-task set | always-expensive | 16/16 | 81 | 138,526 | $0.1505 | 2717s |
| 16-task set | adaptive | 16/16 | 71 | 136,436 | $0.0581 | 812s |
| 5 real OSS repos | always-expensive | 2/5 | 71 | 329,438 | $0.3059 | 2992s |
| 5 real OSS repos | adaptive | 3/5 | 75 | 302,801 | $0.0730 | 581s |
The first version made things worse.
The v1 predictor scored raw message text. Harness prompts are large by construction — system templates plus injected file context — so every call saturated at “hard” and the adaptive arm degenerated to always-expensive: 17 of 17 calls on the expensive tier, costing about twice the baseline. That result drove the v2 redesign: score only the issue portion of the first user message, add a struggle-based escalation signal, and add router-side 429 backoff.
96–97% of ON-arm calls ran on the cheap tier, and success never dropped because of routing.
RESULTS.md:100-101
The cheap tier's speed (p50 22s vs the expensive tier's p95 309s) meant more fix attempts fit the same wall-clock budget.
RESULTS.md:117-120
Escalation never sticks: after any expensive call, routing drops back to cheap unless struggle persists.
RESULTS.md:112-114
Reliability under deliberate abuse.
Stress
45 tasks at concurrency 45, with 8 killed simultaneously mid-run.
Soak
RESULTS.md:139-148
Beyond the fixture set.
Five real OSS repos at pinned SHAs, one genuine bug introduced in each and encoded as a failing regression test. The adaptive arm finished 3/5 against always-expensive's 2/5, at 24% of the cost.
Where it fell short
Absolute success drops on unfamiliar repos. All three failing agents located their bug but ran out of turns before applying the edit — a model-capability limit, not a machinery one.
jaraco/pathREADME.md:102-106; CHANGELOG.md:60-64
python-semverREADME.md:108-110; CHANGELOG.md:65-66
more-itertoolsREADME.md:94-95; RESULTS.md:74
arrowREADME.md:94-95; RESULTS.md:74
inflectREADME.md:94-95; RESULTS.md:74
boltonsREADME.md:94-95; RESULTS.md:74
Repositories with runs still in flight are deliberately absent. Numbers appear here only once the repository records them. The full write-up lives in RESULTS.md.