Skip to content
Benchmarks

Every number, and how it was measured.

Results are only worth the method behind them, so the method comes first. Each figure below links back to the file in the repository that records it.

How the ablation was run.

Paired arms
The same task set runs twice: once with every call pinned to the expensive model, once with adaptive routing on. Nothing else differs between arms.
The full real stack
Scheduler subprocesses, the real harness loop, real model calls, and Docker-sandboxed pytest verification. No mocks anywhere in the measured path.
Per-call ledgers
Every model call appends model, tokens, cost and the routing hint to a JSONL ledger, so the totals are summed from records rather than estimated.
Verifier-gated success
A task counts as a success only when the target test passes AND the full suite shows no regressions. There is no partial credit.

Adaptive routing versus always-expensive.

Task setArmSuccessCallsTokensCostWall
5 fixture bugsalways-expensive5/51738,680$0.0528575s
5 fixture bugsadaptive5/53169,615$0.0237300s
16-task setalways-expensive16/1681138,526$0.15052717s
16-task setadaptive16/1671136,436$0.0581812s
5 real OSS reposalways-expensive2/571329,438$0.30592992s
5 real OSS reposadaptive3/575302,801$0.0730581s
Source: README.md:58-63; cross-checked RESULTS.md:80-84
summed across all three task sets
always-expensive
$0.5092
adaptive
$0.1548
ratio
3.29×

The first version made things worse.

The v1 predictor scored raw message text. Harness prompts are large by construction — system templates plus injected file context — so every call saturated at “hard” and the adaptive arm degenerated to always-expensive: 17 of 17 calls on the expensive tier, costing about twice the baseline. That result drove the v2 redesign: score only the issue portion of the first user message, add a struggle-based escalation signal, and add router-side 429 backoff.

  • 96–97% of ON-arm calls ran on the cheap tier, and success never dropped because of routing.

    RESULTS.md:100-101

  • The cheap tier's speed (p50 22s vs the expensive tier's p95 309s) meant more fix attempts fit the same wall-clock budget.

    RESULTS.md:117-120

  • Escalation never sticks: after any expensive call, routing drops back to cheap unless struggle persists.

    RESULTS.md:112-114

Reliability under deliberate abuse.

Stress

45 tasks at concurrency 45, with 8 killed simultaneously mid-run.

succeeded
45/45
genuine resumes
8/8
leaked containers
0

Soak

tasks
3,600
mid-run kills
120
checks passed
15/15
latency p95 drift
1.02x
artifacts per task
7.06 -> 7.07
leaked workers
0

RESULTS.md:139-148

Beyond the fixture set.

Five real OSS repos at pinned SHAs, one genuine bug introduced in each and encoded as a failing regression test. The adaptive arm finished 3/5 against always-expensive's 2/5, at 24% of the cost.

Where it fell short

Absolute success drops on unfamiliar repos. All three failing agents located their bug but ran out of turns before applying the edit — a model-capability limit, not a machinery one.

  • jaraco/path

    README.md:102-106; CHANGELOG.md:60-64

  • python-semver

    README.md:108-110; CHANGELOG.md:65-66

  • more-itertools

    README.md:94-95; RESULTS.md:74

  • arrow

    README.md:94-95; RESULTS.md:74

  • inflect

    README.md:94-95; RESULTS.md:74

  • boltons

    README.md:94-95; RESULTS.md:74

Repositories with runs still in flight are deliberately absent. Numbers appear here only once the repository records them. The full write-up lives in RESULTS.md.