Running across repositories
Fan a task set out through the scheduler, one supervised agent per bug.
`neo run-benchmark` takes a set of tasks and runs them through the scheduler, one process per task. Each task is a separate repository, a separate sandbox, and a separate log directory; the gate applies to each one independently.
Run a subset
neo run-benchmark --subset tasks.json --concurrency 4
# plumbing check: one bundled fixture task, no key needed
neo run-benchmark --subset smoke`--subset` is either the literal `smoke` or a path to a JSON file. `--concurrency` caps how many run at once and defaults to 10. The run ends with per-task outcomes and a summary line of success, failed, and error/timeout counts with total cost.
The subset file
[
{
"repo": "../more-itertools",
"issue": "chunked() drops the final partial chunk.",
"target_test": "tests/test_more.py::ChunkedTests"
},
{
"repo": "../python-semver",
"issue": "compare() orders prerelease versions incorrectly.",
"target_test": "tests/test_semver.py::test_compare",
"task_id": "semver-compare"
}
]The file must be a JSON list of objects. `repo` and `issue` are required and must be non-empty strings; `target_test`, `test_command`, `task_id` and a `config` object are optional. A malformed entry is a usage error — exit code 2, naming the entry index — rather than a traceback from somewhere inside the scheduler. `--model` and `--provider` on the command line apply to every task in the set.
Where the logs go
logs/
<task_id>/
state.json plan, completed steps, files touched, decisions
trace.jsonl every prompt, response, tool call, verify result
rationale.md the grounded paragraph
git.json branch + commit + PR description (verified fixes only)
<task_id>.runtime/
model_ledger.jsonl per call: model, tokens, cost, hint
checkpoint.json the resume point