Skip to content
Guides

Running across repositories

Fan a task set out through the scheduler, one supervised agent per bug.

`neo run-benchmark` takes a set of tasks and runs them through the scheduler, one process per task. Each task is a separate repository, a separate sandbox, and a separate log directory; the gate applies to each one independently.

Run a subset

neo run-benchmark --subset tasks.json --concurrency 4

# plumbing check: one bundled fixture task, no key needed
neo run-benchmark --subset smoke

`--subset` is either the literal `smoke` or a path to a JSON file. `--concurrency` caps how many run at once and defaults to 10. The run ends with per-task outcomes and a summary line of success, failed, and error/timeout counts with total cost.

The subset file

[
  {
    "repo": "../more-itertools",
    "issue": "chunked() drops the final partial chunk.",
    "target_test": "tests/test_more.py::ChunkedTests"
  },
  {
    "repo": "../python-semver",
    "issue": "compare() orders prerelease versions incorrectly.",
    "target_test": "tests/test_semver.py::test_compare",
    "task_id": "semver-compare"
  }
]

The file must be a JSON list of objects. `repo` and `issue` are required and must be non-empty strings; `target_test`, `test_command`, `task_id` and a `config` object are optional. A malformed entry is a usage error — exit code 2, naming the entry index — rather than a traceback from somewhere inside the scheduler. `--model` and `--provider` on the command line apply to every task in the set.

Where the logs go

logs/
  <task_id>/
    state.json       plan, completed steps, files touched, decisions
    trace.jsonl      every prompt, response, tool call, verify result
    rationale.md     the grounded paragraph
    git.json         branch + commit + PR description (verified fixes only)
  <task_id>.runtime/
    model_ledger.jsonl   per call: model, tokens, cost, hint
    checkpoint.json      the resume point