Your first fix
Fix one bug against the bundled fixture repo, then read every artefact the run leaves behind.
The repository ships a fixture repo with one real bug in it: `mean()` in mathutil.py returns the sum instead of the arithmetic mean. It is the shortest honest path to a verified fix, because the bug, the test that catches it, and the suite around it all already exist.
Run one fix
export MY_KEY=... # your OpenAI-compatible router key
neo fix \
--repo cli/fixtures/smoke_repo \
--issue "mean() in mathutil.py returns the sum, not the average. Fix it so tests/test_mathutil.py::test_mean passes." \
--provider openai \
--model <model> \
--api-key $MY_KEY \
--api-base <base-url>The repository is snapshotted first, so the copy you pointed at is never edited. Artefacts land under `logs/<task_id>/`, with the runtime's own bookkeeping in the sibling `logs/<task_id>.runtime/`. Pass `--log-root` to put them somewhere else.
| Artefact | What it holds |
|---|---|
| rationale.md | One paragraph: what was wrong, what changed, how it ended |
| git.json | branch, commit_sha, commit_message, pr_description |
| state.json | Plan, completed steps, files touched, decisions |
| trace.jsonl | One JSON object per event: prompts, responses, tool calls, verify results |
| model_ledger.jsonl | Per call: model, tokens, cost, difficulty hint |
How to read rationale.md
The rationale is assembled from the trace, not written by a second model call. That is the point: it cannot drift from what happened, because there is no generation step between the trace and the paragraph. Nothing the trace does not support appears in it.
- The failing test the baseline run found, with a short excerpt of its error
- The files the fix touched, read from state.json
- The decisions the harness recorded while working
- A closing verdict, and how many attempts it took