Decision memory
A decision store and a code graph the planner reads before it plans, and the ablation that measured what that changed.
The memory layer is two stores: a SQLite decision store holding what previous tasks concluded, and a tree-sitter code graph holding a repository's structure. Both are scoped per repository, and both are read before planning rather than after — which is the only way they can change a plan.
Before the planner call, run_task asks the store what was recorded against this repository and injects the matches as a `## Relevant past decisions` section of the planner prompt. It sits after `## Retrieved context` and before `## Constraints`, deliberately past the point where the difficulty predictor stops reading, so recalled decisions never shift routing.
- The query runs in-process against the store, not over the MCP wire — same process tree, no server round-trip
- Search is scoped by repo path, so decisions recorded against other repositories are not injected
- The store auto-ingests each task's state.json decisions, so the record accumulates without being written by hand
- It is best-effort: a missing module or an unreadable store degrades to "none recorded yet" plus a trace event, never a failed plan
The code graph
The graph indexes functions, classes, methods, imports and call relationships, and persists to disk, so structural context rides into prompts without re-reading files. It is also what coordinated-change detection walks to find the callers and importers a signature change would break.
What the ablation measured
Five fixture tasks were seeded with decisions genuinely mined from earlier runs, then run twice against the same pinned model with adaptive routing off, so the arms differed only in plan_with_memory. Both arms succeeded on every task in one attempt: memory did not change whether these tasks were solvable. It changed what they cost.
| Measure | Memory off | Memory on |
|---|---|---|
| Harness model calls | 35 | 27 |
| Tokens | 147,916 | 112,390 |
| Repeated past mistakes | 3 | 0 |
Scope limits
- The code graph indexes Python, JavaScript, and TypeScript
- Retrieval is keyword and structural, not semantic or embedding-based
- Context is bounded by file and line counts — there is no tokenizer
- Call edges are best-effort: Python is dynamic, so a bare call resolves to any known symbol of that name