The four layers
Harness, execution, runtime and memory — what each owns, and the contracts that hold them together.
neo is four modules built against fixed contracts rather than one program. The harness runs the agent loop, execution owns the sandbox and the verifier, the runtime schedules and routes, and memory records what was learned and exposes it over MCP. The integration is the point: no single layer fixes a bug on its own.
| Layer | What it owns | Where |
|---|---|---|
| Harness | Planner, step agent, verifier gate, repo snapshot and diff, git output, rationale, resume | harness/ |
| Execution | Docker sandbox — a fresh container per command — plus stateless verify and flake detection | execution/ |
| Runtime | Process-per-task scheduler, checkpoint and resume, approval gate, adaptive router, cost ledger | runtime/ |
| Memory + MCP | tree-sitter code graph, SQLite decision store, the MCP server and the MCP client | memory/, mcp_server/ |
Each layer was built as a separate workstream against boundaries recorded in INTERFACES.md, so a layer depends on its neighbour's documented signature rather than on its internals.
The contracts between them
| Boundary | The contract | What crosses it |
|---|---|---|
| Harness to execution | execute_sandboxed(repo_path, command, timeout_s) | One command in one fresh container; an ExecutionResult back |
| Harness to execution | verify(repo_path, target_test, rerun_for_flake_check) | A VerificationResult for one repo state |
| Harness to runtime | call_model(messages, difficulty_hint, provider, model, api_key) | One model call, with an optional difficulty hint |
| Runtime to harness | run_task(task) -> TaskResult | One whole task, called concurrently with checkpoints around it |
| Harness to memory | logs/<task_id>/state.json | Plan, completed steps, files touched, decisions, repo path |
| Memory to anyone | Five MCP tools over stdio | Code-graph and decision queries, for neo or any MCP client |
Why four and not one
Each layer exists because the verifier gate demands it. The gate needs a test result the agent cannot influence, which is what the sandbox provides. It needs a pristine baseline to compare against, which is why the harness snapshots and diffs instead of editing in place. Running many gated tasks at once and surviving a hard kill mid-run is what the runtime's checkpoints are for. And a gate that refuses a plausible-looking fix produces information worth keeping, which is what the decision store holds.
- verify() is stateless — it evaluates one repo state and knows nothing about attempts
- baseline_passed is always false from verify(); only the harness knows the pristine outcome, so it fills that field itself
- run_task is the contract the scheduler is built around — everything else in the runtime exists to call it safely
- The sandbox raises when Docker is down rather than falling back to running unsandboxed