Skip to content
Concepts

The four layers

Harness, execution, runtime and memory — what each owns, and the contracts that hold them together.

neo is four modules built against fixed contracts rather than one program. The harness runs the agent loop, execution owns the sandbox and the verifier, the runtime schedules and routes, and memory records what was learned and exposes it over MCP. The integration is the point: no single layer fixes a bug on its own.

LayerWhat it ownsWhere
HarnessPlanner, step agent, verifier gate, repo snapshot and diff, git output, rationale, resumeharness/
ExecutionDocker sandbox — a fresh container per command — plus stateless verify and flake detectionexecution/
RuntimeProcess-per-task scheduler, checkpoint and resume, approval gate, adaptive router, cost ledgerruntime/
Memory + MCPtree-sitter code graph, SQLite decision store, the MCP server and the MCP clientmemory/, mcp_server/

Each layer was built as a separate workstream against boundaries recorded in INTERFACES.md, so a layer depends on its neighbour's documented signature rather than on its internals.

The contracts between them

BoundaryThe contractWhat crosses it
Harness to executionexecute_sandboxed(repo_path, command, timeout_s)One command in one fresh container; an ExecutionResult back
Harness to executionverify(repo_path, target_test, rerun_for_flake_check)A VerificationResult for one repo state
Harness to runtimecall_model(messages, difficulty_hint, provider, model, api_key)One model call, with an optional difficulty hint
Runtime to harnessrun_task(task) -> TaskResultOne whole task, called concurrently with checkpoints around it
Harness to memorylogs/<task_id>/state.jsonPlan, completed steps, files touched, decisions, repo path
Memory to anyoneFive MCP tools over stdioCode-graph and decision queries, for neo or any MCP client

Why four and not one

Each layer exists because the verifier gate demands it. The gate needs a test result the agent cannot influence, which is what the sandbox provides. It needs a pristine baseline to compare against, which is why the harness snapshots and diffs instead of editing in place. Running many gated tasks at once and surviving a hard kill mid-run is what the runtime's checkpoints are for. And a gate that refuses a plausible-looking fix produces information worth keeping, which is what the decision store holds.

  • verify() is stateless — it evaluates one repo state and knows nothing about attempts
  • baseline_passed is always false from verify(); only the harness knows the pristine outcome, so it fills that field itself
  • run_task is the contract the scheduler is built around — everything else in the runtime exists to call it safely
  • The sandbox raises when Docker is down rather than falling back to running unsandboxed