Four layers, wired deliberately.
Each layer exists elsewhere in some form. What does not exist elsewhere is the four of them agreeing before anything is called done.
The loop, end to end.
harness/Harness
The agent loop itself: planning, step execution, and the verifier gate.
Bash-only, with a fresh session per step and the task constraints re-injected each time. The original repository is never touched — the harness snapshots it, works on the copy, and diffs the two on the host. On a verified fix it writes a branch, a commit, a PR description and a rationale grounded in the actual trace.
What it guarantees
- Success is claimed only when the target test passes and the suite shows no regressions
- Checkpoint and resume at step granularity, surviving hard kills
- Protected paths: the agent cannot reach .git or the test files
- RECALL pulls older detail back from the task trace on demand
What it does not do
- Failure classification is deliberately not implemented — the chosen novel mechanism was routing, not repair strategy
- The state machine in harness/state_machine.py is a designed contract, not wired into the running loop
execution/Execution
The Docker sandbox and the stateless verifier.
Every command gets a fresh container. Verification is a stateless evaluation of one repository state: the target test runs, then the full suite, with a three-valued outcome so a timeout is distinguishable from a failure. Flake detection reruns the target and flags differing outcomes.
What it guarantees
- read-only rootfs — the image layer is immutable
- --network none — no egress unless explicitly allowed
- --cap-drop ALL — every Linux capability dropped
- mem-limit — OOM-killed rather than starving the host
- pids-limit — fork bombs collapse at the cap
- fresh container — one --rm container per command
What it does not do
- The read-write bind mount lets a container write host disk unquota'd. There is no Docker primitive for bind-mount quotas, so it is inherent to the contract that lets the harness diff on the host
- When Docker is unavailable the sandbox raises rather than silently running unsandboxed
runtime/Runtime
Scheduling, checkpointing, the approval gate, and adaptive model routing.
One process per task, concurrency-capped, with wall-clock and hang supervision and a per-task crash budget. The router predicts difficulty per call from the issue text and from live struggle signals in the conversation tail, then routes to a cheap or expensive tier accordingly.
What it guarantees
- Proven at 45 concurrent tasks with 8 simultaneous mid-run hard kills
- Two-authority checkpointing: the harness owns progress, the runtime owns whether a relaunch is a resume
- Every model call lands in a per-task JSONL cost ledger
- Escalation never sticks — routing drops back to cheap unless struggle persists
What it does not do
- Budget caps are enforced at attempt granularity, so a cap can overshoot by one attempt
- Costs are computed from published price rates, not from bills
memory/, mcp_server/Memory + MCP
The code graph, the decision store, and the MCP surface.
A tree-sitter code graph and a SQLite decision store that auto-ingests every task's structured state. The planner queries the store for decisions recorded against the current repository before it plans, so a documented mistake is not repeated.
What it guarantees
- query_structure — tree-sitter code graph — functions, classes, calls, imports
- query_decisions — decision and pattern memory across tasks
- record_decision — write a decision back into the store
- task_status — structured state for one task
- list_repos — repos the graph has indexed
What it does not do
- The code graph is Python-only.
- Retrieval is keyword and structural, not semantic or embedding-based.
- Context is bounded by file and line counts — there is no tokenizer.
A client as well as a server.
Neo is also an MCP client. It consumes any external stdio MCP server, so the memory layer and outside tooling meet on the same protocol.
Adversarial result
Escape attempts against host mounts, the PID namespace, the Docker socket and cross-container networking all failed. A fork bomb collapsed at the pids limit, a 2 GB memory bomb was OOM-killed, and tmpfs hit ENOSPC at exactly the 256 MB cap.