Skip to content
About

A harness that refuses to claim success.

Most coding agents finish when the model says they have finished. That is a claim about the work, produced by the thing that did the work. neo replaces it with a test.

The thesis.

An agent that grades its own homework will always pass. The only way to know whether a fix is real is to run the tests that were already there, in an environment the agent cannot reach into, and to treat the result as final.

That constraint shapes everything else. Because success must be provable, the sandbox has to be sealed. Because a run can be killed mid-flight, progress has to be checkpointed. Because the loop makes many calls and most of them are easy, routing by predicted difficulty becomes worth doing. The four layers are not a feature list — each exists because the gate demands it.

The same standard applies to this project's own claims. Every number on this site is traceable to a file in the repository, every caveat travels with the figure it qualifies, and the results that came out badly are still here.

The first version of the router made things worse — twice the cost for no gain. That result is on the benchmarks page, because a project that reports only its wins has told you nothing about its losses.

the v1 ablation, kept

What it cannot do.

Stated plainly and up front, because this is the part that decides whether the rest is useful to you.

No SWE-bench numbers
Not run. Benchmark-grade claims need paid tiers, multiple repetitions and confidence intervals, and that work is deferred. Anyone quoting a SWE-bench score for neo is quoting something that does not exist.
Structural retrieval is Python-complete, JS/TS-conservative
The code graph parses Python fully and scans .js/.jsx/.mjs/.cjs/.ts/.tsx conservatively. Symbol extraction is exact for Python functions, classes, and methods; for JavaScript and TypeScript it is name-based and intentionally over-approximates, and a file in any other language can only be claimed whole-file. This entry previously claimed a single-language code graph, which stopped being true when the JS/TS grammars were added.
No semantic retrieval
Retrieval is keyword and structural. There are no embeddings and no vector store.
No real tokenizer
Context is bounded by a token budget, but the budget is metered with a conservative heuristic (a configurable characters-per-token estimate, default 4) rather than a real tokenizer for the target model. The bound is therefore approximate. This entry used to claim there was no token budget at all, which stopped being true when the context compiler landed.
Directional sample sizes
The ablations run n=5 and n=16 at one repetition each. The cost ratios are large enough to be robust; success-rate differences at that n are noise.
Proxy pricing
Neither endpoint bills for usage, so costs use published rates for comparable model classes. The delta between arms is a price-model delta, not an invoice.
One transport
The MCP surface is stdio only, and there is no web UI beyond the read-only dashboard.

Built with.

  • Python 3.10
  • litellm
  • Docker
  • tree-sitter
  • MCP Python SDK
  • argparse
  • SQLite
  • pytest
  • stdlib dashboard

Open source. The full source, the results write-up, and the changelog are all in the repository — including the logs layout that every figure on this site is derived from.