A harness that refuses to claim success.
Most coding agents finish when the model says they have finished. That is a claim about the work, produced by the thing that did the work. neo replaces it with a test.
The thesis.
An agent that grades its own homework will always pass. The only way to know whether a fix is real is to run the tests that were already there, in an environment the agent cannot reach into, and to treat the result as final.
That constraint shapes everything else. Because success must be provable, the sandbox has to be sealed. Because a run can be killed mid-flight, progress has to be checkpointed. Because the loop makes many calls and most of them are easy, routing by predicted difficulty becomes worth doing. The four layers are not a feature list — each exists because the gate demands it.
The same standard applies to this project's own claims. Every number on this site is traceable to a file in the repository, every caveat travels with the figure it qualifies, and the results that came out badly are still here.
The first version of the router made things worse — twice the cost for no gain. That result is on the benchmarks page, because a project that reports only its wins has told you nothing about its losses.
What it cannot do.
Stated plainly and up front, because this is the part that decides whether the rest is useful to you.
Built with.
- Python 3.10
- litellm
- Docker
- tree-sitter
- MCP Python SDK
- argparse
- SQLite
- pytest
- stdlib dashboard
Open source. The full source, the results write-up, and the changelog are all in the repository — including the logs layout that every figure on this site is derived from.