Case study · October 7, 2026
RealityRouter beat GPT-6 Astra, DeepSeek V4.1 Flash and Jev Router
Ten randomly sampled DeepSWE tasks, four configurations, the same models available to each. RealityRouter solved 8 of 10 for about $9.
Ten randomly sampled DeepSWE tasks. Four configurations. The same models available to each.
EnlargeWe ran RealityRouter against a frontier model, a strong cheap model, and the router everyone is arguing about. Ten DeepSWE tasks, drawn at random with a published seed, same model pool for every configuration.
| Configuration | Full solves | F2P | P2P | Cost |
|---|---|---|---|---|
| RealityRouter | 8/10 | ~99.5% | 1.000 | ~$9 |
| DeepSeek V4.1 Flash | 7/10 | 99.2% | 1.000 | ~$4 |
| GPT-6 Astra | 6/10 | ~69.6% | 1.000 | ~$20 |
| Jev Router | 4/10 | ~88.5% | ~1.000 | ~$16 |
The frontier model cost 2.2× what we did and solved two tasks fewer. Jev Router solved half as many as we did, for 44% more.
The cheap model is the interesting one. DeepSeek V4.1 Flash solved 7 of 10 for four dollars — genuinely good, and cheaper than us. Then dasel: it passed 142 of 146 tests and still failed the task. Four tests short is a branch nobody can merge.
RealityRouter passed 146 of 146. It started cheap, exactly as DeepSeek did, and brought in a stronger model when the evidence said the cheap one would not close it out.
That is the whole argument. A cheap model's problem is not its average — its average is excellent. It is that you cannot tell in advance which task it will quietly miss. Routing is how you stop guessing.
Run the identical ten tasks yourself:
pier run -p deep-swe/tasks --agent mini-swe-agent --n-tasks 10 --sample-seed 0
Every task
The full breakdown, so the numbers above can be checked rather than taken on trust.
| Task | RealityRouter | DeepSeek V4.1 Flash | Jev Router | GPT-6 Astra |
|---|---|---|---|---|
| testem-bail-on-test-failure | 90/90 | 90/90 | 87/90 | 86/90 |
| effect-sse-httpapi-streaming | 46/47 | 46/47 | 45/47 | 0/47 |
| httpx-streaming-json-iteration | 108/108 | 108/108 | 108/108 | 108/108 |
| ts-pattern-match-each | 85/85 | 85/85 | 0/85 | 0/85 |
| python-statemachine-state-data-scoping | 70/72 | 70/72 | 69/72 | 0/72 |
| dasel-html-document-format | 146/146 | 142/146 | 144/146 | 146/146 |
| katex-multicolumn-array-spans | 94/94 | 94/94 | 92/94 | 94/94 |
| task-task-graph-export | 20/20 | 20/20 | 20/20 | 20/20 |
| tengo-callable-instance-isolation | 23/23 | 23/23 | 23/23 | 23/23 |
| sql-formatter-bigquery-pipe-formatting | 26/26 | 26/26 | 26/26 | 26/26 |
Method, and why the cheap model's variance is the whole point: the full write-up