Case study · October 7, 2026

RealityRouter beat GPT-6 Astra, DeepSeek V4.1 Flash and Jev Router

Ten randomly sampled DeepSWE tasks, four configurations, the same models available to each. RealityRouter solved 8 of 10 for about $9.

Ten randomly sampled DeepSWE tasks. Four configurations. The same models available to each.

Tasks fully solved on DeepSWE, out of ten. RealityRouter 8, DeepSeek V4.1 Flash 7, GPT-6 Astra 6, Jev Router 4.Enlarge

We ran RealityRouter against a frontier model, a strong cheap model, and the router everyone is arguing about. Ten DeepSWE tasks, drawn at random with a published seed, same model pool for every configuration.

ConfigurationFull solvesF2PP2PCost
RealityRouter8/10~99.5%1.000~$9
DeepSeek V4.1 Flash7/1099.2%1.000~$4
GPT-6 Astra6/10~69.6%1.000~$20
Jev Router4/10~88.5%~1.000~$16

The frontier model cost 2.2× what we did and solved two tasks fewer. Jev Router solved half as many as we did, for 44% more.

The cheap model is the interesting one. DeepSeek V4.1 Flash solved 7 of 10 for four dollars — genuinely good, and cheaper than us. Then dasel: it passed 142 of 146 tests and still failed the task. Four tests short is a branch nobody can merge.

RealityRouter passed 146 of 146. It started cheap, exactly as DeepSeek did, and brought in a stronger model when the evidence said the cheap one would not close it out.

That is the whole argument. A cheap model's problem is not its average — its average is excellent. It is that you cannot tell in advance which task it will quietly miss. Routing is how you stop guessing.

Run the identical ten tasks yourself:

pier run -p deep-swe/tasks --agent mini-swe-agent --n-tasks 10 --sample-seed 0

Every task

The full breakdown, so the numbers above can be checked rather than taken on trust.

TaskRealityRouterDeepSeek V4.1 FlashJev RouterGPT-6 Astra
testem-bail-on-test-failure90/9090/9087/9086/90
effect-sse-httpapi-streaming46/4746/4745/470/47
httpx-streaming-json-iteration108/108108/108108/108108/108
ts-pattern-match-each85/8585/850/850/85
python-statemachine-state-data-scoping70/7270/7269/720/72
dasel-html-document-format146/146142/146144/146146/146
katex-multicolumn-array-spans94/9494/9492/9494/94
task-task-graph-export20/2020/2020/2020/20
tengo-callable-instance-isolation23/2323/2323/2323/23
sql-formatter-bigquery-pipe-formatting26/2626/2626/2626/26

Method, and why the cheap model's variance is the whole point: the full write-up

→ realityrouter.dev