Mentlio Route whitepaper
SWE-Bench Pro routing methodology and results
Mentlio Route retained 98.3% of the Always Opus result while reducing modeled spend by 21.7% on the same SWE-Bench Pro traces.

By Ahmet Demirbas
The routing question
Sending every coding task to Opus gives the strongest static result, but it also applies the highest model rate to routine work. Moving everything to Sonnet or Haiku costs less and lowers completion quality. A router should keep difficult tasks on Opus and move only the work that a less expensive model can handle.
Evaluation contract
We evaluated Opus, Sonnet, and Haiku on the same SWE-Bench Pro task set. Mentlio Route then selected a model for each prompt, and the replay applied that model's result to the routed task.
The same replay rule was used for the external router baselines. Spend applies the model-specific rates to each routed mix and indexes Always Opus at 100. This keeps the task outcomes fixed while the routing policy changes.
Results
| Policy | Accuracy | Spend |
|---|---|---|
| Opus | 69.8% | 100.0 |
| Mentlio Route | 68.6% | 78.3 |
| Prev. SoTA Router | 66.0% | 85.7 |
| vLLM-SR | 62.0% | 64.3 |
| Sonnet | 60.0% | 60.0 |
| Haiku | 50.8% | 20.0 |
Mentlio Route reached 68.6% accuracy at a spend index of 78.3. Always Opus reached 69.8% at an index of 100. Route therefore finished 1.2 points behind Opus and reduced modeled spend by 21.7%.
Route also finished above the cheaper static baselines: 8.6 points above Sonnet and 17.8 points above Haiku. The previous state-of-the-art router reached 66.0% accuracy at a higher spend index of 85.7.
Validity and limits
This design compares policies on identical traces, but model outputs vary between samples. Repeated evaluations can therefore move the result. The spend index is also a model-rate comparison rather than an invoice estimate; actual savings depend on token usage and contracted prices.
SWE-Bench Pro measures repository-level coding work. It does not represent every agent workload. Production use still requires a shadow replay on the customer's own prompts and an agreed quality floor.
Mentlio's impact
Mentlio makes the tradeoff visible before routing is enabled: the quality retained, the spend removed, and the tasks that still require the strongest model. The policy can then be monitored against a fixed quality floor.
On this benchmark, Route retained 98.3% of Opus accuracy at 78.3% of Opus spend. That is the useful operating point: keep the stronger model available without paying its rate on every task.
Test the policy on your workload.
Mentlio starts in shadow mode and measures quality and cost against your existing model traces.
