Mentlio Route benchmark

Routing in the Fable Era: Frontier Quality Without the Frontier Tax

On the same SWE-Bench Pro replay, Mentlio Route kept 96.32% route sufficiency and reduced average model cost by 20.66% compared with Always Fable.

Ahmet Demirbas

By Ahmet Demirbas

July 20267 min readSWE-Bench Pro replay

Why Fable changes the baseline

Fable 5 raises the ceiling on coding performance, but it also raises the cost of sending every task to the strongest model. In this replay, Always Fable scored 80.4 at a relative cost of 10.00. Always Opus cut that cost in half, but its score fell to 69.8.

Routing is useful in the space between those two defaults. The goal is to reserve Fable for work that needs it and use a less expensive tier when the expected result remains strong.

Evaluation contract

Every policy is replayed against the same recorded model outcomes. For each task, the router selects a model tier and receives that tier's recorded result and cost. The model traces do not change between policies, so the comparison isolates the routing decision.

Sufficiency measures whether the chosen tier was strong enough for the task. Actual score measures the result produced by the selected tier. Cost is indexed to Always Fable at 10.00, which makes savings comparable without tying the result to one vendor contract.

Results

6872768010865ACTUAL SCOREAVERAGE RELATIVE COST · LOWER →Always Fable80.4 · 10.00Always Opus69.8 · 5.00Mentlio Route78.4 · 7.93NadirClaw (tuned)74.6 · 7.62
PolicySufficientCostScore
Always Fable100.00%10.0080.4
Mentlio Route96.32%7.9378.4
NadirClaw (tuned)91.32%7.6274.6
Always Opus86.84%5.0069.8

Mentlio Route scored 78.4 at a relative cost of 7.93. It finished 2.0 points behind Always Fable while reducing cost by 20.66%. Always Opus was cheaper, but it gave up 10.6 points of score.

Router comparison

The tuned NadirClaw configuration reduced cost by 23.79%, slightly more than Mentlio. At that operating point, it reached 91.32% sufficiency and a 74.6 score. Mentlio reached 96.32% sufficiency and a 78.4 score. The extra quality came with 3.13 percentage points less savings.

This is the practical routing tradeoff: the lowest bill is not automatically the best result. A production policy needs a quality floor, then should find the lowest cost that stays above it.

Validity and limits

This is a replay benchmark, not a live production trial. It compares routing policies on the same model results and task set. Absolute savings also depend on the customer's model prices and token mix.

The result covers software-engineering tasks from one benchmark family. Before automatic routing, the same policy should be evaluated in shadow mode on the customer's own traffic and quality requirements.

Mentlio's impact

Mentlio turns model choice into a policy that can be measured before it is enforced. Teams can set a quality floor, compare the routed result with an always-frontier baseline, and review when a stronger model was selected.

In this replay, that control preserved most of Fable's result while removing 20.66% of its modeled cost. The same process can be repeated on private workloads without exporting source code or raw prompts.

Test the policy on your workload.

Mentlio starts in shadow mode and measures quality and cost against your existing model traces.

Get a demo