The right model for every task, at the lowest price that keeps the quality.
Some models are far better at code than at maths. Some are a fiftieth of the price and just as good at the job in front of them. We work out which is which, per request, and send your work to the one that wins — while you keep a hard floor on quality that we are not allowed to cross.
The same six requests, routed
Every row below is generated by running the actual router when this page is built — not written by a marketer. The baseline is Claude Fable 5, the sort of model you might reasonably default to for everything.
| Request | Detected as | Routed to | Quality | Cost | vs baseline |
|---|---|---|---|---|---|
| Acknowledging a message | Simple chat | DeepSeek V4 ProDeepSeek | 85% | £0.000007 | −97% |
| Summarising a long document | Summarisation | Qwen3.7 PlusAlibaba | 90% | £0.001176 | −97% |
| Writing a new function | Code generation | Claude Opus 5Anthropic | 94% | £0.001492 | −53% |
| Debugging a stack trace | Code review & debugging | Claude Opus 5Anthropic | 93% | £0.00199 | −53% |
| A hard maths problem | Maths & logic | Claude Opus 5Anthropic | 100% | £0.000498 | −53% |
| Drafting a short story | Creative writing | Qwen3.8 MaxAlibaba | 96% | £0.000135 | −87% |
Quality is that model’s published benchmark score on the detected task type, as a percentage of the best model available for that task. Costs are estimates for a 1,024-token reply — a different model writes a different number of words, so the exact figure moves. The median saving across these six is 87%. We publish the whole method, including where the benchmark numbers come from and what we cannot prove yet.
Cheaper is only worth it
if the answer still lands
A router that quietly downgrades your output to save a penny is worse than no router. So the price dial in Progressive Labs is not allowed to touch quality at all.
The floor is a constraint, not a preference
You set a minimum — "never below 92% of the best model for this task". Models that fail it are removed from consideration before price is even looked at. The dial only ever chooses between survivors.
When we are unsure, we get more expensive
If the classifier is not confident about what kind of work this is, the floor rises automatically. An uncertain router should be a cautious one — misrouting an agentic coding task costs you a failed task, which dwarfs the tokens it saved.
You can always pin a model
Choose a specific model and that is exactly what runs, every time, with no second-guessing. We still show you afterwards what routing would have done, so the choice stays yours and stays informed.
Then we cut the tokens
themselves
Picking a cheaper model is only half of it. The other half is not paying for the same tokens twice.
Repeat answers cost a tenth
Ask the identical question twice at temperature zero and the second one is served from cache. It costs the upstream provider nothing, so we charge a flat 10% instead of full price rather than pretending we did the work.
Warm prompt caches, managed for you
Long system prompts and conversation history can be cached upstream at up to 99% off. We track which prefixes actually repeat and only pay the cache-write premium when the maths says it will earn its money back.
Routing that knows about the cache
Switching model throws away a warm cache. The router prices that in, so it will not move you to save 15% on paper while quietly losing a 90% discount.
One conversation, one model
Within a session we keep you on the model we picked, so the prefix stays warm. Consistency is usually worth more than chasing the cheapest option on every turn.
What you get
One account, one balance, in pounds.
Smart routing
Send your request to auto and we choose the model. You see which one ran, why, what it scored, and what it saved.
Compare
Two models, one prompt, one screen — or put Auto against your current default and see whether we actually beat it on your real work.
API
One OpenAI-compatible endpoint. Change the base URL, pass auto as the model, and keep the rest of your code exactly as it is.
Usage & billing
Every run itemised: which model, which task type, what it scored, what it cost, and what your baseline would have cost. Exportable as CSV.
Fair by construction
The things we decided up front so you don't have to take our word for them later.
Prepaid credit, never a surprise bill
You top up; we draw down. There is no invoice at the end of the month and no way to spend money you have not already put in.
Every charge is itemised
Each run writes a ledger entry with the running balance, so any charge reconciles back to the exact request that caused it.
Prompts are not stored
We keep a hash for abuse detection and nothing else. Your prompts and completions are not retained, not logged, and never used for training.
We show our working
Benchmark sources, dates, the normalisation formula and the limits of what we can prove are all published. A saving you cannot audit is a saving you should not believe.
Models we route between
Retail prices per million tokens in GBP. The full catalogue lists every model we carry.
Claude Fable 5
AnthropicAnthropic's most capable model. Top of SWE-bench Verified and top of the creative-writing arena — the one to reach for when being right matters more than what it costs.Input /M£9.18Output /M£45.90
Claude Opus 5
AnthropicComplex agentic coding and long-horizon work at half the price of Fable 5. Leads Terminal-Bench 2.1 and the maths arena.Input /M£4.59Output /M£22.95
Claude Sonnet 5
AnthropicNear-Opus coding quality at Sonnet money — 0.852 on SWE-bench Verified for a fifth of Fable 5. One of the best value-per-point models in the catalogue.Input /M£2.87Output /M£14.34
Qwen3.8 Max
AlibabaAlibaba's newest flagship — 2.4T-parameter MoE, natively multimodal, third in the creative-writing arena. Flat pricing across the whole 1M window, with no long-context surcharge.Input /M£1.91Output /M£5.74
GPT-5.6 Sol
OpenAIOpenAI's flagship. Leads Terminal-Bench 2 and tops the GPQA science leaderboard.Input /M£4.59Output /M£27.54
Gemini 3.6 Flash
GooglePunches far above its price on maths — fourth in the arena, level with models costing several times more.Input /M£1.43Output /M£7.17
DeepSeek V4 Pro
DeepSeekFrontier-adjacent quality at a tenth of frontier prices, and by far the most aggressive prompt-cache rate on the market — cached input costs under 1% of fresh.Input /M£0.466Output /M£0.932
MiniMax M3
MiniMaxStatistically level with Gemini 3.1 Pro on SWE-bench Verified at a fraction of the price. The clearest example of why routing pays.Input /M£0.643Output /M£2.57
Stop paying frontier prices
for everyday work.
Free to sign up. No card required to look around. Add credit only when you want to run something.