friday, september 18, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

When to let the tool pick the model for you

Every major AI tool now chooses a model on your behalf. Learn where that saves real money, where it quietly costs you, and what to keep on manual.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

Somewhere in the last year the model picker stopped being yours. Cursor’s Auto now has three modes [2]. OpenRouter rewrote its Auto router in August 2026, so a lightweight classifier reads your prompt and ranks models by what the market is actually spending on [4]. ChatGPT can add reasoning to an answer while you have Instant selected, and that may draw on your reasoning allowance [5]. The Gemini app, when a subscriber reaches their limit, carries the conversation on with Flash-Lite [6]. Something is now choosing a model on your behalf in most of the tools you open, and the setting you are looking at does not always describe what ran [5][6].

Most of the time that is fine and it saves you money. Occasionally it is the reason a piece of work came back flatter than last week when you changed nothing. The useful skill is not deciding whether routing is good. It is knowing which of your work you are willing to let a classifier gamble on, and which work you pin to a named model and leave pinned. This guide is for a solo operator or a small team using these tools daily on a paid plan. It is not for anyone building their own routing layer against raw APIs, where you want the vendors’ routing references and your own evaluation suite rather than a guide.

Three different things are now called routing

Untangling them matters, because only one of the three is a decision you make on purpose.

The first is a deliberate router. Cursor’s Auto has three modes, Cost, Balance and Intelligence [2], and Cursor’s own descriptions trade quality against token spend: Intelligence for frontier quality matching the most expensive models, Balance for the quality most people daily drive, Cost for reaching the highest available intelligence while optimising token spend [1]. OpenRouter’s version is the same idea with a dial exposed: you send model: 'openrouter/auto' and set a cost_tier of low, medium, high, xhigh or max, which filters candidates into a price band before selection, and the band is not a ceiling, so models cheaper than it are excluded along with models above it [3]. In both cases you chose to hand over the decision, and you set the budget it operates under.

The second is an effort dial inside one model, which people mistake for routing because it lives in the same dropdown. ChatGPT Business shows Instant, Medium, High and Extra High, and all four run GPT-5.6 Sol at different reasoning efforts, with Pro using GPT-5.6 Sol Pro or GPT-6 Pro as it becomes available [5]. Claude Code exposes effort as low, medium, high, xhigh and max, defaulting to high on most models [7]. Nothing is being routed. The same model is being asked to think for longer, and the cost moves accordingly.

The third is routing you did not ask for, and it is the one that surprises people. Gemini’s help page says that a subscriber who reaches their limit can continue the conversation with Flash-Lite [6]. Claude Code re-runs a request on a different model when a classifier flags certain content categories, and shows a notice in the transcript when it does [7]. It can also fall down a configured chain when the primary model is overloaded or unavailable, for the current turn only [7]. And ChatGPT’s own documentation is blunt about the cost side: automatic reasoning may use your reasoning allowance even while Instant remains selected [5]. None of that is a setting you turned on. It is the product protecting its own availability and its own limits, and it moves your outputs while it does so.

The saving is large and the quality cost is small but real

The published numbers are consistent in shape across vendors, which is the useful part.

Cursor launched its router on 22 July 2026, trained on more than 600,000 live requests and evaluated in an online A/B test across millions more [1]. Its own cost-per-commit figures put Fable 5 at $12.69, Intelligence mode at $6.76 and Balance mode at $4.63 [1]. OpenRouter published a straighter comparison because it benchmarked its new router against its old one: on MMLU Pro, the new default setting of cost_tier=low scored 85.2% at $140.93, against the old default’s 86.6% at $393.34 [4]. That is roughly 1.4 percentage points of accuracy for just over a third of the cost.

Read those two together and the trade is legible. You are not buying equivalent quality more cheaply. You are buying slightly less quality much more cheaply, on a distribution of tasks where most of them never needed the expensive model in the first place. For a summarisation job, a first draft, a rename, a formatting pass, 1.4 points is invisible. For the one client deliverable a week where being subtly wrong costs you a re-do and some credibility, it is not.

Worth knowing on the billing side: OpenRouter charges no extra fee for the router itself, and you pay the standard rate for whichever model is selected [3]. Cursor bills all Auto modes at the list price of the model each request is routed to, and a specific-model selection draws from the Other Models pool at that model’s API rate [2]. The router is not a product you buy. It is a decision about which meter runs.

A router optimises for the median request, not for yours

Look at how the selection is actually made and the limit becomes obvious.

OpenRouter’s classifier assigns each prompt one of about 30 fine-grained task types, then ranks candidate models using the community’s real spending over a trailing 7-day window, so aggregate usage rather than manual curation drives the choice [3][4]. Cursor’s router analyses each request on query, context, task complexity and domain, combined with what Cursor knows about each model’s behaviour, sending simple work to the most price-efficient models and more complex, long-horizon problems to frontier reasoning models [1].

Both are genuine advantages of scale. A classifier trained on hundreds of thousands of real requests [1] has seen edge cases your own rules never anticipated, and no small team has enough traffic of its own to build the equivalent. But notice what the machinery optimises. OpenRouter’s ranking is in part a record of what everyone else paid for this week [3][4]. Cursor’s router was trained on the live traffic of its own user base [1]. If your work sits at the edge of that distribution, a niche stack, a house style, regulated copy that has to be exactly right, then the router’s judgement about what counts as hard enough to deserve the expensive model is a judgement formed on somebody else’s work.

The way to find out is dull and it works. Run the router in its default mode alongside whatever fixed model you use now, on a week of genuine work rather than a set of test prompts, and pay attention only to where the two answers diverge. The divergence is the information. The vendor’s aggregate saving figure is a ceiling measured under its conditions, not a forecast for yours.

Find out what it actually picked

A router you cannot audit turns every bad output into an unsolvable question, because you cannot tell a weak prompt from a weak route.

Check what your tool reports before you rely on it. OpenRouter returns a model field in the response telling you which model answered, and it remembers the model a conversation landed on and prefers it on later turns, so a multi-turn thread does not lurch between models mid-task [3]. Claude Code names both the requested and the substituted model when an allowlist forces a substitution, and shows a transcript notice when a flagged request is re-run elsewhere [7]. Those are the good cases.

The weaker case is a chat interface that changes what it is doing without changing what it says. ChatGPT’s documentation states that automatic reasoning may use your reasoning allowance even while Instant remains selected [5], which means the setting in the dropdown is not a complete description of what ran. Gemini’s continuation on Flash-Lite once a subscriber hits a limit is stated on the help page [6], which is not where you are looking when the answer arrives. Neither vendor is being dishonest. Both mean that when you compare this week’s output to last week’s, the model may not be the constant you assumed.

The practical habit is small. When an output is worse than you expected, check what answered before you rewrite the prompt. If the tool will not tell you, that is itself a reason to pin a model for that class of work.

Keep a manual lane for the work that is expensive to get wrong

Every tool covered here still lets you take the decision back, and the point of leaving a router on is that you have decided in advance where it does not apply.

In Cursor you can still select a specific third-party model instead of Auto, with usage drawn from the Other Models pool at that model’s API rate [2]. The current list spans Claude Fable 5.1, Claude Opus 5 and Claude Sonnet 5, Gemini 3.1 Pro and Gemini 3.8 Flash, and GPT-5.6 Luna, Sol and Terra [2]. Claude Code has a default that follows your plan, Opus 5 on Max, Team Premium, Enterprise and the Anthropic API, Sonnet 5 on Pro and Team Standard, and an Enterprise admin can set an organisation default that the default option resolves to instead [7]. On OpenRouter, setting cost_tier to a higher band is the lighter-touch version of the same move: you keep the router and raise the floor [3].

Write the list of pinned categories down once, in whatever your team reads. Anything customer-facing on the day it ships. Anything you would have to explain if it were wrong. Anything where the cost of the model is trivial next to the cost of redoing the work. Everything else goes to the router, which is most of it.

Cost per request is the number that decides it, not the vendor’s percentage

Percentage savings are the wrong unit, because 60% of a number you have not measured is still a number you have not measured. Work out what one representative request costs on each end of your range instead.

The published API prices give you the spread. Claude Fable 5.1 is $10 per million input tokens and $50 per million output; Opus 5 is $5 and $25; Sonnet 5 is $2 and $10; Haiku 4.5 is $1 and $5 [8]. On a request of roughly 10,000 input tokens and 1,000 output tokens, that is about 7.5 cents on Opus 5 and about 1.5 cents on Haiku 4.5. Six cents. Multiply it by how many of those you actually send in a month and you get a real figure, and usually it is either large enough to justify twenty minutes of setup or small enough that you should stop thinking about it and pin the better model.

calculator
What routing is worth per month
$ / month

6 cents is the gap on a 10,000-in, 1,000-out request between Claude Opus 5 at $5/$25 per million tokens and Claude Haiku 4.5 at $1/$5 per million [8]. 22 working days. Computed in the page; nothing is sent anywhere.

checklist
Before you leave a router on
0 of 7 · saved in this browser only

What still goes wrong

The honest problem with a router is that its mistakes are less legible than yours. When you pick a model and the output is poor, you know the one variable that changed. When a classifier picks and the output is poor, you are debugging a decision you did not see, made on grounds that are only summarised in a blog post. Vendors have improved this at the API level, where OpenRouter returns the model it used [3] and Claude Code notices its own substitutions in the transcript [7], but the consumer chat interfaces are still the weak spot, and that is where most of this guide’s readers work.

The published figures also have a shelf life measured in weeks. Cursor’s numbers come from its own A/B tests against its own user base [1], OpenRouter’s benchmark compares its new router to its old one rather than to an independent baseline [4], and both were true on the day they were published. Model names churn faster than any of it. Everything named in this guide was read from the vendor’s own page on 5 September 2026, and a pinned model is a maintenance item, not a permanent decision. Whatever you pin, put a reminder somewhere to check it in a quarter.

Finally, a caveat about measuring at all. The side-by-side test works cleanly for tasks with a right answer and badly for tasks without one. If your work is drafting, naming, positioning or anything where you are the judge, you will find divergence and no way to score it, and you will end up making a taste decision with a spreadsheet open next to it. That is allowed. Just do not tell yourself the spreadsheet decided.

sources
  1. 01Cursor — Introducing Cursor Routercursor.com
  2. 02Cursor Docs — Modelscursor.com
  3. 03OpenRouter Docs — Auto Routeropenrouter.ai
  4. 04OpenRouter — Model Routing Powered by Wisdom of the Marketopenrouter.ai
  5. 05OpenAI Help Center — ChatGPT Business models and limitshelp.openai.com
  6. 06Gemini Apps Help — Gemini Apps limits & upgrades for Google AI subscriberssupport.google.com
  7. 07Claude Code Docs — Model configurationcode.claude.com
  8. 08Claude Platform Docs — Models overviewplatform.claude.com
next guide
Where your AI rules actually come from
10 min · verified 2026-09-05
related guides