saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

How to tell whether a new model is worth switching to

Turn every model release into a short, repeatable decision: check the price and the retirement date, run your own tasks, and switch only when your numbers move.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

A new model lands with a chart. The chart has your current model on the left, the new one on the right, and a set of benchmark names you have never run. Somewhere below it there is a price, usually lower than you expected, and a paragraph about agentic workflows. You are sitting on work that already runs fine, and the question you actually have is small and specific: does any of this change what I should be sending my tasks to on Monday. The chart does not answer that, and it is not designed to.

This guide is a standing procedure for that moment. It is for someone who pays for AI by the token or by the seat, runs a handful of tasks repeatedly, and has neither the time nor the appetite to re-evaluate their setup every three weeks. It is not for anyone whose model is fixed by a procurement or compliance decision, and it is not worth the effort if you use a chat product a few times a week, where the vendor picks the model for you and the switch happens without your involvement. Every price and date below was fetched from the vendor’s own page on 4 September 2026, which is itself part of the point.

The releases arrive faster than anyone can test them

The cadence has settled into something you cannot keep up with by reacting, and it is worth seeing the shape of it before designing a response.

Google announced Gemini 3.6 Flash on 21 July 2026, with a set of benchmark gains and a price of $1.50 per million input tokens and $7.50 per million output tokens [3]. Gemini 3.7 Flash followed on 13 August 2026 [2]. Gemini 3.8 Flash reached general availability on 2 September 2026, described as engineered for long-horizon software engineering, autonomous agents and complex enterprise workflows [2]. That is three generations of one model tier in six weeks, and it does not count the Omni Flash, Transcribe and Lyria releases in the same window [2]. Anthropic’s own status table currently lists 11 active models at once, from Haiku 4.5 up to Fable 5.1 [5].

If you treat each of those as a decision, you have acquired a part-time job with no output. The alternative is not ignoring releases. It is having a procedure short enough that running it is cheaper than worrying about whether you should, and a default of doing nothing when the procedure comes back inconclusive. The rest of this guide is that procedure, in the order the steps are worth taking.

The announcement price is not the price

Start on the pricing page, not the launch post, because the launch post is a snapshot of one day and the pricing page is what you will be billed against.

Gemini 3.6 Flash is the clean illustration. It launched at $1.50 and $7.50 per million tokens [3]. Today the same model is listed at $0.75 per million input tokens and $3.75 per million output through 31 December 2026, and it appears on the free tier at no charge [1]. The model it replaced, Gemini 3.5 Flash, is listed at $1.50 input and $9.00 output with no expiry note attached [1]. Read those two rows together and the conclusion is uncomfortable for anyone who sat still: staying on the older model now costs twice as much on input and more than twice as much on output as moving to its successor.

Three things on that page matter more than the headline rate. The first is the expiry date, because a price quoted “through 31 December 2026” is a promotional rate with a documented end, and several Gemini rates revert on 1 January 2027 [1]. The second is the batch discount, listed as a 50% reduction on both input and output [1], which for queued work moves more money than most model swaps do. The third is the shape of the ladder inside one vendor, which tells you what a tier is worth. Anthropic prices Fable 5.1 at $10 and $50 per million tokens, Opus 5 at $5 and $25, Sonnet 5 at $2 and $10, and Haiku 4.5 at $1 and $5 [8]. Gemini 3.5 Flash-Lite sits at $0.30 and $2.50 [1]. A new release is only interesting to you if it changes your position on one of those ladders.

A benchmark score is a fact about the benchmark

The scores in a launch post are real measurements. They are just measurements of something other than your work, and the gap between those two things is larger than most people assume.

OpenAI published the clearest evidence of this when it built SWE-bench Verified. Annotating 1,699 samples from the original benchmark, it found 38.3% flagged for underspecified problem statements and 61.1% flagged for unit tests that might unfairly mark a valid solution incorrect, and it filtered out 68.3% of samples overall [4]. On the cleaned set, GPT-4o’s score moved from 16% to 33.2% [4]. The model did not change. The measuring instrument did, and the headline number doubled.

Hold the July Gemini figures against that standard. Google reported Gemini 3.6 Flash at 49% on DeepSWE against 37% for 3.5 Flash, 83.0% on OSWorld-Verified against 78.4%, and 63.9% on MLE Bench against 49.7% [3]. For 3.5 Flash-Lite it reported 54% on Terminal-Bench 2.1 against 3.1 Flash-Lite’s 31%, and 54.2% on SWE-Bench Pro against Gemini 3 Flash’s 49.6% [3]. Those are genuine and they point in a consistent direction, which is useful information. What they cannot tell you is whether the thing you do all day, summarising client documents in a particular format or extracting fields from a particular kind of invoice, moved at all. Every comparison in a launch post is against the vendor’s own previous model on the vendor’s chosen tasks, which is the correct way to demonstrate progress and the wrong way to make a routing decision.

Twenty of your own tasks settle it in an afternoon

The replacement for the leaderboard is smaller and duller than people expect, and you only have to build it once.

Anthropic’s own guidance on writing evaluations gives the design rules plainly. Be task-specific and mirror your real-world task distribution, including edge cases. Structure questions so grading can be automated, whether by string match, code or another model. And prioritise volume over quality, because more questions with slightly lower-signal automated grading beats fewer questions graded by hand [7]. It also ranks the grading methods: code-based is fastest and most reliable but lacks nuance, model-based grading is flexible and scalable once you have checked it is reliable, and human grading is the highest quality and the one to avoid where possible [7].

In practice that means a folder. Twenty to thirty real inputs pulled from work you have already done, each paired with the output you were happy with, stored as files you own rather than inside one product’s saved projects. When a release lands, you run the set through the candidate model, read the results next to the old ones, and record two numbers: how many outputs you would have shipped without editing, and what the run cost. The first time takes an afternoon. Every time after that it takes the length of a coffee, which is the entire reason it survives contact with a busy month.

Record the cost properly while you are there. Token counts differ between models even at identical prices, and Google’s claim for 3.6 Flash was framed as a 17% reduction in output token usage against 3.5 Flash on the Artificial Analysis Index rather than as a rate cut [3]. A cheaper per-token price on a chattier model is not a saving, and your own run is the only place that shows up.

Price the switch, not the model

The saving is the easy half of the arithmetic. The cost of moving is the half that gets left out, and it is what makes most small improvements not worth taking.

Switching costs you the prompts you tuned to one model’s habits, the downstream formats that quietly depend on how the old model laid things out, and the week of low-grade suspicion where you check outputs you used to trust. None of that appears on a pricing page. All of it is real, and it scales with how deeply the model is wired into anything automated. The honest way to handle it is to put a number on the migration and see how long the saving takes to repay it.

calculator
How long a model switch takes to pay for itself
months to pay back

Migration cost divided by monthly saving. Computed in the page; nothing is sent anywhere.

Read the answer against how long you expect the new model to stay current, which on the cadence above is not long. If the payback is under two months, the switch is easy. If it is over six, you are doing unpaid work for a model that will be superseded before it breaks even, and the right move is to note the candidate and wait for the next release to make the case properly. The exception is a task whose volume is growing, where the saving compounds and the same arithmetic flips within a quarter.

The retirement date is the only switch you do not choose

Everything above assumes staying put is free. It is, until the model you are on is turned off, and that date is published well in advance by everyone who matters.

Anthropic commits to at least 60 days’ notice before retiring a publicly released model, and publishes a table of tentative retirement dates for models that are still active [5]. Haiku 4.5 currently shows a retirement date of not sooner than 15 October 2026, while Fable 5.1 shows not sooner than 1 September 2027 [5]. The same table records what the end of the process looks like: claude-opus-4-1 was deprecated on 5 June 2026 and retired on 5 August 2026, after which requests to it fail [5]. OpenAI’s policy is longer and more granular, at least 6 months for generally available models and at least 3 months for specialised variants, but with a sharp exception: preview models, identified by “preview” in the name, may be retired with much shorter notice, such as 2 weeks [6]. Its current table shows the older GPT-5 and o3 snapshots announced on 11 June 2026 for shutdown on 11 December 2026, with the 5.6 family as the replacement [6]. Google’s release notes carry the same pattern, including the gemini-omni-flash-preview endpoint, which the notes say will be deprecated on 30 September 2026 [2].

Two habits follow. Pin an explicit model version in anything automated rather than an alias that moves under you, so an upgrade is something you did on purpose. And keep the retirement date for every model you depend on in the same calendar you keep client deadlines in, set 60 days early. A forced migration you have three months to plan is an afternoon. The same migration discovered through a wave of failed requests is a bad week, and the notice period was sitting on a public page the whole time.

checklist
Before you switch to a newly released model
0 of 8 · saved in this browser only

What still goes wrong

Twenty tasks is a small sample, and it will not catch a model that is slightly worse in a way that only shows up on the unusual inputs. That is the known trade in Anthropic’s own advice, which prefers volume with automated grading over a handful of carefully hand-graded cases [7], and it is the right trade for a small operation. It is still a trade. If the work carries real consequences, keep a sample of outputs from before the switch and compare against them after a month of ordinary use, because the failure you are looking for is drift you stopped noticing rather than an error you would have caught on day one.

The published figures also move faster than any guide can. Gemini 3.6 Flash launched at one price in July and sits at half that price in September [1][3], and the promotional rates on that page have an expiry date written into them [1]. Retirement dates are described as tentative and can shift [5]. Everything cited here was live on 4 September 2026, and the procedure matters more than the numbers precisely because the numbers will be wrong by the time you need them.

Finally, this procedure is deliberately biased towards inaction, and that bias has a cost. A model two generations back will keep working, keep producing plausible output, and keep costing more than its replacement without ever failing loudly enough to prompt a review. The counterweight is not vigilance, which nobody sustains. It is a date in the calendar, once a quarter, to open the pricing pages of the vendors you use and re-run the test set on your two highest-volume tasks. Fifteen minutes, four times a year, catches almost everything that matters, including the releases you were right to ignore.

sources
  1. 01Google — Gemini API pricingai.google.dev
  2. 02Google — Gemini API release notes / changelogai.google.dev
  3. 03Google — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyberblog.google
  4. 04OpenAI — Introducing SWE-bench Verifiedopenai.com
  5. 05Anthropic — Model deprecationsplatform.claude.com
  6. 06OpenAI — Deprecationsdevelopers.openai.com
  7. 07Anthropic — Develop test casesplatform.claude.com
  8. 08Anthropic — Claude pricingclaude.com
next guide
Only hand an agent work you can check
10 min · verified 2026-09-04
related guides