friday, september 18, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

When the top model leaves your plan

Arrange your AI work so a model moving between subscription tiers costs you an afternoon of rerouting instead of a month of improvising.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You picked the plan because of one model. It handled the messy client brief, the contract review, the code that three other models mangled, and you quietly built a week around it. Then a support page changed. Claude Fable 5 was included for up to 50% of weekly usage limits on Pro, Max, Team and select Enterprise plans through 7 July 2026, after which it was available via usage credits [5], billed at standard API pricing rates [3]. Earlier the same year, on 13 February 2026, OpenAI deprecated GPT-4o, GPT-4.1, GPT-4.1 mini, o4-mini and GPT-5 in ChatGPT, and existing conversations and projects were defaulted to their GPT-5.3 Instant and GPT-5.4 Thinking and Pro equivalents [7]. Nobody asked you.

This guide is for a solo operator or small team paying for one or two AI subscriptions, who has real work sitting on top of a specific model and wants that work to survive the next tier change. It is not for enterprise buyers with a negotiated contract and a named account manager, and it is not for anyone still on a free tier, where the answer is simpler: nothing is promised to you at all. Prices below are US list prices fetched on 4 September 2026.

Your subscription buys usage, not a particular model

Look at what the money actually purchases. Claude’s Free plan gives you Sonnet and Haiku. Pro is $17 a month billed annually at $200 up front, or $20 billed monthly, and adds Opus. Max starts at $100 a month, lets you choose 5x or 20x more usage than Pro, and adds Fable at 50% of weekly limits [1]. The same shape appears at Google: the free tier gets access to Gemini 3.6 Flash and varying access to 3.1 Pro, Google AI Pro at $19.99 a month buys 4x higher usage access than free, and Google AI Ultra, starting at $99.99 a month, buys up to 20x higher usage limits in Gemini than the Pro plan along with first access to advanced features like Deep Think [8].

The pattern is consistent enough to plan around. The tier is the product. The newest and most expensive model is the thing that distinguishes the expensive tier from the cheap one, which is exactly why it is the first thing to move when the economics move. Serving the frontier model at flat rate to everyone who pays $20 is the least defensible line on the vendor’s cost sheet, so that line gets adjusted, and it gets adjusted more often than the price of the tier does.

Today, Fable 5 and Fable 5.1 are reachable on every paid Claude plan, but the terms differ by seat. On Max, premium Team and premium Enterprise seats they are included as a standard part of the plan, and you can use up to 50% of your weekly usage limits on them at no extra cost. On Pro plans and standard Team seats they are not included in the plan’s usage limits at all, and you reach them with usage credits [2], which bill at standard API pricing rates [3]. That arrangement is two months old. Fable 5 was pulled on 12 June 2026 when the US government applied export controls, redeployed globally on 1 July, and included at up to 50% of weekly limits only through 7 July [5]. The version you read on a vendor page is a snapshot, not a contract.

The notice you get depends on which door you came in through

There are two entrances to the same models, and only one of them has a written promise attached.

On the API side, the promises are published and dated. Anthropic notifies customers with active deployments for models with upcoming retirements, providing at least 60 days’ notice before model retirement for publicly released models, and states that impacted customers will always be notified by email and in the documentation [4]. OpenAI publishes longer windows: at least 6 months for generally available models, at least 3 months for specialized variants such as chat, Codex and deep research versions, and as little as about 2 weeks for anything with preview in the name [6]. Those pages carry actual tables. Anthropic’s currently lists claude-sonnet-4-5-20250929 as active with a tentative retirement date not sooner than 29 September 2026, and records that claude-opus-4-1-20250805 was deprecated on 5 June 2026 and retired on 5 August 2026 [4]. OpenAI’s records the transcription models whisper-1, gpt-4o-transcribe and gpt-4o-mini-transcribe shutting down on 26 February 2027, and gpt-realtime shutting down on 20 January 2027 [6].

On the subscription side there is no equivalent commitment. A consumer plan promises capacity, features and support. It does not promise that a named model stays inside that capacity, and the vendors do not pretend otherwise. When ChatGPT retired that group of models in February, some paid tiers got short extensions: Enterprise workspaces retained access to GPT-5 models through 19 February 2026, GPT-5 Pro remained available to all paid users through the same date, and Business, Enterprise and Edu customers could keep using GPT-4o inside custom GPTs until 3 April 2026 [7]. Those were concessions of a few weeks. The models continued to be available through the OpenAI API [7], which is the point worth internalising. The model did not disappear. It moved behind a meter.

Most of your work does not need the top tier

The price ladder is unusually clean, which makes the audit easy. Claude Fable 5.1 costs $10 per million input tokens and $50 per million output tokens. Opus 5 is $5 and $25, Sonnet 5 is $2 and $10, and Haiku 4.5 is $1 and $5 [1]. Opus is exactly half the frontier price on both sides, Sonnet exactly a fifth, Haiku exactly a tenth. So the question for every recurring task is not which model is best. It is whether the fifth-price model changes the outcome of this particular work.

Run the test while you still have both. Take your two highest-volume recurring tasks, the ones you do weekly or daily, and run each on the tier below with the same inputs. Read the two outputs beside each other and ask what you would have had to fix. If the answer is a comma and a heading, that task belongs on the cheaper tier permanently. If the answer is that the reasoning went shallow in a way a client would notice, you have found a task that genuinely needs the top model, and now you know what you are protecting.

This matters even when you are not paying per token. On Max, Fable usage draws on the same weekly allowance as everything else, up to half of it [2]. Sending routine drafting to the frontier model does not just waste money in the abstract. It spends the capacity you will want on Thursday afternoon for the thing that actually needed it.

The output of this exercise is one short document: each recurring task, the model tier it runs on, and the reason. Keep it somewhere you will find it. When a tier change lands, that document is the difference between rerouting five workflows in an afternoon and rediscovering your own stack from memory.

Metering changes the shape of the risk, not just the size

A subscription is a capped liability. You know the worst case before the month starts. Usage credits are a different instrument. When you reach your session usage limit, you see a notification, and if usage credits are enabled and you have funds available you can choose to continue working, with subsequent usage billed at standard API pricing rates [3].

Set the ceiling before you need it, not after the first surprising statement. You can set a maximum amount you are willing to spend on usage credits each month, and there is a daily redemption limit of $2,000 [3]. Read that second number correctly: it is a hard stop the vendor imposes, not a budget anyone should regard as normal. Auto-reload will automatically add funds when your balance drops below a threshold you configure, which is convenient and is also the mechanism by which a misbehaving loop turns into a real bill [3]. If you decide the metered path is not for you, usage credits can be turned off at any time through Settings then Usage [3].

Before you enable anything, price the workload. One recurring workflow is usually enough to decide the whole question.

calculator
One workflow at metered rates
$ / month

Priced at Claude Fable 5.1 rates of $10 per million input tokens and $50 per million output tokens. Opus 5 is exactly half these figures, Sonnet 5 a fifth, Haiku 4.5 a tenth [1]. Computed in the page; nothing is sent anywhere.

Then compare that number against the next tier up rather than against zero. If one workflow on frontier rates comes to more than $100 a month, and Max starts at $100 a month with Fable at 50% of weekly limits [1][2], the metered path is the expensive one and you are paying for flexibility you are not using. If it comes to $12 a month, credits with a modest cap are the cheaper and less committed answer. The mistake is treating the tier price as the only number and the credits as free-floating.

Cheap switching is something you build in advance

Portability is not a philosophy here. It is three concrete habits.

Keep the material outside the chat. System prompts, house-style notes, the reference documents you paste at the start of every session, the examples of good output: these belong in files you own, not in a vendor’s conversation history. A workflow whose instructions live in a saved project inside one product is a workflow you cannot test anywhere else, which means you cannot price the alternative and cannot move when you need to.

Name the model in one place. If a task is described in a document, the document should say which tier it runs on, so changing it is an edit rather than an archaeology project. If anything you run calls an API, pin the dated model ID rather than an alias, and write the retirement date next to it from the vendor’s table [4][6]. Aliases move under you silently. Dated IDs fail loudly, and loud is better.

Check the deprecation page on a schedule. Once a month, five minutes, both vendors you use. The notice goes to customers with active deployments [4], which is not necessarily the address you read, and a page you check deliberately beats a message you might filter.

One habit that is worth resisting: rebuilding a prompt around a specific model’s quirks. Prompts that depend on one model’s tics are the ones that break when the tier changes, and they break in ways that are hard to see because the output still looks fine.

It is also worth remembering that price is not the only reason a model moves. Fable 5’s July disruption started with export controls applied by the US government on 12 June 2026 and lifted on 30 June, after which the model was redeployed globally on 1 July with an improved safety classifier [5]. No pricing page would have predicted that. Availability can change for regulatory reasons, safety reasons and capacity reasons, and the plan that survives all three is the one that does not depend on any single model being there on Monday.

checklist
Before the next tier change lands
0 of 7 · saved in this browser only

What still goes wrong

The cheerful version of this guide says everything runs fine one tier down. It does not. Some work genuinely needs the frontier model, and if that work is yours and you are on Pro, you are now paying standard API rates for it while someone on Max gets it inside their weekly allowance [1][2][3]. The honest answer in that case is to move tiers, not to pretend Sonnet is Fable. Tier discipline saves money on the eighty per cent of work that never needed the top model. It does not conjure a cheaper version of the twenty per cent that did.

The side-by-side test is also weaker than it sounds. Two outputs, read once, by the person who wrote the prompt, is not an evaluation. It catches large quality drops and misses small ones, and small ones compound quietly across a few hundred runs before anyone notices the tone drifted or the summaries started missing the caveat. If a task carries real consequences, keep a sample of the old model’s outputs before you switch, and check the new ones against it in a month rather than trusting the first impression.

Finally, the published notice periods cover retirement, not behaviour. A model can be updated under the same name, and no deprecation table records that. Vendors also reverse themselves inside weeks: Fable 5 went from withdrawn on 12 June 2026 to redeployed on 1 July to metered for Pro seats after 7 July [2][5], so an afternoon spent migrating away can be an afternoon wasted. The defence is not prediction. It is keeping the cost of switching low enough that being wrong about the direction stops mattering very much.

sources
  1. 01Claude — Pricingclaude.com
  2. 02Anthropic Help Center — Claude Fable models on your plansupport.claude.com
  3. 03Anthropic Help Center — Manage usage credits for paid Claude planssupport.claude.com
  4. 04Anthropic — Model deprecationsplatform.claude.com
  5. 05Anthropic — Redeploying Claude Fable 5anthropic.com
  6. 06OpenAI — API deprecationsdevelopers.openai.com
  7. 07OpenAI Help Center — Retiring GPT-4o and other ChatGPT modelshelp.openai.com
  8. 08Google — Gemini subscription plansgemini.google
next guide
When to let a coding agent act without asking you
9 min · verified 2026-09-04
related guides