saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

What your plan includes is not a price

Tell the two meters on your AI bill apart, put a dollar figure on the usage a seat includes, and know what to do the month a plan's terms move.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You bought a seat, the charge settled into the statement, and you stopped thinking about it. Then one month the work changes shape without the invoice moving. The model you use for the hard jobs is unavailable halfway through a Thursday. Someone runs the thing they run every morning and is told they are out of usage. A plan that covered a model at no extra charge starts asking for credits instead. Nothing you agreed to has been broken, and nothing on your card has changed. What moved is the part of the deal that was never expressed as a number.

This guide is for whoever pays for AI tools at a company of 1 to 50 people and has at least one seat-based subscription doing real work. It is not for teams on a committed-spend enterprise contract with a named vendor rep, who get notice and a renegotiation instead of a surprise. Every figure below was read from a vendor’s own page on 5 September 2026, and the pages themselves say the figures move: Anthropic’s pricing page carries the line that price and plans are subject to change at Anthropic’s discretion [1].

Your AI bill has a published price and an unpublished allowance

The metered half of the market is priced in public, per million tokens, down to cents. Claude Fable 5.1 is $10 input and $50 output, Opus 5 is $5 and $25, Sonnet 5 is $2 and $10, Haiku 4.5 is $1 and $5 [2]. On OpenAI’s list, GPT-5.6 Sol is $5 and $30, Terra is $2 and $12, and Luna is $0.20 and $1.20 [4]. Gemini 3.5 Flash-Lite is $0.30 and $2.50 [6]. You can multiply these by your volume and get a number that is right.

The subscription half is not priced that way. It is priced in multipliers. Claude Max is from $100 per month for 5 times or 20 times more usage per 5-hour session than Pro, where Pro is $17 per month billed annually at $200 up front, or $20 monthly [1]. A Claude Team premium seat is $100 per seat annually or $125 monthly for 5 times more usage than a standard seat, which is $20 annually or $25 monthly [1]. ChatGPT Business has the identical shape, $20 and $25 for standard and $100 and $125 for premium [5], and describes everyday text chats as unlimited with an asterisk that resolves to a sentence saying usage must be reasonable and comply with the policies [5].

Read that chain again. A premium seat is defined against a standard seat. Max is defined against Pro. Pro is defined as more than free. At no point does the chain touch a unit. There is no token count, no message count, no hours. The one thing under your control is observation after the fact, through the usage view in settings [1], which tells you how much of an unnamed quantity you have consumed.

That asymmetry is the whole subject. Half your bill is a price, and prices are defended in public because customers compare them. The other half is an allowance, and an allowance can be tightened without a single published number changing.

Convert the allowance into dollars before you are forced to

You can put a figure on the allowance, because the vendors publish the rate that applies the moment it runs out. Anthropic’s pricing page says that when you reach a limit you can wait for it to reset, move to a higher plan, or, on paid plans, turn on usage credits to keep working at standard API rates [1], and the standard API rates are the same list you already have [2]. Cursor documents the same handover in plainer words: when you exceed your included monthly usage you either add on-demand usage and continue at the same API rates with pay-as-you-go billing, or upgrade your plan for more included usage [7].

So the metered price list is not just the API’s price list. It is the shadow price of your subscription, and it is the price you fall back to when the terms move. That makes the arithmetic worth doing once, in a quiet month, before anyone is annoyed.

Take a typical month for one heavy user. Estimate the input and output tokens they actually push through the plan, price that volume at the list rate for the model they default to, and compare the result to what the seat costs. If the metered figure comes out near or below the seat price, the subscription is a convenience purchase and a change in terms is an inconvenience. If it comes out at several times the seat price, the seat is carrying a subsidy, and subsidies are the line items that get adjusted.

calculator
What your seat covers, priced at metered rates
$ / month of exposure

Price defaults are Claude Fable 5.1 list rates; enter 5 and 25 for Opus 5, or 0.20 and 1.20 for GPT-5.6 Luna. A large positive number is what would land on your card if the same work moved to metered credits. Computed in the page; nothing is sent anywhere.

The included access to the expensive model is the least stable thing you buy

The spread across a single vendor’s menu is a factor of 10 from bottom to top: Haiku 4.5 at $1 and $5 against Fable 5.1 at $10 and $50 [2]. On OpenAI’s list it is a factor of 25, from Luna at $0.20 and $1.20 to Sol at $5 and $30 [4]. A flat monthly seat that includes the top of that menu is the vendor absorbing the expensive end for a fixed fee, which is a fine offer and an unstable one.

You can see the fences built to make it survivable. On Claude’s plan comparison, Sonnet and Haiku run on every individual plan and Opus starts at Pro, while Fable is reachable on Pro only through usage credits and on Max 5x and 20x at 50 percent of weekly limits [1]. Those weekly limits sit above the 5-hour session window the plans are defined against, and the same page reserves the right to limit usage in other ways, such as weekly and monthly caps or model and feature usage, at Anthropic’s discretion [1]. That is 3 separate constraints on one model: which plan, what share, what window. OpenAI fences with purchase instead, noting that Enterprise and Business can purchase credits for more access, and that credit-based and token-based pricing are available for Enterprise plans [5]. Cursor fences by splitting the pool, with its own models in one bucket and third-party models from Anthropic, Google and OpenAI in another, and adds a Cursor Token Rate of $0.25 per million tokens on third-party models for Teams and Enterprise [7].

Every one of those fences is a dial. Moving a dial does not require changing a headline price, does not show up on a comparison table, and is covered by the sentence about plans being subject to change at the vendor’s discretion [1]. When you hear that a premium model has moved out of plan limits and into separately purchased credits, that is a dial turning, and the announcement is describing a decision that was always available to them.

The practical response is not to distrust the offer. It is to notice which work is currently sitting inside a fenced allowance. If a job earns money, has a deadline, or would embarrass you if it stopped, it should have a path that does not depend on how a multiplier is defined this quarter, even if you never use that path.

Published prices move in both directions, and the dates are printed

The reason to price the allowance rather than panic about it is that the direction of travel is genuinely unpredictable, including to the vendor. Anthropic’s pricing documentation currently states that the $2 and $10 per million token pricing for Sonnet 5, announced at launch as introductory pricing through 31 August 2026, is now the standard price, and that the previously scheduled increase to $3 and $15 on 1 September 2026 will not occur [2]. A promotion ended by becoming permanent. Anyone who rebuilt a pipeline in August to escape the increase spent the effort for nothing.

The opposite case is on Google’s page, printed just as plainly. Gemini 3.8 Flash is $0.75 input and $3.75 output on the paid tier through 31 December 2026, and $1.50 and $7.50 from 1 January 2027 [6]. The doubling is not a rumour. It has a date and a page. The same entry moves that model’s context caching rate from $0.075 to $0.15 per million tokens on the same date, and the hourly cache storage price from $0.50 to $1.00 per million tokens [6], so the increase reaches the parts of a bill that are easy to forget. Cursor’s documentation carries a comparable expiry, noting that Legacy Enterprise Auto is priced per million tokens regardless of model only until 7 September 2026 [7].

The habit that follows is small. When a price goes into a forecast, write the date you read it next to it, and put a reminder on the earliest date that appears anywhere in that vendor’s own terms. Then check whether the date moved rather than watching the number.

Free tiers deserve the same treatment with an extra clause, because free is a price too and it is paid in something other than money. Google’s pricing page states that on the free tier content is used to improve their products, and that on the paid tier content is not [6], and it lists access to context caching as a paid-tier feature [6]. A free allowance is the most adjustable of all allowances, and the exchange rate is not published anywhere.

A retirement date is a scheduled price change with a deadline

The last dial is the model itself. Anthropic documents that it notifies customers with active deployments of upcoming retirements, providing at least 60 days’ notice before retirement for publicly released models [3]. A deprecated model still works but is no longer recommended, with a replacement named and a retirement date assigned; a retired model is no longer available, and requests to it fail [3].

The current table is worth checking against your own routing, with one caution about how to read it. Every model below is listed as active with no deprecation date, and the dates are tentative earliest dates rather than scheduled retirements. Fable 5.1 is listed as retiring not sooner than 1 September 2027, Opus 5 not sooner than 24 July 2027, Sonnet 5 not sooner than 30 June 2027, and Haiku 4.5 not sooner than 15 October 2026 [3]. Haiku 4.5 is the cheapest model still active on the price list at $1 and $5 [2], which is exactly why it tends to be the one carrying overnight classification jobs and other high-volume work nobody looks at. The earliest date on which that work could need rebuilding is now printed, and the replacement is not obliged to arrive at the same rate.

The instruction in the documentation is that once a model is deprecated you migrate all usage to a suitable replacement before the retirement date, and that you should test replacement models on your own tasks well before that date rather than at it [3]. Read as a budget matter, that is a request to re-price the job as well as re-test it, because a migration that keeps the output identical and moves the cost per run is still a change to your bill.

Run a fixed drill when the terms move, not a rebuild

When a vendor announces a change to what a plan includes, the failure mode is not paying too much. It is spending a week reacting to an announcement whose details are still moving, as the Sonnet 5 reversal shows [2].

Give it a fixed shape instead. First, read the change against your own usage view rather than against the announcement’s framing [1], because most changes affect a share of usage you may not be near. Second, price the affected work at the metered rate for the model it runs on [2][4][6], which is the number you would actually pay if you did nothing. Third, decide once between 3 options: pay the metered rate, move that job down the menu, or stop doing it. Fourth, write down which choice you made and the date, so the next announcement is an edit rather than an investigation.

That drill should take an afternoon, and it works whichever direction the change goes. If the promotion is extended or made permanent, you have lost an afternoon and gained a document. If the fence really does move, you already know which jobs are behind it and what they cost on the other side.

checklist
When a plan's terms change
0 of 8 · saved in this browser only

What still goes wrong

Every number here was read on 5 September 2026 and carries the vendors’ own caveat that prices and plans change at their discretion [1]. The structure outlasts the figures. Two meters, one published and one not; the metered rate as the shadow price of the unpublished one; a date written next to every price; a fixed drill instead of a rebuild.

The dollar conversion flatters the subscription more than it should. Metered rates buy API calls, while a seat buys an application, connectors, mobile apps, and the model access bundled inside them, so the comparison is not like for like and the exposure figure is not a purchase recommendation. It answers one narrow question, which is what would land on your card if the same volume of work moved to credits at list price. Treat a large answer as a reason to know where your fallback is, not as a reason to cancel anything.

The harder limit is that allowances are adjusted quietly by design. There is no changelog for a multiplier, and the first signal is usually a person hitting a wall in the middle of the afternoon and assuming they did something wrong. No amount of budgeting arithmetic detects that on its own. What detects it is one named person whose job is to notice that several people mentioned limits in the same week, and to go and read the pricing page rather than guess. That is a small assignment, and it is the only part of this that has to be a person.

sources
  1. 01Claude — Pricingclaude.com
  2. 02Claude Docs — Pricingplatform.claude.com
  3. 03Claude Docs — Model deprecationsplatform.claude.com
  4. 04OpenAI — API pricingopenai.com
  5. 05OpenAI — Business pricingopenai.com
  6. 06Google — Gemini API pricingai.google.dev
  7. 07Cursor Docs — Models and pricingcursor.com
next guide
Routing tasks between AI models
9 min · verified 2026-09-05
related guides