What cheap intelligence changes, and what it does not
Model prices fall fast and your bill does not. Learn where the saving actually lands, and how to plan for the next price cut before it arrives.
on this page · 0 / 0 checked
Every few months a lab announces that its models got cheaper, and the announcement reads as though it were written for you. In late July 2026 OpenAI cut the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, and published a post naming the strategy behind it [1]. Google’s free Gemini tier charges nothing for input or output on its Flash models [8]. ChatGPT’s free plan includes unlimited text chats with Luna, subject to abuse guardrails [4]. If you run a small business on these tools, the sensible reaction is to check your own invoice, notice that it did not fall, and assume you missed something.
You did not. Falling token prices are real and the trend is durable, but almost none of the saving arrives as a smaller bill, and the parts of your work that cost the most were never priced in tokens to begin with. What follows is what to actually do when the floor drops again, which it will. This is for solo operators and small teams paying for AI out of their own revenue. If you are negotiating committed-spend contracts, or running enough inference that a 20% serving-cost improvement is a line in a board deck, read the vendors’ pricing pages directly instead.
The cut lands on the cheap tier, not the frontier
Read the July announcement carefully and it contains two opposite price moves. Luna, the budget model, fell 80%, to $0.20 per million input tokens and $1.20 per million output tokens. Terra, the middle tier, fell 20%, to $2 and $12. In the same post, a Fast mode arrived at up to 2.5 times the speed of standard processing for twice the price [1]. One number dropped by four-fifths and another doubled, in the same announcement.
The current published list makes the shape obvious. OpenAI’s flagship table now opens with GPT-6 Astra at $10 and $50 per million input and output tokens, then GPT-5.6 Sol at $4 and $20, Terra at $2 and $12, and Luna at $0.20 and $1.20 [3]. Anthropic lists Haiku 4.5 at $1 and $5, Sonnet 5 at $2 and $10, Opus 5 at $5 and $25, and Fable 5.1 at $10 and $50 [6]. Google’s paid tier prices Gemini 3.8 Flash at $0.75 and $3.75, while Gemini 3 Pro Image bills image output at $120 per million tokens [8].
Two things are happening at once and only one of them is a price cut. The bottom of each list falls toward zero, so that Luna’s input now costs one-fiftieth of Astra’s. The top does not fall. It moves up, because each new flagship enters above the one it replaces and the previous flagship slides down a rung. You can watch the slide inside a single vendor’s list: Claude Sonnet 5 costs $2 and $10, undercutting the older Sonnet 4.6 at $3 and $15 for the same slot in the lineup [6]. Even the slide is provisional. OpenAI marks Sol’s $4 and $20 as promotional pricing available at least through 21 November 2026 [3].
The practical consequence is that price announcements have to be read by tier, not by headline. If your work runs on a flagship model because it genuinely needs to, an 80% cut to the budget tier changes nothing for you. If it runs on the budget tier out of habit, it changes a great deal, and you should be checking whether the tier above is now affordable rather than banking the difference.
Cheaper tokens get spent, not saved
OpenAI’s own post supplies the reason your bill is flat. It reports that six months after signing up, people send roughly 50% more messages each day and use ChatGPT for about twice as many kinds of work [1]. The same post says that across OpenAI itself, agentic work through Codex now accounts for 99.8% of weekly output tokens [1]. That second figure describes the vendor’s own staff rather than its customers, and that is what makes it useful: the organisation closest to free inference is the one whose token consumption went vertical. Cost per token fell. Tokens per job rose, because the same money buys a model that reasons for longer and an agent that runs unattended for an hour.
This is not a trick played on you. OpenAI states the mechanism plainly, writing that when the cost of useful intelligence falls, more work becomes worth doing [1]. But it does mean the honest way to read a price cut is as a budget for more attempts rather than as a discount. At Terra’s published rates, 30 million input tokens and 6 million output tokens a month costs $132 [3]. If those prices halve, you can bank $66 or you can run every draft twice and keep the better one. Only one of those changes what leaves your business.
Decide which before the cut lands, because the default is drift. Usage quietly expands to fill the new price, and three months later you have spent the money without ever choosing to.
Defaults are GPT-5.6 Terra's published rates [3]. Halve both prices to see what the next cut is worth, then double the token counts to see what usually happens instead. Computed in the page; nothing is sent anywhere.
Free access is a distribution channel with published mechanics
Free tiers are now genuinely useful rather than a demo, and each one has terms. ChatGPT’s free plan includes unlimited text chats with GPT-5.6 Luna, subject to abuse guardrails [4]. Google’s free Gemini tier charges nothing for input or output on its Flash models, with limited access to certain models and lower rate limits than the paid tier [8]. OpenAI’s academic programme, announced on 29 July 2026, gives researchers at qualifying institutions free access to frontier models including GPT-5.6 Sol Pro at launch, beginning with 10,000 researchers and scaling to 100,000 through 2027, with each approved researcher able to invite up to four collaborators from their institution who count toward that total [5].
Read the mechanics before you build on any of them. The academic programme is capped, restricted to recognised degree-granting colleges and universities with a high level of research activity, and requires applicants to verify their institutional affiliation and describe their intended scientific use [5]. OpenAI places it inside a stated commitment of more than $250 million through 2027 to support external scientific research and discovery, which also covers NextGenAI and the Department of Energy’s Genesis Mission [5], so that figure is not a budget for this programme alone. Google is blunter about what its free tier costs you: content on the free tier is used to improve Google’s products, and content on the paid tier is not [8].
Two things follow for a small business. First, if any part of what you sell is access to a capable model, check whether your buyers already have one for nothing. That pitch has an expiry date, and for the academic market the date has passed. Second, treat a free tier as a cost you have not been invoiced for yet, payable either in money later or in data now. Build so the same job runs on a paid tier with a configuration change, and price your offer so it survives the month the free tier narrows.
The savings that matter are not on the price list
The two levers that actually move a small operator’s bill are published but rarely read. Batch processing halves both input and output prices in exchange for accepting latency. Anthropic states the 50% discount outright, and OpenAI’s batch table is exactly half its standard table, with Terra at $1 and $6 rather than $2 and $12 [3][6]. Prompt caching is larger still. Luna’s cached input costs $0.02 per million against $0.20 uncached, a tenth of the price [3], and Claude cache hits are billed at 0.1 times the base input rate, or 0.025 times on Fable 5.1 and Mythos 5.1 [6]. If you send the same 40-page contract alongside each of 20 questions, caching is the difference between paying for that contract once and paying for it 20 times.
Then there is the cut you never see on a pricing page. OpenAI reports that GPT-5.6 Sol helped reduce its end-to-end serving costs by 20% and improved speculative decoding, increasing token-generation efficiency by more than 15% [1]; that improvements to retained reasoning and context management raised Sol’s score on the public ARC-AGI-3 task set from 13.3% to 38.3% while using six times fewer output tokens [1]; and, in a later post, that Sol with maximum reasoning reached a new high while using 54% fewer output tokens than another leading model [2]. A model that reaches the answer in fewer output tokens is a price cut in everything but name, and because output is billed at 5 to 6 times input across OpenAI’s and Anthropic’s current lists [3][6], it is often the bigger one.
Prices also move the other way, and those moves are scheduled in public. OpenAI and Anthropic both add 10% for data residency [3][6], and both bill web search at $10 per 1,000 searches on top of tokens [3][6]. Fast mode doubles OpenAI’s standard rates, and crossing into long-context billing doubles the input rate and adds half again to output, so Terra’s long-context request costs $4 and $18 [3]. Google’s paid Flash pricing is set to rise on 1 January 2027, from $0.75 and $3.75 to $1.50 and $7.50 [8]. That the moves go both ways is not guaranteed in your favour either: Anthropic had scheduled Sonnet 5 to rise to $3 and $15 on 1 September 2026 and then cancelled the increase [6], which is welcome and was never yours to plan around. Speed, jurisdiction and tool use are what you now pay a premium for, not raw text generation.
Whatever is cheap today has a retirement date
Abundance has a shelf life, and the vendors document it. Anthropic notifies customers with active deployments and gives at least 60 days’ notice before retiring a publicly released model [7]. The record shows the policy running to the day: Anthropic notified developers about Claude Opus 4.1 on 5 June 2026 and retired it on 5 August 2026 [7]. The same tables show Claude Haiku 3 retired on 20 April 2026, and Claude Sonnet 4 and Claude Opus 4 both retired on 15 June 2026 [7].
Read the live rows the same way. Haiku 4.5, the cheapest generally available Claude at $1 and $5 per million [6], is listed as active with no deprecation date and an earliest retirement of not sooner than 15 October 2026, and Claude Sonnet 4.5 shows not sooner than 29 September 2026 [7]. Those dates are floors rather than plans. The 60-day notice is what actually starts your clock, so a model with no deprecation notice today cannot disappear on you next month.
Behaviour changes as well as availability. The temperature, top_p and top_k parameters are deprecated, and setting one to a non-default value returns a 400 error on Claude 4.7 and later models [7]. Code that worked in January can fail on a model released in June without anyone having changed a line of it.
So the cheap model you standardise on has a working life measured in months, and the migration is unpaid work that arrives on a week of the vendor’s choosing rather than yours. Two defences are enough for a small operation. Keep the model name in exactly one place in your configuration, never scattered through prompts and scripts. And keep a short evaluation set, 10 to 20 real inputs with the outputs you consider acceptable, so that swapping models is an afternoon of comparison rather than a fortnight of suspecting it feels worse.
Price the part that does not get cheaper
If what you sell is work that a model now does in seconds, then falling token prices are falling prices for your deliverable too, and no amount of efficiency on your side changes that. The parts of the job that did not get cheaper are knowing which question to ask, holding the client’s context, and being accountable for the answer when it is wrong. Reviewing an output still costs one person’s attention for as long as the output takes to read, and no price cut has ever appeared on that line of your costs.
The concrete move when your token cost per deliverable falls is to leave the deliverable’s price alone and spend the difference on more attempts and a tighter review. Note what OpenAI itself claims to be selling: intelligence that keeps getting more capable, more affordable and more valuable to the people who use it, with infrastructure that is valuable not because it is large but because of what it makes possible [1]. That is a claim about output, not about invoices. Read your own numbers the same way.
What still goes wrong
An announced price is not your price. What you actually pay depends on the share of your input that hits a cache, whether you can tolerate batch latency, whether your requests cross into long-context billing, and how many output tokens the model spends thinking before it answers. Two businesses on the same model, the same tier and the same nominal workload can differ by a factor of five, and only some of those variables appear on a pricing page. The prices quoted throughout this guide are the ones published on 5 September 2026, in a category that changes on a few weeks’ notice, and several of them carry explicit end dates. Check them before you plan against them.
The efficiency numbers come from the companies that benefit from them. A benchmark score, a token count or a serving-cost improvement published by a lab about its own model is not an independent measurement, and none of the figures in the sections above were verified by anyone outside the vendor [1][2]. The 99.8% figure in particular describes usage inside OpenAI, by staff with no marginal cost of inference, which is the least representative population available. Treat all of it as directionally interesting and test on your own work, which is the only benchmark that bills you.
Finally, cheap is not the same as suitable. The temptation after a cut is to move a workload down a tier because the arithmetic suddenly works, and for summarising or classifying that is usually right. For anything where a wrong answer reaches a client, the saving is small and the failure is expensive. Free tiers narrow, cheap tiers get retired, and the model that costs $0.20 per million tokens today is a different model in eight months. Build so you can leave.
- 01OpenAI — Building abundant intelligenceopenai.com
- 02OpenAI — The full stack behind abundant intelligenceopenai.com
- 03OpenAI — API pricing (developer docs)developers.openai.com
- 04OpenAI — ChatGPT pricingchatgpt.com
- 05OpenAI — ChatGPT for Academic Researchersopenai.com
- 06Anthropic — Claude API pricingplatform.claude.com
- 07Anthropic — Model deprecationsplatform.claude.com
- 08Google — Gemini API pricingai.google.dev