Planning around AI compute scarcity
Learn the three ways a compute shortage actually reaches a one-person business, then move the work you can defer off peak and keep the rest working.
on this page · 0 / 0 checked
You run a small business on a few AI subscriptions and a couple of automations, and every so often something changes underneath you without warning. The assistant announces you have hit a weekly limit on a Tuesday morning. A model your workflow names by version stops answering. A price you budgeted around goes up in January. Each of these reads like a vendor decision made in a meeting you were not in, and in a narrow sense it is.
Underneath a lot of it is something less negotiable. The models you rent run in buildings that need land, cooling, and a connection to an electricity grid, and all three take longer to arrange than a product roadmap does. When the buildings arrive slower than the demand for them, the shortfall gets passed down to you as terms. This guide is about reading those terms early and arranging your work so a squeeze costs you a bit of money rather than a week. It is not for anyone negotiating a data centre lease or buying reserved capacity, which is a procurement problem with entirely different tools. Every price and limit below was read from the vendor’s own page on 4 September 2026, and all of them move.
Compute is a building before it is an API
The demand side of this is no longer speculative. The IEA’s Electricity 2026 report says United States electricity use “is set to add more than 420 TWh in total over the next five years”, and that data centre expansion “is expected to make up about 50% of demand growth out to 2030” [3]. That is a large share of a large number, and it has to be delivered through a physical connection that a grid operator schedules.
Texas is the clearest public example of what happens when the scheduling breaks down. On 3 August 2026, ERCOT told the state’s distribution and transmission providers that it “will not notify each Interconnecting Distribution Service Provider and Transmission Service Provider of how any Large Load is classified in the Batch Zero Interconnection Study by August 7, 2026” [1]. The notice names its own cause. That morning ERCOT had received a letter from Governor Greg Abbott directing it to run a verification process before advancing any data centre Large Loads through interconnection, and on that basis ERCOT said it would file a request for a good cause exception to the timelines and process for Batch Zero set out in Sections 5 and 9 of its own Planning Guide [1].
The queue this applies to holds more than 1,800 projects requesting over 474 gigawatts of potential demand, about 90% of it data centres. Roughly 250 to 300 projects go through the audit, and the 200 gigawatts of future demand attached to them is more than twice the peak demand record the ERCOT grid set in July 2026 [2].
The lesson is not that Texas is a bad place to build. It is the shape of the gap. Filing an interconnection request is cheap and early. Energising a site is expensive and late. The distance between those two events is where a vendor’s capacity roadmap actually lives, and it is measured in years and political decisions rather than in sprints. Any promise about compute that lands in the future is a promise about that gap closing on schedule.
A shortage reaches you as a limit, not an outage
You will not get an email saying compute is short this quarter. You will get a slightly different account.
Anthropic’s page on Claude Pro usage is unusually direct about the mechanism. The session-based message limit “will reset every five hours”, and during peak hours the plan offers “at least 5x the usage per session compared to our free service” [4]. The worked example on that page is around 45 messages every five hours for relatively short conversations on a less compute-intensive model, and the page is explicit that the real number depends on the length of your message, the length of any files you attach, how long the current conversation has run, and which model or feature you use [4]. Then comes the sentence worth writing down: “In addition, to manage capacity and ensure fair access to all users, we may limit your usage in other ways, such as weekly and monthly caps or model and feature usage, at our discretion” [4].
That is the whole argument of this guide, printed on a help page by the vendor. Your allowance is a function of somebody else’s spare capacity, and it can be adjusted without a product announcement. The phrase “peak hours” in the same document tells you the adjustment is not uniform across the day either.
The practical response is small and boring. For one month, note the date, the hour and the task every time you hit a limit. Four or five data points is enough to tell you whether your constraint is genuinely volume, or whether you are doing your heaviest work in the same two hours as everyone else in your time zone. The second problem has a free fix and the first one does not.
Every price you are planning on has an expiry date
Treat a published rate as a quote with a date on it, because increasingly the vendors write the date down themselves.
Gemini 3.8 Flash is listed at $0.75 per million input tokens “through December 31, 2026” and $1.50 “starting January 1, 2027”, with output moving from $3.75 to $7.50 on the same day [6]. That is a doubling, published in advance, in the ordinary course of business. OpenAI’s GPT-5.6 Sol sits at $4 per million input tokens and $20 per million output for short context, and the pricing page says only that this “promotional pricing is available at least through November 21, 2026” [5]. No replacement rate is published. One of those is a change you can put in a diary. The other is an open date, which is harder to plan around, not easier.
The tiering tells you as much as the headline number. On OpenAI’s pricing page, batch runs at exactly half the standard rate for that model, $2 and $10 against $4 and $20. Cached input is $0.40 against $4, a 90% discount. Fast mode doubles both, to $8 and $40, and regional processing endpoints “are charged a 10% uplift for models released on or after March 5, 2026” [5]. Google’s paid tier carries the Batch API at a “50% cost reduction”, and its per-model tables bear that out: Gemini 3.8 Flash batch input is $0.375 against $0.75 standard [6]. Read those multipliers together and the picture is clear enough. Speed is sold at a premium, patience is sold at a discount, and choosing where your tokens are processed costs extra. That is not a pricing quirk. That is a scarce resource being allocated by price, which is what happens when there is not enough of it to hand out evenly.
Model choice is the same lever at a bigger scale. On Claude’s own pricing, Haiku 4.5 is $1 per million input and $5 per million output, Sonnet 5 is $2 and $10, Opus 5 is $5 and $25, and Fable 5.1 is $10 and $50 [8]. The top and bottom of that list differ by a factor of ten. If a job is classification, extraction or reformatting, the ten-times model is not buying you anything you can measure.
The lever you own is when the work runs
You cannot change how much capacity exists. You can change which queue your work sits in, and every major vendor now sells you that choice.
Split your AI use into two piles. The first is you, in a chat window, waiting for an answer before you can do the next thing. That work is interactive by definition and none of what follows applies to it. The second is work that produces something nobody reads for hours: overnight summaries, weekly digests, bulk tagging of a backlog, first-pass drafts queued for a Monday review. That pile is deferrable, and every one of the three vendors here prices batch processing at about half the standard rate [5][6][8]. The trade is that you give up control of exactly when in the window it runs, which for work you were not going to look at until tomorrow is not a trade at all.
The second lever is repetition, and it is worth reading the caching rates carefully rather than assuming they are free money. Cached input on GPT-5.6 Sol is $0.40 per million against $4 standard, and Claude’s cache reads on Sonnet 5 are $0.20 against $2 standard input [5][8]. But both vendors charge more to write a cache entry than to send the same text uncached: $5 per million against $4 on OpenAI, $2.50 against $2 on Sonnet 5 [5][8]. Caching pays when one block of instructions goes out often enough that the cheap reads outweigh the single dear write. It does nothing at all for a prompt you send once.
The third lever is honest model selection per job, using the price ladder above. Picking one model and using it for everything is the low-friction default, and it is defensible while the bill is small. Once the bill is not small, the question worth asking is which jobs on your list ever needed the expensive end of that ladder.
Batch tiers at OpenAI, Google and Anthropic are priced at about half the standard rate. Computed in the page; nothing is sent anywhere.
Model retirement is the capacity event you will actually feel
The version of a model you built on will be switched off, and Anthropic states the reason plainly: it “currently deprecates and retires models to ensure capacity for new model releases” [7]. That is the shortage arriving as a calendar entry rather than a price.
The notice is shorter than most people expect. Anthropic commits to “at least 60 days’ notice before model retirement for publicly released models” [7], and the record shows this is routine rather than exceptional. Developers using claude-sonnet-4-20250514 and claude-opus-4-20250514 were notified on 14 April 2026 that both would retire on 15 June 2026, and users of claude-opus-4-1-20250805 were notified on 5 June 2026 of retirement on 5 August 2026 [7]. Both windows were just over 60 days, and not a day more than that.
Sixty days is comfortable if you find out on day one and know where the exposure is. It is not comfortable if you find out from a failing automation on day sixty. The fix is a list. Keep one record of every place a specific model version is named, which for a small operation is usually a few automation steps, a script or two, and a saved prompt. Make sure vendor emails about deprecations reach a human inbox you read, not a shared address nobody opens. Anthropic also lets you export your API usage by key and model from the Console to find deprecated models you have forgotten about [7]. Then, when a retirement notice arrives, you are doing a scheduled swap rather than an emergency one.
There is a design point hiding in this. A workflow that names a model version has a life measured in months. A workflow that names a capability, with the version set in one place you can edit, does not.
Dates that belong to somebody else
Whenever a vendor gives you a date, the useful question is what it is contingent on. Features shipping from existing capacity are one kind of promise. Anything that depends on new physical capacity coming online is a forecast about a queue, and the Texas sequence is a fair illustration of how those go. ERCOT published a timeline in its own Planning Guide, then said it would ask the regulator for an exception to it after the Governor’s letter landed, adding that it would consult the commission to determine next steps [1].
For a business your size, the exposure is rarely a contract. It is commitment. Do not buy an annual plan on the strength of a capability that has been announced but not shipped, because monthly billing is exactly the premium you pay for that uncertainty and it is usually worth it. Do not rebuild a client-facing process around a region, a tier or a model that is described as coming soon. And when you do adopt something new, give yourself a written answer to one question before it becomes load-bearing: if this were withdrawn or repriced next quarter, what would you actually do that week.
What still goes wrong
Every number here has a shelf life. The prices, limits and retirement dates were read from vendor pages on 4 September 2026. Google has already published the January 2027 doubling for Gemini 3.8 Flash [6], and OpenAI’s promotional rate for GPT-5.6 Sol is guaranteed only to 21 November 2026 with nothing published about what follows [5]. Re-read the pages rather than trusting this one in six months.
Diversification is the obvious hedge and it is weaker than it looks. Running two assistants splits your context, doubles the admin, and does not necessarily halve your exposure, because both vendors may be renting capacity in the same handful of regions from the same handful of suppliers. What a second account genuinely buys you is a way to keep working through one vendor’s bad afternoon. That is worth having on a free tier. It is not a strategy.
The batch and caching advice also has a narrower reach than the arithmetic suggests. Neither applies to a person sitting in a chat window waiting for an answer, and that is where a lot of one-person AI usage sits. Work out your own split before you count on the saving. If your deferrable share is 10%, halving its cost saves you 5%, which is real but will not rescue a budget. The larger saving is usually picking a cheaper model for jobs that never needed an expensive one, and the obstacle there is not price. Switching models means re-testing prompts, and nobody enjoys that week.
Finally, nobody can tell you how the grid constraint resolves. ERCOT asked for an exception to its own published Batch Zero timelines rather than replacing them with a new date [1], the audit covered 250 to 300 projects when it was announced in August [2], and a policy change could loosen or tighten that queue faster than any of the numbers above suggest. Plan for the terms of your accounts to keep moving. Do not plan on a specific direction.
- 01ERCOT — Market Notice M-A080326-01, Update Regarding Batch Zero Timelines and Processesercot.com
- 02The Texas Tribune — Texas will audit up to 300 projects, mostly data center proposalstexastribune.org
- 03IEA — Electricity 2026, Demandiea.org
- 04Anthropic Help Center — About Claude's Pro plan usagesupport.anthropic.com
- 05OpenAI — API pricingdevelopers.openai.com
- 06Google — Gemini API pricingai.google.dev
- 07Anthropic — Model deprecationsplatform.claude.com
- 08Claude — Pricingclaude.com