saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

What a compute shortage actually does to your AI bill

Find the four places where scarce electricity and hardware reach your invoice, and cut what you pay using the discounts vendors already publish.

Published 2026-09-04 · Updated 2026-09-04 · Read 10 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

Your bill went up and nothing you changed explains it. Maybe an introductory rate quietly expired. Maybe your documents got longer and crossed a threshold where input costs double [3]. Maybe a job that used to finish started returning errors in the afternoon and you assumed the model had got worse. Behind every one of those is a supply chain of turbines, transformers, buildings and grid connections that you will never touch, that nobody can hurry, and that is currently the tightest part of the whole industry.

This guide is about the short distance between that physical constraint and the invoice in front of you. Scarce compute does not reach a small operator as a blackout. It reaches you as four things published on a vendor’s own website: a price, a modifier sitting next to the price, a rate limit, and an expiry date. Each of them has a lever attached, and none of the levers require you to have an opinion about semiconductors. This is written for a solo operator or small team spending tens or hundreds of dollars a month on API calls and subscriptions. If you are negotiating reserved capacity, running your own GPUs, or signing a contract with a committed spend, your constraints are different and this is too light an instrument for them.

The constraint is electricity, and its clock runs in years

Data centres used around 415 terawatt-hours of electricity in 2024, about 1.5% of world consumption, and the International Energy Agency’s base case has that more than doubling to around 945 TWh by 2030 and rising to around 1,200 TWh by 2035 [6]. In advanced economies, data centres account for more than 20% of electricity demand growth to 2030; globally they account for around one-tenth of it [6].

The number that actually explains your invoice is smaller and duller. The IEA reports that “around 20% of planned data centre projects could be at risk of delays” [6]. The reasons given are not about chips. Building new transmission lines can take four to eight years in advanced economies, wait times for critical grid components such as transformers and cables have doubled in the past three years, and turbine deliveries for new gas-fired power plants now face lead times of several years [6].

That is the whole mechanism. Demand for inference can double in a year. A substation cannot. When supply is set by manufacturing lead times rather than by willingness to pay, the gap gets closed by rationing, and a company that sells you tokens rations in the only two ways it has: what it charges, and how fast it will let you go.

None of this means your unit costs are heading up. On the same page where Anthropic lists its current models, the retired Claude Opus 4.1 sits at $15 per million input tokens and $75 output, while Claude Opus 5 is $5 and $25, and Claude Sonnet 5 is $2 and $10 against Sonnet 4.6’s $3 and $15 [1]. Per-token prices have been falling across generations even while the physical build-out strains. Both things are true, and confusing them is how people end up making infrastructure decisions based on a mood.

The price is the small number; the modifiers beside it are the big one

Open any vendor pricing page and the headline table is the least interesting part. The money is in the multipliers printed underneath, because those are the surfaces where capacity gets priced.

Long context is the clearest example, and vendors have split on it. OpenAI charges GPT-6 Astra at $10 per million input tokens and $50 output, but $20 input and $75 output once a request qualifies as long context [3]. GPT-5.6 Terra goes from $2 and $12 to $4 and $18 the same way [3]. Google prices Gemini 3.1 Pro Preview at $2 input for prompts of 200,000 tokens or under and $4 above that [2]. Anthropic has gone the other direction and states that Claude 4.6 and later models include the full 1 million token context window at standard pricing, with a 900,000-token request billed at the same per-token rate as a 9,000-token one [1]. If your prompts have been growing, that difference is worth more than any comparison of headline rates, and it is invisible unless you go and read.

Speed is priced separately too. Anthropic’s fast mode for Opus 5 and 4.8 costs $10 input and $50 output against the standard $5 and $25, exactly double [1]. OpenAI’s fast mode also doubles standard pricing [3]. Where you process is priced as well: Anthropic applies a 1.1x multiplier on all token categories for US-only inference [1], and OpenAI adds a 10% uplift for regional processing on models released on or after 5 March 2026 [3]. Those are all capacity charges wearing different names. You are paying for a scarcer version of the same computation.

The last modifier is a date. Google lists Gemini 3.8 Flash at $0.75 input and $3.75 output “through December 31, 2026”, and $1.50 and $7.50 starting 1 January 2027 [2]. OpenAI’s GPT-5.6 Sol at $4 and $20 is marked as promotional pricing through 21 November 2026 [3]. An introductory rate is a real price and a temporary one. If a model’s economics only work at the promotional number, you do not have a working system, you have a countdown, and the honest move is to put the expiry date in your calendar the day you adopt the model rather than discovering it on an invoice.

Half a token bill is a scheduling decision, not a purchasing one

The single largest discount available to a small operator has nothing to do with negotiating and nothing to do with switching vendors. It is agreeing to wait.

All three major vendors price asynchronous work at half. Anthropic’s Message Batches API states that “all usage is charged at 50% of the standard API prices”, which puts Claude Sonnet 5 at $1 input and $5 output [1][8]. OpenAI’s Batch API applies a 50% discount across its flagship models, on input, cached input and output alike [3]. Google’s batch mode does the same, putting Gemini 3.8 Flash at $0.375 and $1.875 through the same end-of-2026 date [2]. The catch is smaller than it sounds. Anthropic documents most batches completing within 1 hour, with results available when all messages have completed or after 24 hours, whichever comes first, and batches expiring if processing does not complete inside that window [8]. A batch is capped at 100,000 requests or 256 MB, whichever is reached first, and results stay downloadable for 29 days [8].

So the practical question is not whether you can afford the model. It is which of your jobs actually needed an answer while you watched. Overnight summarisation, bulk classification, backfilling a dataset, running your own evaluations, generating a month of drafts: none of that needs a live connection, and all of it is half price. Anthropic’s own documentation says as much, listing large-scale evaluations, content moderation, data analysis and bulk content generation as the intended cases [8].

Caching is the second half of the same idea, and it rewards structure rather than patience. Anthropic prices a 5-minute cache write at 1.25x the base input rate and a 1-hour write at 2x, then charges cache reads at 0.1x, falling to 0.025x on Claude Fable 5.1 and Mythos 5.1 [1]. OpenAI lists cached input for GPT-6 Astra at $1.00 against $10.00 standard [3]. Google prices context caching for Gemini 3.8 Flash at $0.075 per million tokens through the end of 2026 [2]. If you send the same instruction block, style guide or reference document on every call, you are paying full rate for text the model already read this minute. Putting the stable material at the front of the prompt, ahead of the part that changes, is what makes it cacheable.

The rate limit and the service tier are where rationing is written down

Price is the surface people watch. The limit is the surface that stops your work, and it arrives without warning because it is not on the invoice.

Anthropic’s API places organisations in usage tiers with monthly spend caps: Start at $500, Build at $1,000, Scale at $200,000, and Custom with no cap [5]. Placement is automatic, “based on usage history and account standing”, and new organisations may start in an evaluation tier with limits below the standard ones while account history is established [5]. Within a tier, three separate limits apply per model: requests per minute, input tokens per minute and output tokens per minute. On the Start tier, Claude Opus 5 is documented at 1,000 RPM, 2,000,000 ITPM and 400,000 OTPM [5]. Capacity replenishes continuously under a token bucket algorithm rather than resetting at fixed intervals [5].

Two details in that page are worth more than the rest. First, hitting the spend cap returns a 429 with no retry-after header, and “API usage pauses until 00:00 UTC on the first day of the next month, unless you request a higher limit sooner” [5]. That is a hard stop, not a slowdown, and it is the failure most likely to take a small operation offline for real time. Second, cached input tokens do not count toward your input-token-per-minute limit on most models, so the same caching that cut your bill also raises your effective ceiling. Anthropic’s own example is a 2,000,000 ITPM limit with an 80% cache hit rate behaving like 10,000,000 total input tokens per minute [5].

The service tier page is where scarcity is stated in plain language. Anthropic runs three tiers. Standard is the default, and the API prioritises those requests “alongside all other requests with best-effort availability” [4]. Priority tier prioritises requests over all others and “helps minimize ‘server overloaded’ errors, even during peak times”, but is “available only to organizations with an existing capacity commitment”, and the page adds that “Priority Tier capacity commitments are no longer available for purchase” [4]. Requests beyond committed capacity automatically fall back to standard tier [4]. Read that as what it is. The guaranteed lane is closed to new entrants, and if you are reading this you are on best effort. Design for the afternoon when best effort is thinner than usual, which mostly means retrying on 429 with the retry-after header, and moving anything that can wait into the batch lane so it is not competing with the work you are watching.

The orbital bets are a thermometer, not a product you can buy

When a constraint gets expensive enough, well-funded people start attacking it with physics, and the seriousness of the attempts tells you how hard the constraint is.

Google Research has published a design study, Project Suncatcher, envisioning compact constellations of solar-powered satellites carrying Google TPUs and connected by free-space optical links [7]. The reasoning is entirely about the same bottleneck the IEA describes: in the right orbit, Google writes, a solar panel can be up to 8 times more productive than on Earth and produce power nearly continuously, reducing the need for batteries [7]. Google’s analysis of historical and projected pricing puts launch prices at less than $200 per kilogram by the mid-2030s, and its next milestone is a learning mission with Planet, slated to launch two prototype satellites by early 2027 [7]. Its own summary is careful: “significant engineering challenges remain”, naming thermal management, high-bandwidth ground communications and on-orbit system reliability [7].

There is nothing here to act on, and that is the point. You cannot buy orbital inference, the prototypes are not the product, and the cost projection is for the middle of the next decade. What the study is good for is calibration. A company with access to as much terrestrial power and land as anyone alive is publishing serious work on getting its compute above the atmosphere, because grid connections, siting and cooling on the ground are hard enough to make that worth studying. Treat announcements in this category as a reading on how tight terrestrial capacity is, and then go back to the pricing page, which is where the tightness will actually reach you.

Portability is what turns a price change into a config edit

Everything above assumes the vendor moves and you respond. The cheap version of responding is having built the ability to move before you need it.

Keep the model name, the base URL and the API key in environment variables rather than inline, so a switch is one edit in one place. Keep a working key with a second provider even if you never send it traffic, because opening an account under time pressure is how people end up accepting whatever tier they are given. Keep 20 real tasks from your actual work, with the outputs you were happy with, and run them through the cheaper model before you assume you need the expensive one. The published price gap is wide enough to be worth an afternoon: Gemini 3.5 Flash-Lite lists at $0.30 input and $2.50 output [2], GPT-5.6 Luna at $0.20 and $1.20 [3], and Claude Haiku 4.5 at $1 and $5 [1], against flagship rates of $10 input and $50 output on both Claude Fable 5.1 and GPT-6 Astra [1][3].

Then price a real month before you decide anything. Take your actual token volumes off the vendor’s usage console, not from memory, and multiply. Most alarming price changes turn out to be a few dollars on a small operation, and the ones that are not become obvious in about a minute. Do this before you migrate, because a migration you did not need is paid for entirely in your own time.

checklist
Before you accept next month's AI bill
0 of 8 · saved in this browser only
calculator
What moving work to the batch lane saves
$ / month saved

Monthly spend at standard rates, times the share you can send asynchronously, times the 50% batch discount. Defaults are Claude Sonnet 5's published standard rates. Computed in the page; nothing is sent anywhere.

What still goes wrong

The biggest limit of this approach is that it only works on the parts vendors have chosen to publish. A price, a multiplier and a documented notice period are commitments you can read. Capacity is not. Nothing on any of these pages tells you how much headroom exists behind your best-effort request on a Tuesday afternoon, and the service tier page is explicit that standard means best-effort availability rather than a guarantee [4]. You will occasionally be slow or overloaded for reasons that never appear in writing, and no amount of reading prepares you for that beyond having retries and a second provider.

Batch and caching also have a cost that does not appear in the price table, which is design work. Moving a job to the batch endpoint means restructuring it around a submission and a later collection, handling the 24-hour expiry when a batch does not complete [8], and accepting that a workflow which used to be one function is now two. Caching means keeping prompts in a stable order, which quietly constrains how you compose them. Both are worth it at volume and are not worth it at ten calls a day. If your monthly bill is $6, the correct action after reading this is to close the tab.

The other thing that goes wrong is timescale confusion, which is the specific trap this topic sets. Transmission lines and gas turbines on multi-year lead times [6], satellite prototypes flying by early 2027 [7], and launch costs projected for the mid-2030s [7] all describe a decade. Your batch endpoint, your cache hit rate and your rate limit describe this week. It is tempting to treat the first set as the reason to do something about the second, and the two are barely connected. The physical constraint is real, it is slow, and the only part of it you can act on is which lane you send your own work down.

sources
  1. 01Anthropic — Pricingplatform.claude.com
  2. 02Google — Gemini API pricingai.google.dev
  3. 03OpenAI — API pricingdevelopers.openai.com
  4. 04Anthropic — Service tiersplatform.claude.com
  5. 05Anthropic — Rate limitsplatform.claude.com
  6. 06IEA — Energy and AI, executive summaryiea.org
  7. 07Google Research — Exploring a space-based, scalable AI infrastructure system designresearch.google
  8. 08Anthropic — Message Batches APIplatform.claude.com
next guide
What custom AI chips mean for what you pay
9 min · verified 2026-09-04
related guides