saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

Read your vendor's deprecation page, not its funding announcements

Infrastructure deals rarely move your AI bill, but model retirement dates will break your scripts, so here is how to check your exposure and cut it.

Published 2026-09-04 · Updated 2026-09-04 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

Every few months a chip company and a model company announce something with a number in it larger than a mid-sized country’s annual budget, and the coverage tells you it changes everything. Then you go back to the thing you were actually doing, which is an invoicing summariser that runs against one model, four times a day, for about eleven dollars a month. The honest question is whether that eleven dollars is still eleven dollars next year, and whether the script still runs. Infrastructure announcements are a bad place to look for the answer.

There is a good place to look, and almost nobody reads it. This guide is about telling the two apart: which parts of the AI buildout actually reach a one-person or five-person business, and which are noise you can skip. It is not for anyone evaluating semiconductor equities, negotiating colocation contracts, or sizing a training run. It is for someone with three or four workflows wired into a model API or a paid assistant seat who wants to know what to watch, what to ignore, and what a bad week would cost.

The Ohio announcement says less than the headline number does

On 17 August 2026, OpenAI said it had entered an agreement to secure approximately 8 gigawatts-IT at the PORTS-Pike Technology Campus in Pike County, Ohio, on remediated land controlled by the Department of Energy at the former Portsmouth Gaseous Diffusion Plant [1]. The structure is worth reading slowly, because it is not a purchase. SB Energy will build, own and operate the data center under a 20-year lease to OpenAI [1][2]. NVIDIA will invest $1.5 billion in SB Energy and provide credit support on land, power and shell buildout to secure the initial 4.25 IT-GW, with an option to take the remaining 3.75 [2]. SB Energy and SoftBank will invest at least $4.2 billion in new regional grid infrastructure through a partnership with AEP Ohio [2]. The first 800 megawatts are expected to become available in 2028, and the project runs a six-year buildout through 2032 [1].

Three parties, three different kinds of exposure. The landlord owns the concrete. The chip vendor owns equity and a credit backstop, which is a standing obligation rather than a completed sale. The model company owns a twenty-year lease and a first delivery date two years out. Notice what is not in either announcement: any statement about what an API call costs in 2029. The nearest thing to a number aimed at users is $84 million in Codex credits through ChatGPT for Ohio college students, plus an incremental $40 million on top of SB Energy’s existing $40 million community benefits fund [1][2].

That gap is the point. Capacity deals set how much compute exists in five years. They do not set your unit price next quarter, and they contain no promise that any particular model stays reachable. The financing pages and the pricing pages are different documents, written by different teams, on different clocks. Reading the first to predict the second is how people end up either panic-buying annual commitments or assuming prices are frozen forever.

The volatility that reaches you is a model name with an expiry date

Here is what actually breaks a small operation. Your workflow does not depend on “OpenAI” or “Anthropic” in the abstract. It depends on a string like gpt-5-2025-08-07 or claude-sonnet-4-20250514, and that string has a shutdown date on a public page [3][4].

OpenAI states a minimum notice period: “Unless safety or compliance concerns require a faster timeline, we provide the following minimum notice periods before model retirement: Generally available models: At least 6 months” [3]. Specialised variants of generally available models get at least three months, and preview models “may be retired with much shorter notice, such as 2 weeks” [3]. The published schedule is specific. The Videos API and the sora-2 and sora-2-pro models shut down on 24 September 2026, announced on 24 March 2026. A batch of legacy snapshots including gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-turbo and o1-2024-12-17 goes on 23 October 2026. Then gpt-5-2025-08-07, o3-2025-04-16 and o3-pro-2025-06-10 go on 11 December 2026, notified on 11 June 2026 [3]. That last one is the six-month floor honoured to the day.

Anthropic publishes the same kind of page with a shorter commitment: “Anthropic notifies customers with active deployments for models with upcoming retirements, providing at least 60 days’ notice before model retirement for publicly released models” [4]. That page shows Claude Sonnet 3.7 and Claude Haiku 3.5 retired on 19 February 2026, Claude Haiku 3 on 20 April 2026, Claude Opus 4 and Claude Sonnet 4 on 15 June 2026, and Claude Opus 4.1 on 5 August 2026 after a notice dated 5 June 2026 [4]. Sixty-one days. Requests to models past the retirement date fail [4].

Google’s Gemini changelog shows a third pattern, which is that preview models go on much less warning. The entry dated 30 July 2026 announced that gemini-robotics-er-1.6-preview would be shut down on 31 August 2026, which is 32 days [5]. Its predecessor, gemini-robotics-er-1.5-preview, was announced on 14 April 2026 for shutdown on 30 April 2026, which is 16 [5]. And gemini-omni-flash-preview, released in public preview on 30 June 2026, was announced on 27 August 2026 as being deprecated on 30 September 2026 in favour of the generally available gemini-omni-1.1-flash [5]. Three months old at deprecation, with a month’s notice.

So the notice you are entitled to ranges from six months down to a fortnight, depending on which vendor you picked and which tier of model you wired in. That range is knowable today, for free, from three public pages. It is a far better predictor of your next unplanned work weekend than any financing structure in Ohio.

Published prices move in both directions, and they carry dates

The other half of the worry is price. Read the three rate cards side by side and the picture is more interesting than a straight line down.

Within a tier, replacements have been cheaper than what they replace. Claude Opus 4.1 was $15 per million input tokens and $75 per million output; it retired on 5 August 2026 with claude-opus-4-8 named as its replacement [4], and Opus 4.8 is $5 and $25 [7]. Claude Sonnet 5 is $2 and $10 where Claude Sonnet 4.6 is $3 and $15 [7]. On Google’s card, Gemini 3.5 Flash is $1.50 and $9.00 while the newer Gemini 3.8 Flash is $0.75 and $3.75 [8].

The top of the range moves the other way. On OpenAI’s card, GPT-5.6 Sol is $4.00 per million input tokens and $20.00 per million output, with cached input at $0.40. GPT-5.6 Terra is $2.00 and $12.00, and GPT-5.6 Luna is $0.20 and $1.20. The newer GPT-6 Astra sits above all of them at $10.00 and $50.00 [6]. A new flagship is not a discount.

Then the dates, which are printed on the cards and mostly ignored. OpenAI notes that GPT-5.6 Sol’s promotional pricing is available at least through 21 November 2026 [6]. Google lists Gemini 3.8 Flash at $0.75 and $3.75 through 31 December 2026, then $1.50 and $7.50 starting 1 January 2027 [8]. That is a published doubling, already on the page you are reading the cheap number from. It can go the other way too: Anthropic’s $2 and $10 for Claude Sonnet 5 was announced as introductory pricing through 31 August 2026, and the page now records that the scheduled increase to $3 and $15 will not occur [7]. Treat every rate card as a quote with an expiry, and go and read the expiry.

If cost is your problem, the fix is almost never a vendor migration. All three cards price asynchronous work at half the standard rate. Anthropic states a 50% discount on both input and output tokens [7], Google lists the Batch API at a 50% cost reduction [8], and OpenAI’s batch column is exactly half its standard column across the flagship range [6]. A cheaper tier, prompt caching and the batch endpoint will beat a rewrite, and none of them require you to trust a forecast.

What the public record does not contain is any statement from any of these vendors about what a token costs in 2029. The Ohio documents are about capacity and the rate cards are quotes with dates on them, and nothing bridges the two. The cheapest way to be wrong about that gap is to keep the cost of moving low.

What a forced migration actually costs

The number worth having is not your monthly bill. It is the one-off cost of being told, with 60 days of notice, that the model your workflows are pinned to stops answering [4]. That cost is mostly your own hours: rewriting prompts that were tuned against the old model, re-running whatever checks you have, and fixing the two or three outputs that quietly changed shape.

Work it out once, honestly, and the answer usually lands somewhere between an annoying Tuesday and a fortnight you did not budget for. Either answer is useful. If it is the first, stop worrying about vendor risk and go do your actual job. If it is the second, the fix is not to pick a better vendor, it is to reduce the number of hours a migration takes.

calculator
Cost of a forced model migration
one-off cost

workflows × hours each × your rate. Computed in the page; nothing is sent anywhere.

Pin the version, keep the eval set, hold a second key

The habits that cut that number are unglamorous and take an afternoon. Pin explicit dated model names rather than floating aliases, so an upgrade happens when you choose it rather than mid-run. Keep ten to twenty real inputs with the outputs you consider correct, saved as a plain file, so that checking whether the new model still works is a twenty-minute job instead of a week of vibes. Keep prompts in files rather than pasted into a vendor’s web console, because portable text is the entire difference between switching vendors and rebuilding.

Do not put anything load-bearing on a preview model. OpenAI says preview models may go with as little as two weeks’ notice and does not recommend them for business-critical production workloads unless you can migrate on short notice [3], and Google’s own robotics previews went with 32 and 16 days [5]. Preview is for experiments you would not mind losing.

Finally, keep a second vendor’s key working on one low-stakes workflow. Not as a hedge against a company failing, which is unlikely and slow, but as a way of knowing, before you are forced to find out, whether your prompts survive contact with a different model. The mid tiers on all three cards sit close together, at $2.00 and $12.00 for GPT-5.6 Terra, $2 and $10 for Claude Sonnet 5, and $0.75 and $3.75 for Gemini 3.8 Flash [6][7][8], so a genuine second option costs very little to keep alive.

checklist
Quarterly vendor-risk pass
0 of 8 · saved in this browser only

What still goes wrong

The deprecation pages tell you the date, not the difference. A replacement model can be live, cheaper and better on benchmarks while still producing subtly different output in the one format your downstream spreadsheet expects. Anthropic’s page names a recommended replacement for each retired model [4], and that recommendation is a good starting point, not a guarantee that your prompts transfer unchanged. This is the entire reason the eval file matters, and it is also why the file needs to hold your real inputs rather than tidy examples.

Notice periods are floors, not promises about calm. OpenAI’s own wording carves out safety and compliance concerns as grounds for a faster timeline [3], and Anthropic’s 60-day commitment covers publicly released models [4]. Neither commitment covers a capability changing inside a model you still have access to, which is the more common and much harder-to-detect failure. If your workflow depends on a specific behaviour rather than a specific endpoint, no policy page protects you and the eval file is the only warning system you get.

And the honest limit on the financing story: nobody outside these companies can tell you what compute costs in 2030. The Ohio buildout will not deliver its first 800 megawatts until 2028 and runs through 2032 [1], which is a longer horizon than most small businesses plan on. Watching those announcements is entertainment. Watching your own vendor’s retirement schedule, and keeping the cost of moving low enough that you would not flinch, is the part you can actually control.

sources
  1. 01OpenAI — OpenAI joins PORTS-Pike projectopenai.com
  2. 02SB Energy — PORTS-Pike Technology Campus to exclusively host NVIDIA AI computesbenergy.com
  3. 03OpenAI — API deprecationsdevelopers.openai.com
  4. 04Anthropic — Model deprecationsplatform.claude.com
  5. 05Google — Gemini API changelogai.google.dev
  6. 06OpenAI — API pricingdevelopers.openai.com
  7. 07Anthropic — Pricingplatform.claude.com
  8. 08Google — Gemini Developer API pricingai.google.dev
next guide
How to read an AI vendor's adoption number
9 min · verified 2026-09-04
related guides