saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

The limit in the spec sheet is often a price

How to check whether an AI tool's context window, rate limit or message cap is a technical ceiling or a billing boundary, and what to do about each.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

GPT-5.6 Sol has a context window of 1,050,000 tokens [1]. The Codex client ships with a default of 272,000 tokens for that same model [3]. The same OpenAI page that carries the 1,050,000 figure also carries this one: “Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request” [1]. The product’s default stops at the exact number where the price doubles.

That is the pattern this guide is about, and it is not really about Codex. You read a headline number on a model page, you plan a workflow around it, and the tool you actually use enforces a different, smaller number that nobody announced. Sometimes the smaller number is an engineering constraint. Often it is a billing boundary wearing engineering clothes, which matters because engineering constraints stay put and billing boundaries move. This guide is for anyone who pays for AI by usage or by plan and has designed a piece of work around a stated limit. It is not for teams on a negotiated enterprise contract, where the limits are terms you can read and argue with, and it is not a guide to choosing a model.

Every tool has three limit numbers, and you were shown one

The advertised context window is what the model will accept through the API. The product default is what your client will use unless you change it. The product maximum is what your client will use at all. These are three different numbers and vendors publish them in three different places.

For GPT-5.6 Sol the three are 1,050,000, 272,000 and 872,000. The first comes from the model reference [1]. The other two come from the bundled metadata file inside the Codex repository, which sets context_window to 272000 and max_context_window to 872000 for GPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna alike [3]. GPT-6 Astra’s model page also states a 1,050,000 window [2]. So the client will not use the top 178,000 tokens of the model’s window under any setting, and it will not use the top 778,000 without you changing something.

None of that is hidden, exactly. It is in a JSON file in a public repository. It is also not in any place a person budgeting a week of work would think to look. The practical move is to find the third number before you design around the first, and to write down which file or page each of your three numbers came from.

Compare the default against the vendor’s own price break

The check takes about ten minutes per tool. Take the number the product enforces. Then look at the vendor’s pricing for a rate that changes at the same number.

OpenAI puts the same threshold on GPT-5.6 Sol and GPT-6 Astra. Above 272,000 input tokens, the Sol page says a prompt is “priced at 2x input and 1.5x output for the full request” [1], and the Astra page says “2x input and cache rates and 1.5x output for the full request” [2]. Either way the surcharge lands on the whole request rather than the excess. Codex’s default context window is 272,000 [3]. Google publishes the same shape at a different number: Gemini 3.1 Pro Preview costs $2.00 per million input tokens and $12.00 per million output for prompts up to 200k tokens, and $4.00 and $18.00 above that; Gemini 2.5 Pro runs $1.25 and $10.00 below the same 200k line and $2.50 and $15.00 above it [6].

When a default cap and a price break carry the same number, you do not need to prove intent. You only need to change how you treat the number. A limit that sits on a billing boundary behaves like a billing number: it will be defended when costs rise, it will move when pricing moves, and it will not be explained in engineering terms because it was never an engineering answer.

A vendor without the price break does not need the cap

Anthropic sells the same class of long-context capability with no context-length tier at all. Its documentation states that for every model with a 1M-token context window, “1M is the default: you don’t need a beta header, and long-context requests are billed at standard pricing” [5]. A 900,000-token request costs the same per token as a 9,000-token one. Eleven models carry that window, among them Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5, Claude Opus 4.6 and Claude Sonnet 5, while other models, “including Claude Sonnet 4.5, have a 200k-token context window” [5]. Output is capped at 128k tokens per request on the 1M models [5].

That is the useful comparison. Two vendors selling long context to the same buyers, and one of them has a hard number at 272,000 while the other does not have that number anywhere. The gap is a commercial choice about how to sell long context, and it tells you the shape of the ceiling is negotiable in a way a token limit sounds like it is not.

Use this when you compare tools. The honest comparison is the effective window at the price you will actually pay, on the client you will actually run, not the two headline figures side by side. A model advertised at 1,050,000 tokens that your client caps at 272,000 [1][3] is, for your purposes, a 272,000-token model with an expensive upgrade path.

Numbers set by pricing move quietly, in files not announcements

On July 18, 2026, a one-file pull request titled “Backport refreshed bundled model metadata to 0.144” merged into the Codex repository’s release/0.144 branch [4]. It changed context_window and max_context_window for gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna from 372000 to 272000 [4]. One JSON file, 64 lines added and 54 removed [4]. The models themselves did not change, and the model reference page still states 1,050,000 today [1].

The metadata file on the main branch now sets those models to a 272,000 default with an 872,000 maximum [3], so the number has moved more than once and in more than one direction inside a few months. Nothing about that is unusual or improper. It is just a different kind of change from the kind you get an email about.

The operational consequence is that “it truncated earlier than it did last month” is a configuration question before it is a model question. Pin your client version if a workflow depends on the limit. Know which file holds the number for the tools you rely on, and check it after an update instead of re-tuning a prompt that was never the problem.

Consumer plans run the same mechanism with the numbers removed

If you use ChatGPT or Claude through a subscription rather than an API key, the pricing boundary is still there. It is just expressed as something less countable.

OpenAI’s plan documentation lists Free at $0 per month, Go at $8, Plus at $20, and Pro from $100, where you “Choose 5x or 20x higher rate limits than Plus” and the 20x option is the $200 tier; Business is $20 per user per month billed annually, or $25 per user per month billed monthly [8]. The page then says: “The estimates below show local messages per five-hour period” [8]. On Plus, GPT-6 Astra is estimated at 5 to 45 messages and GPT-5.6 Sol at 10 to 100 [8]. A range where the top is 9 times the bottom is not a message count. It is a compute budget being reported in messages, and it varies with how long your messages are.

The credits table underneath says the quiet part in the right unit. GPT-6 Astra is billed at 250 credits per million input tokens, 25 per million cached input tokens, and 1,250 per million output tokens [8]. That is the real meter. The message estimate is a translation of it.

Google’s API tiers are the least disguised version of this. The free tier asks for an “Active project or free trial”; Tier 1 asks you to “Set up and link an active billing account”; Tier 2 asks for “Paid $100 + 3 days from first successful payment”; Tier 3 asks for “Paid $1,000 + 30 days from first successful payment” [7]. On top of the usual requests and tokens per minute, the three paid tiers carry a rolling 10-minute spend cap of $10, $50 and $200 respectively, and crossing it returns a “429 RESOURCE_EXHAUSTED” error [7]. Not one criterion in that ladder is technical. It is money and elapsed time, published as a rate limit.

Working with a limit once you have decided it is a price

Design to the default rather than the maximum. The default is what the tool does on a machine you have not configured, which includes every teammate’s machine and every fresh install. If the workflow only survives at the maximum, it is a workflow with a setup step, and the setup step will be skipped.

Know the escape hatch and price it properly. Codex will go to 872,000 tokens [3], and above 272,000 input tokens OpenAI charges 2x input and 1.5x output on the full request [1]. That last phrase is the trap. A 300,000-token prompt is not billed as 272,000 tokens at the normal rate plus 28,000 at double. The whole request reprices. Crossing the line by a little costs almost as much as crossing it by a lot, which means the right strategy is usually to stay well under or go well over, not to sit next to it.

Leave headroom for the next move. The last change to the Codex default was 100,000 tokens [4]. If a number that governs your work sits on a price break, budget as though it can move by roughly the size of the last move, in either direction, without notice. Then track cost per finished task rather than cost per token, because a smaller window that forces three passes over a document is more expensive than a larger window that does it once, and the per-token price will never tell you that.

checklist
Before you build on a stated limit
0 of 8 · saved in this browser only
calculator
What crossing the long-context line costs
$ / month extra

The surcharge only, on top of standard rates. Assumes 2x input and 1.5x output applied to the full request, and defaults to GPT-5.6 Sol's $4 and $20 [1]. Computed in the page; nothing is sent anywhere.

What still goes wrong

Matching numbers are evidence, not proof. A vendor can land on 272,000 for serving reasons and price at the same line because that is where serving gets expensive, which is a perfectly ordinary way for a business to behave. The reason to run the check anyway is that the two cases produce identical advice: treat the number as mutable, keep headroom, and do not build a customer promise on top of it. Run the check to calibrate your planning, not to write an accusation.

Every figure here was read off a vendor page or a repository file on September 5, 2026, and this topic moves faster than most. One of the numbers in this guide changed by 100,000 tokens in a single commit [4] and has since changed again [3]. Use the method and re-read the sources; do not quote the figures back at anyone in six months.

Absence of a price tier is not the same as cheap. Anthropic bills its 1M-token window at standard rates [5], and you still pay for every token you put into it. A larger window mostly removes a constraint on your design, and removing that constraint tends to increase what you send. Cheaper per token and cheaper per task are different measurements, and only the second one shows up on the invoice in a way you can act on.

Finally, the consumer plans cannot really be audited. “5 to 45 messages” [8] is not a number you can staff a day around, and there is no meter in the interface that reconciles it. If your work depends on a guaranteed volume rather than a probable one, the honest answer is that you need the API, where the unit is a token and the limit is a number you can read.

sources
  1. 01OpenAI — GPT-5.6 Sol model referencedevelopers.openai.com
  2. 02OpenAI — GPT-6 Astra model referencedevelopers.openai.com
  3. 03openai/codex — bundled model metadata (models.json, main)raw.githubusercontent.com
  4. 04openai/codex — Backport refreshed bundled model metadata to 0.144 (PR #33972)github.com
  5. 05Anthropic — Context windowsplatform.claude.com
  6. 06Google — Gemini API pricingai.google.dev
  7. 07Google — Gemini API rate limitsai.google.dev
  8. 08OpenAI — ChatGPT pricing and usage limitslearn.chatgpt.com
next guide
Giving an AI agent write access without losing your work
9 min · verified 2026-09-05
related guides