How to read a big AI capability claim
Big AI claims arrive most months, and this is how to separate the part you can grade from the part that costs nothing to say.
on this page · 0 / 0 checked
Every few months someone who runs a frontier lab says something enormous. The takeoff has started. We are past the point of no return. The next model changes everything. It lands on your phone as a headline, you feel something for about a minute, and then you go back to the invoice you were writing. Nothing about your Tuesday changed. A week later a client forwards you the same headline and asks whether they should be worried, and you find that you have no way to answer that is not a guess about a stranger’s sincerity.
This guide gives you a way to read those statements that does not require you to decide whether the person is honest. It works by sorting one statement into parts: the part that has a date or a number attached, the part that somebody is contractually accountable for, and the part that costs nothing to say. It is written for solo operators and small-team owners who use these tools for ordinary work. It is not a forecasting method, it is not an investment thesis, and it will not tell you whether the singularity is near. If you are trying to decide how to price an option on a lab’s future, this is the wrong document.
The claim and the commitment live in different documents
Sam Altman’s essay “The Gentle Singularity” opens with the line “We are past the event horizon; the takeoff has started” [1]. There is no threshold attached to that sentence, no test that could fail it, and no consequence for the author if it turns out to describe nothing. It is a sincere sentence, probably, and it is also free.
The same company publishes a document where sentences are not free. OpenAI’s Preparedness Framework names three Tracked Categories — Biological and Chemical, Cybersecurity, and AI Self-improvement — and defines two thresholds inside each. High capability means “capabilities that significantly increase existing risk vectors for severe harm.” Critical capability means “capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent” [2]. Crossing those lines triggers obligations: a system that reaches High must have safeguards that sufficiently minimize the associated risk of severe harm before it is deployed, and a system that reaches Critical needs safeguards during development, whether or not anyone plans to deploy it [2].
Anthropic keeps the equivalent document, its Responsible Scaling Policy, currently version 3.4, effective 8 July 2026. It hangs an ASL-3 Security Standard and an ASL-3 Deployment Standard on capability thresholds, and the July 2026 revision specifically changed the automated AI research and development threshold to better track the threat model of concern [3]. That is a company quietly rewriting the line it has promised not to cross, in public, with a version number.
So when a big claim arrives, the first move is not to assess the claim. It is to find out which kind of document it came from. An essay, a podcast and an earnings call carry no cost when wrong. A safety framework, a policy version and a system card carry a launch delay when wrong, which is why they are written in much more careful language and revised much more often.
A forecast has a date, a vibe does not
The same essay that opens with an unfalsifiable sentence also contains three you can grade. “2025 has seen the arrival of agents that can do real cognitive work; writing computer code will never be the same.” “2026 will likely see the arrival of systems that can figure out novel insights.” “2027 may see the arrival of robots that can do tasks in the real world” [1].
Those have years on them. They are hedged, and the hedges matter: “will likely” and “may” are doing defensive work, and a prediction that arrives pre-defended should be weighted accordingly. But they are still gradeable in a way that “past the event horizon” never will be. The 2025 claim is the strongest of the three because it is the only one stated flatly, and it is also the one closest to something you can check against your own week.
The practical habit is unglamorous. Keep one file. When a lab leader makes a public statement, paste in only the sentences that contain a year, a number, a named capability or a named product, and write the date each one becomes checkable. Ignore everything else. In eighteen months that file is worth more than any amount of commentary, because it tells you, from evidence you gathered yourself, how much to discount the next statement from the same person. Nobody else is going to keep this record for you, and the outlets that reported the original claim are unlikely to report the grade.
Somebody eventually has to write the threshold as a number
The reason vague claims persist is that nothing forces them to become specific. Regulation forces it. Under Article 51 of the EU AI Act, a general-purpose AI model is presumed to have high impact capabilities, and therefore systemic risk, when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25 [4].
That number is a crude proxy for capability and everyone involved knows it. Training compute is not intelligence, and a smaller model trained better can beat a larger one. But it has the property that matters here: it is checkable by a third party, it does not depend on anyone’s self-report about how historic the moment feels, and obligations follow automatically from it. Someone had to pick a number because the alternative was a law that meant whatever the regulated party said it meant.
Apply the same test to any claim you are handed. Ask what number, if anyone bothered to write it down, would settle the question. If a number exists and the claim is above or below it, you have a fact. If no number exists and nobody is trying to write one, you are looking at a description of a mood. That is not the same as the claim being false. It means the claim is not yet the kind of thing that can be right.
The measurement has error bars and the speech does not
There is a small industry of people trying to make capability claims measurable, and their output is far more useful than the claims themselves, mostly because it comes with qualifiers.
METR measures what it calls the time horizon at 50% success: the length of task, measured by how long a human professional takes to do it, that a frontier model agent completes half the time. Its widely cited finding is that this length had been doubling approximately every 7 months for the last 6 years, with 2024 and 2025 data alone suggesting a faster pace of 1 to 4 doublings per year [5]. In that analysis Claude 3.7 Sonnet sat at a horizon of about one hour, models succeeded on close to 100% of tasks that take a human under 4 minutes, and success dropped below 10% on tasks over about 4 hours [5].
Read what the qualifier is doing. A 50% success rate means half the attempts fail, so a doubling of the horizon is not a doubling of work you can hand over unsupervised. And the post itself now carries a banner saying that some of its text and figures are out of date, pointing readers at revised results [5]. A number that flags its own staleness is behaving correctly. A speech never does that.
The measurements have their own problems, which is the second reason to hold them loosely. A 2025 study of agentic benchmarks found that SWE-bench Verified uses insufficient test cases and that TAU-bench counts empty responses as successful, and concluded that issues of this kind can lead to under- or overestimation of agents’ performance by up to 100% in relative terms; applying the authors’ checklist to CVE-Bench reduced performance overestimation by 33% [6]. When the instruments used to measure progress can be off by a factor like that, a claim built on top of them inherits the error, and a claim built on nothing at all inherits everything.
The pages that actually change your week
While the discourse runs on adjectives, the documents that decide what you can build run on digits, and they change constantly without anyone making a speech about it.
OpenAI’s API price list currently shows standard short-context rates of $10 per million input tokens and $50 per million output for gpt-6-astra, $4 and $20 for gpt-5.6-sol, $2 and $12 for gpt-5.6-terra, and $0.20 and $1.20 for gpt-5.6-luna. Batch processing halves both directions, long-context requests double the input rate and add half again to the output rate, and regional processing endpoints for data residency carry a 10% uplift for eligible models released on or after 5 March 2026 [7]. The page also commits to a date: Sol’s promotional pricing is available at least through 21 November 2026 [7].
Anthropic’s list shows Claude Opus 5 at $5 and $25, Sonnet 5 at $2 and $10, Haiku 4.5 at $1 and $5, and Fable 5.1 and Mythos 5.1 at $10 and $50, with the Batch API taking 50% off both directions and a cache hit billed at 0.1 times the base input price, or 0.025 times on Fable 5.1 and Mythos 5.1 [8]. The same page records a price rise that was called off: Sonnet 5’s $2 and $10 introductory rate is now the standard price, and the increase to $3 and $15 previously scheduled for 1 September 2026 will not occur [8].
Those numbers decide whether the automation you have been considering is worth building this quarter. A promotional rate with an expiry date tells you when to re-run your costing. A cancelled increase tells you a budget you had written off is still good. A cache discount tells you to restructure a prompt rather than shop for a cheaper model. None of it appears in any essay about the takeoff, and all of it is directly actionable by someone running a business on a laptop. Ten minutes a month on vendor pricing and changelog pages will change more of your decisions than ten hours of interviews.
The only test that resolves it is what you would do differently
The last step takes one minute and settles most arguments. Assume the claim is true. Write down, in a sentence, what you would do differently on Monday. Then assume it is false and write the same sentence.
For nearly every claim of the “we are past the event horizon” type, the two sentences are identical, and that identity is the answer. You would use the same tools, at the same prices, with the same verification habits, for the same work. The claim is real information, but it is information about the speaker and the company’s messaging, not about your operations.
Occasionally the two sentences differ, and then you have something worth acting on. A price change alters what you build. A deprecation date alters your migration schedule. A capability threshold crossed in a safety framework alters what an agent is allowed to touch in your business. Notice that all three of those arrive in documents with dates and numbers, not in the headline. The headline is how you learn there is a document. The document is the thing you read.
METR measured about one doubling every 7 months over six years, with 2024–25 data suggesting 1 to 4 per year. Three doublings from a 1 hour horizon is an 8 hour horizon. Computed in the page; nothing is sent anywhere.
What still goes wrong
This method grades statements, not futures. A claim can be entirely unfalsifiable and also entirely true, and sorting it into the “costs nothing to say” pile does not make it wrong. People who were right early about AI capability were mostly working from intuitions they could not operationalise at the time. The method protects you from acting on noise; it does not give you foresight, and anyone selling you foresight is making a claim that fails its own test.
The gradeable half is shakier than it looks too. “Systems that can figure out novel insights” has a year attached and no agreed test, so in 2027 the same sentence will be scored as a hit by some people and a miss by others, and both will be arguing in good faith [1]. Measurement lags reality by months, benchmark design flaws can move reported performance by up to 100% in relative terms [6], and even the careful sources revise themselves and say so [5]. Treating a measured number as ground truth is a smaller error than treating a speech as one, but it is still an error.
The pricing pages have a trap of their own, and it is the one I fell into while writing this. A vendor page shows the same model at four different rates, one column each for standard, batch, flex and fast processing, and it is easy to read the cheap column and quote it as the price. Check which row you are on before you build a costing on it.
The honest limit is that none of this tells you when to change your business, only when a specific statement fails to justify changing it. That decision still comes from your own work: what you have automated, what broke, what a month of tokens costs you, and what your clients will accept. If you are looking for permission to ignore the discourse entirely, this guide is close to giving it, with one exception. Read the pricing pages, the deprecation notices and the safety frameworks, because those are the places where a company writes something it can be held to.
- 01Sam Altman — The Gentle Singularityblog.samaltman.com
- 02OpenAI — Preparedness Framework v2cdn.openai.com
- 03Anthropic — Responsible Scaling Policyanthropic.com
- 04EU AI Act — Article 51, classification of general-purpose AI models with systemic riskartificialintelligenceact.eu
- 05METR — Measuring AI ability to complete long tasksmetr.org
- 06Zhu et al. — Establishing Best Practices for Building Rigorous Agentic Benchmarksarxiv.org
- 07OpenAI — API pricingdevelopers.openai.com
- 08Anthropic — Claude pricingplatform.claude.com