How to read an AI lab's capability claim
Turn an unfalsifiable headline about AI capability into a small test you can run on your own work, and know when to ignore it entirely.
on this page · 0 / 0 checked
On 3 September 2026 OpenAI shipped a model called GPT-6 Astra, and the company’s president, Greg Brockman, told Axios “Welcome to the AGI era”, saying he thought it “might be about this model” [8]. If you run a business with fewer than ten people, your Tuesday did not change. It rarely does. But claims like that keep arriving, and every one of them carries an implied instruction: rethink the plan, the staffing, the pricing, now.
What follows is a method for turning that kind of announcement into a decision in about twenty minutes. Usually the decision is to do nothing. Occasionally it is to spend an afternoon running a small test. This is not for people who evaluate frontier models professionally, and it is not a forecast about whether AGI is close. It is for the person who has to decide whether a headline changes what they do on Thursday, and who has no evaluation harness and no intention of building one.
AGI is a definition someone picked, not a line a machine crosses
Start with the word, because it does less work than it appears to. OpenAI’s own Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work” [1]. Every load-bearing term there is a choice. “Most” is a threshold somebody sets. “Economically valuable work” is a category somebody draws. “Outperform” needs a task and a scoring rule, and the choice of both sits with whoever is making the claim.
You can watch the choosing happen in the commercial paperwork. When OpenAI and Microsoft restructured their partnership on 28 October 2025, the announcement said that once AGI is declared by OpenAI, “that declaration will now be verified by an independent expert panel”, and extended Microsoft’s IP rights over models and products through 2032 [2]. That is a definition with money attached and a procedure for settling arguments about it. Reasonable for a contract term. Odd for a scientific fact.
The money was attached quite literally. The October announcement tied the revenue share to the finding: “The revenue share agreement remains until the expert panel verifies AGI, though payments will be made over a longer period of time” [2]. Six months later that link was cut. On 27 April 2026 the two companies announced that Microsoft’s license would become non-exclusive, and that revenue share payments from OpenAI to Microsoft would continue through 2030 “independent of OpenAI’s technology progress, at the same percentage but subject to a total cap” [3]. The financial consequence that had been hanging on an AGI declaration was written out.
Take the lesson rather than the gossip. A term whose consequences two companies rewrote twice in six months is not a measurement. When a lab’s leadership says the era has arrived, they are exercising judgment, and the careful ones admit it. No formal declaration accompanied this launch. Brockman said he personally believes OpenAI has reached AGI, and left users to decide whether Astra meets the definition [8]. Treat “AGI” in a headline as a mood. The data is somewhere further down the page.
Split the headline from the claim you could check
Underneath nearly every unfalsifiable headline sits a narrower claim that somebody could test. Find it, and ignore everything above it.
A checkable claim has three parts: a named task, a success rate, and somebody outside the company who could run the same thing. The launch post for that same model carries several. It says Astra “saturates FrontierMath Tier 4 with a 98% score” and “saturates ARC-AGI-3 with a 99.9% score”. On Terminal-Bench 4.0 it reports 57.9%, against 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1. On OSWorld 2.0, a computer-use suite, it reports 72.6% at roughly 40 minutes per task, up from 65.7% at roughly 75 minutes [4].
Put those side by side and the headline gets quieter. The same system scores 99.9% on a benchmark with AGI in its name and 57.9% on a suite of terminal tasks. That gap is not a contradiction. It is the entire picture. Problems with clean, gradeable answers saturate first. Work that runs through a messy environment, a real filesystem and a long chain of dependent steps does not.
So rank the reported benchmarks by how closely they resemble your work, and read only the closest one. If your day is drafting and checking documents, a coding suite tells you almost nothing. If your day is a long sequence of small actions inside software you did not build, the computer-use number is yours. Read its label before you read its value. OpenAI reports that OSWorld 2.0 result on an offline set as a partial score [4], which credits progress through a task rather than counting clean finishes, so 72.6% does not mean that roughly three attempts in four came out finished and correct. It means roughly a quarter of the available credit went unearned, at about 40 minutes per attempt. Whether that clears your bar depends on what a half-done attempt costs you to inherit.
Shipped, priced and documented beats previewed
Claims about internal, unreleased systems cannot be checked by anyone, including the journalists who print them. Before a model ships there is nothing to test, only a description of a thing that exists on someone else’s servers. The correct response to a preview is to note it and wait.
Three artifacts mark the moment a claim becomes checkable: an availability line, a price, and a system card. The availability line is more restrictive than the headline almost every time. Astra was “rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users” [4]. On launch day, the thing announced was not the thing most subscribers had.
The price is the part that converts a capability claim into a budget line. Astra’s API pricing is $10 per million input tokens and $50 per million output tokens [4]. A model that works through a 40-minute computer-use task [4] is producing a great many output tokens while it does so, and the bill scales with the autonomy, not with the impressiveness. Any claim of the form “it can now do a whole day of work” is also a claim about a day’s worth of tokens. Until there is a price and a plan name attached, you are reading about someone else’s property.
The safety document is the honest half of a launch
A launch post selects. A system card is the document a company writes knowing that regulators, review boards and its own safety staff will read it, which gives it a reason to record what went wrong. Read the limitations before the results.
Astra’s card, published the same day, states that “Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework” [5]. It also says the model “shows a substantial decrease in chain-of-thought monitorability compared to previous models” and is “significantly more able to control its own chain-of-thought” [5]. In plain terms: the reasoning trace is a less reliable record of what the model actually did. If any part of your review process is “read the model’s explanation and sanity-check it”, that part just got weaker, and no launch chart would have told you.
The card also does something launch posts never do, which is warn you off its own numbers. It notes that because the evaluations focus on difficult cases and long-tail risks, “their results should not be interpreted as estimates of how frequently these behaviors occur in typical production use” [5]. That sentence is the vendor telling you the benchmark does not transfer to your traffic. Believe it, in both directions. A frightening number on an adversarial suite is not your Tuesday, and neither is a reassuring one.
Independent measurements exist, and the error bars are the finding
There is a small number of outfits that measure models without selling them, and their numbers are worth more than any launch post. METR’s time horizon is the most useful single figure: the length of task, measured by how long humans take, that a model completes at a given success rate. In METR’s March 2025 paper the measured length had been doubling approximately every 7 months for the preceding 6 years [6]. METR now marks parts of that post as outdated and points readers to its newer results, which is itself the behaviour you want from a measurement outfit [6].
The January 2026 update is where the caution lives. METR expanded its suite from 170 to 228 tasks and doubled the number of tasks lasting 8 hours or more, from 14 to 31 [7]. At a 50% success rate it put Claude Opus 4.5 at 320 minutes, GPT-5 at 214 minutes, and GPT-4o at 6 minutes [7]. The trend is real and steep. Then read the brackets: Opus 4.5’s interval runs from 170 to 729 minutes, and METR notes it measured human baseline times for only 5 of its 31 long tasks, with the rest estimated [7].
Two things follow. First, “50%” means half the attempts fail at that length, so the honest translation of a 320-minute figure is “it sometimes gets through a day’s work and sometimes does not”. Second, an interval that spans from under 3 hours to over 12 is telling you the measurement is young. When a number arrives with an interval, the interval is the finding. Anyone quoting the midpoint without the brackets has taken a research result and turned it into a slogan.
A folder of ten real tasks beats any benchmark you did not run
The only evaluation that changes what you should do is one built from work you have already done. Pull ten tasks out of the last quarter, ones you finished and were paid for, where you still have both the input and the answer you shipped. Strip your answer out, keep it as the grading key, and write a one-line yes/no rule for each.
Run each task three times, because the same prompt does not reliably give the same output twice, and one lucky run has convinced a lot of people of a lot of things. Grade by hand. Ten tasks and three runs is small enough to finish in an afternoon and large enough to separate “useful on my work” from “not useful on my work”, which is the only distinction you need.
Decide the threshold before you run it. Write down, in advance, the result that would make you change how you work and the result that would make you change nothing. Done afterwards, this exercise becomes a search for confirmation of whatever the headline put in your head.
tasks × runs × (tokens + your grading time). The tokens are the cheap part. Computed in the page; nothing is sent anywhere.
What still goes wrong
This method is slow and it is biased towards inaction. If a capability turns out to be real and broadly available, you will be a few weeks behind the people who moved on the announcement. Most of the time that is the cheap error, because most announcements do not survive contact with real work. Occasionally it is the expensive one, and nothing in this guide will tell you which kind you are looking at in advance.
Your own evaluation is small, and small evaluations are noisy. Ten tasks run three times will not distinguish two models a few points apart, and it will not catch a failure mode that shows up once in fifty runs, which is exactly the sort that hurts when a task runs unsupervised for 40 minutes. Benchmarks have the opposite problem: they saturate, and once a suite is old enough that frontier models cluster near the ceiling, a new record tells you very little. OpenAI’s own card says it believes HealthBench, now more than a year old, “is approaching a noise ceiling for frontier models” [5]. When a number stops moving, you cannot tell from outside whether progress stopped or the ruler ran out.
And the documents are still vendor documents. A system card is more honest than a launch post because of who reads it, not because of who writes it. The Charter definition [1] and the expert-panel procedure [2] were both drafted by parties with a direct interest in the answer, and the clause that made the answer expensive was later renegotiated away [3]. Independent measurement is the only real check, there is not much of it, and it lags every release by months. Assume you are always reading about the world as it was one model ago.
Prompts from this guide
eval-set-from-your-own-work
Below are {n} pieces of finished work I produced, each with the input I
started from. Turn them into a test set. Do not attempt the tasks and do
not improve my answers.
For each piece, write:
1. The task, stated as an instruction, with no reference to my answer
2. The inputs needed, quoted from the material below
3. A grading rule I can apply in under a minute, phrased as a yes/no
check against my answer
4. One thing a wrong answer would most likely get wrong
If a piece of work depends too heavily on context you cannot see, say so
and skip it rather than guessing.
---
{finished_work_with_inputs} - 01OpenAI — Charteropenai.com
- 02OpenAI — The next chapter of the Microsoft–OpenAI partnershipopenai.com
- 03OpenAI — The next phase of the Microsoft partnershipopenai.com
- 04OpenAI — GPT-6 Astra: A new generation of intelligenceopenai.com
- 05OpenAI — GPT-6 Astra System Carddeploymentsafety.openai.com
- 06METR — Measuring AI ability to complete long software tasksmetr.org
- 07METR — Time Horizon 1.1metr.org
- 08Axios — OpenAI releases new model GPT-6 Astra, says it may represent AGIaxios.com