How to read an AI lab's safety claims
Work out what a lab's safety framework actually promises, why it never covers your failure modes, and which controls you have to set yourself.
on this page · 0 / 0 checked
A new model lands and the announcement reads well. It was evaluated against the lab’s safety framework, tested by internal and external red teams, and shipped with new refusals in place. Somewhere near the top is a sentence about this being the most carefully evaluated model the lab has shipped. Then you have to decide something the post never addresses: whether to give it your inbox, your client folder, and a credential that can send email under your name.
That gap is structural rather than dishonest. Frontier safety frameworks measure whether a model could help someone build a weapon or run an autonomous cyber operation. Your question is whether it will overwrite the wrong row in a spreadsheet or send a half-finished draft to a client who has already complained once. Both matter. Only one of them has a published rubric, and it is not yours. This guide is for solo operators and small teams deciding how much access to hand a capable model. If you run critical infrastructure or employ security staff, go to the frameworks and system cards directly; this is the version for people with neither the time nor the regulatory obligation.
Every frontier safety framework is written and graded by the company it constrains
OpenAI’s Preparedness Framework tracks 3 capability areas, described as “Biological and Chemical capabilities, Cybersecurity capabilities, and AI Self-improvement capabilities”, and sorts findings into 2 thresholds. High capability could “amplify existing pathways to severe harm”; Critical capability could “introduce unprecedented new pathways to severe harm” [2]. A system at High needs safeguards that minimise the associated risk before deployment, and a system at Critical needs them during development as well [2]. The judgment sits with the Safety Advisory Group, described as a “cross-functional team of internal safety leaders”, whose “guidance goes to OpenAI Leadership for final decisions” [2].
Anthropic’s Responsible Scaling Policy does the same work with different vocabulary. Reaching certain Capability Thresholds “requires us to upgrade our safeguards to the ASL-3 Security Standard or the ASL-3 Deployment Standard” [1]. Google DeepMind’s Frontier Safety Framework defines Critical Capability Levels as “capability levels at which, absent mitigation measures, frontier AI models or systems may pose heightened risk of severe harm”, and commits to safety case reviews “prior to external launches when relevant CCLs are reached” [3].
Notice how much these documents move. Anthropic’s policy is at version 3.4, effective 8 July 2026, and describes itself as a living document [1]. Google published the third iteration of its Frontier Safety Framework on 22 September 2025, then updated it on 17 April 2026 to FSF 3.1, adding Tracked Capability Levels to help it spot “potential less extreme risks sooner” [3]. A rubric its own author revises this often is not a standard. It is a policy, and the same organisation writes it, applies it, scores itself against it, and amends it when the score is inconvenient.
That is not a reason to ignore them. A written commitment you can quote back is worth more than no commitment, and all 3 labs have published enough detail to be embarrassed by later. It is a reason to read them as a description of an internal process, not as a warranty attached to a model.
The thresholds are sized for catastrophe, not for your Tuesday
Look at what the thresholds actually cover. Anthropic’s capability thresholds include AI research and development, framed as “the ability to fully automate entry-level AI research work, and the ability to cause dramatic acceleration in the rate of effective scaling”, and chemical and biological capability that could “substantially uplift the development capabilities of moderately resourced state programs” [1]. Google’s Critical Capability Levels concern “heightened risk of severe harm” and now include a level for harmful manipulation [3]. OpenAI’s 2 tiers are both defined by their relationship to “severe harm” [2].
None of that is a claim about whether the model will hallucinate a clause into a contract, misread a PDF invoice, or take an instruction from a web page it was asked to summarise. A model can be comfortably below every published threshold and still be the reason a client receives the wrong file. The frameworks and your operational risk are measured on different instruments, and only one of the two gets a public score.
So when a release post says a model was rated safe to deploy, translate it: a committee inside the company concluded that this model does not clear their bar for helping produce mass-casualty harm, given the mitigations they built. That is a real statement. It is also not about you.
Only a narrow slice of this is law, and it binds the vendor rather than you
There is one place where safety language stops being a blog post and becomes an obligation. The EU AI Act’s rules on general-purpose AI models became effective in August 2025, and from 2 August 2026 “the AI Office and authorities of the Member States are responsible for implementing, supervising and enforcing the AI Act” [4]. The AI Office “can request technical documentation, evaluate models, require corrective measures and issue fines for non-compliance” [4].
Article 55 sets out what providers of models with systemic risk have to do. They must “perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing”; assess and mitigate possible systemic risks at Union level; “keep track of, document, and report, without undue delay, to the AI Office” information about serious incidents and corrective measures; and “ensure an adequate level of cybersecurity protection for the general-purpose AI model” and its physical infrastructure [5]. Compliance can be demonstrated through approved codes of practice or harmonised standards, or by showing adequate alternative means for the Commission to assess [5].
Read the limits of that carefully. The duties fall on the model provider, the risk in question is systemic and assessed at Union level, and the reporting goes to a regulator rather than to you. It raises the floor on documentation, which is genuinely useful when you want evidence instead of adjectives. It says nothing about whether a model connected to your accounting system will behave.
When one company sells both the capability and the remedy, the assessment is not independent
The pattern worth learning shows up whenever a lab publishes a risk and a product for that risk in the same week. On 3 September 2026 OpenAI announced Daybreak for Frontline Defenders, committing “$1 billion in subsidized Daybreak access” to help “resource-constrained cyber defenders, starting with the United States”, and targeting it “to be consumed over the next six months” [8]. The named beneficiaries are water and wastewater systems, electric grid operators, state and local governments, community and regional banks, nonprofits and open-source maintainers. Access splits into a Daybreak Blue tier that “supports common defensive work with our mainline models” and a Daybreak Red tier giving approved organisations “specialized cyber models for more sensitive and technically demanding work” [8].
Take that at face value and it is a real offer to organisations that cannot hire a security team. It is also a document produced by the company that built the offensive capability the programme defends against, denominated in credits against its own products, gated by its own eligibility rules. Both readings are correct at once, and holding both is the durable skill.
Generalise it: when the same organisation defines the threat, grades the severity, and sells the mitigation, its assessment carries a commercial interest in a particular conclusion. That is not an accusation of bad faith, and it applies to every lab equally. It is a reason to want a second reader before you act on the assessment. Your candidates are the regulator’s documentation requirements, an independent evaluator, or a test you run yourself on your own data. The third is slower than the first two and the only one you fully control.
The system card is where the caveats live
All 3 labs produce something more detailed than the release announcement, and the announcement is not the document you want. OpenAI says it will “continue to publish our Preparedness findings with each frontier model release” [2]. Anthropic’s Risk Reports “aim to provide a direct, candid, and informative description of how we see the risks of our systems and our state of preparedness for them” [1]. Google commits to safety case reviews, described as “detailed analyses demonstrating how risks have been reduced to manageable levels”, before external launches at relevant capability levels [3].
Those documents are longer, duller, and far more useful. Look for 3 things in one. First, what was actually tested, and on which version or checkpoint of the model. Second, whether a mitigation is training-time refusal behaviour or a control enforced outside the model, which is the difference between a habit and a rule. Third, the sentences where the lab says it could not establish something. A line beginning “we were unable to rule out” tells you more than the entire summary above it.
Give one of these 20 minutes before a model touches anything that matters, and skip it entirely for a model you only use to rewrite paragraphs. The point is proportion, not diligence theatre.
The controls sized to your risk are the ones you set
Everything above is context. The part that changes your exposure is what you grant, and the guidance there is unglamorous and well documented. The Model Context Protocol’s security best practices name the failure directly: broad tokens create “expanded blast radius”, where “stolen broad token enables unrelated tool/resource access”, and list “using wildcard or omnibus scopes (*, all, full-access)” among the common mistakes [7]. The recommended pattern is a “progressive, least-privilege scope model” that starts from a “minimal initial scope set” containing “only low-risk discovery/read operations” and elevates only when privileged operations are first attempted, with elevation events logged [7].
For a small team that translates into 4 moves. Connect read-only first and stay there until something breaks. Use a separate account with its own credentials rather than your own login, so revoking access does not mean locking yourself out. Put an approval step in front of any action that spends money, sends something externally, or deletes. And keep a log you can actually read afterwards, because an audit trail nobody opens is a file, not a control.
Reviewing every action does not scale, and the calculator below prices that in hours rather than in good intentions. Decide up front which class of action gets a human, and accept that the rest runs unwatched.
The last piece is responsibility, and it is already assigned. Anthropic’s Usage Policy, effective 15 September 2025, requires that in high-risk domains including legal, healthcare, insurance, finance, employment, housing, academic testing and media, “a qualified professional in that field must review the content or decision prior to dissemination or finalization”, and that you tell end users you are using AI to help produce your advice or decisions [6]. It then states plainly: “You or your organization are responsible for the accuracy and appropriateness of that information” [6]. No safety framework transfers that sentence to the vendor.
actions × share × seconds × 22 working days, converted to hours. Computed in the page; nothing is sent anywhere.
What still goes wrong
The frameworks can move faster than your understanding of them. Anthropic’s policy reached version 3.4 in July 2026, and Google’s framework has had 3 iterations plus an April 2026 update [1][3], which means a threshold you read about last year may have been redefined, split, or folded into a new category since. There is no notification when a rubric changes underneath a claim you already acted on, so a safety assessment you read once has a shelf life you have to guess at.
Independent verification barely exists at the level you need. The EU AI Act creates documentation and incident-reporting duties owed to the AI Office rather than to customers, and covers systemic risk at Union level [4][5]. Nothing in that regime tells a 2-person agency whether a particular model, connected to a particular inbox, with a particular set of scopes, is a sensible risk. That judgment stays with you, and the evidence available to make it is mostly written by the party selling the thing.
And least privilege is a ceiling, not a fix. Narrow scopes constrain what goes wrong, they do not prevent it [7], and a read-only connection can still leak. If the consequence of a bad action is genuinely unacceptable, such as an irreversible payment, a filing with a deadline, or a message to a client relationship you cannot repair, the correct control is not a tighter scope. It is not connecting that system at all until you have watched the model work on a copy for long enough to know how it fails.
- 01Anthropic — Responsible Scaling Policy (v3.4)anthropic.com
- 02OpenAI — Our updated Preparedness Frameworkopenai.com
- 03Google DeepMind — Strengthening our Frontier Safety Frameworkdeepmind.google
- 04European Commission — AI Act regulatory frameworkdigital-strategy.ec.europa.eu
- 05EU AI Act — Article 55, obligations for providers of GPAI models with systemic riskartificialintelligenceact.eu
- 06Anthropic — Usage Policyanthropic.com
- 07Model Context Protocol — Security Best Practicesmodelcontextprotocol.io
- 08OpenAI — Daybreak for Frontline Defendersopenai.com