saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

How to read an AI lab's safety report

Read a frontier lab's safety report the way it deserves, tell a verified claim from a self-graded one, and pick vendors on things you can actually check.

Published 2026-09-04 · Updated 2026-09-04 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

A vendor’s safety page is the easiest document to read and the hardest one to check. It has ratings, a framework with a version number, a list of capability thresholds, and a tone of institutional care. You skim it, decide the company looks responsible, and move a client workflow onto its model. Nothing in that sequence is unreasonable. It is just that the document you treated as evidence was written, scored, published and signed off by the company being assessed.

This guide is about reading those documents properly: what a frontier lab’s safety report genuinely establishes, which sentence in it is worth more than the headline rating, and what to check instead when the decision carries real money. It is written for a solo operator or small team deciding where to put paid work — drafts, code, client files, automations. It is not a compliance procedure. If you are putting AI into hiring, credit, insurance, medical or legal decisions, you need a risk process and someone accountable for running it, and NIST’s generative AI profile is a better starting point than this page [7].

The rating is the vendor’s own homework

Every frontier lab now publishes a safety framework, and they are more serious than most people assume. Anthropic’s Responsible Scaling Policy is on version 3.4, effective 8 July 2026, and states that reaching certain Capability Thresholds requires the company to upgrade its safeguards to the ASL-3 Security Standard or the ASL-3 Deployment Standard; the same document commits to an Executive Risk Council “sponsored by executive leadership to oversee security programs” [2]. OpenAI’s Preparedness Framework, version 2, last updated 15 April 2025, tracks three categories — biological and chemical, cybersecurity, and AI self-improvement — each with a High threshold covering “capabilities that significantly increase existing risk vectors for severe harm” and a Critical threshold for “capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent” [4]. Google DeepMind’s Frontier Safety Framework reached version 3.1 on 17 April 2026, adding Tracked Capability Levels alongside its Critical Capability Levels [5].

These are detailed, versioned, and revised in public. They are also administered end to end by the company they govern. The lab defines the threshold, builds the evaluation, runs it, scores it, interprets the score, and decides whether to ship. In OpenAI’s framework, an internal Safety Advisory Group “reviews the Capabilities Report and decides on next steps” [4]. That is not a scandal, and it is not hypocrisy. Nobody outside these companies currently has the weights, the compute and the access to do the job instead. But it does fix what a published rating can mean to you: it is a company’s considered opinion of itself, released voluntarily, on a scale it invented.

The model you can buy is not the most capable model that exists

Labs run internal-only models, and they say so. Anthropic’s August 2026 risk report, whose coverage date is 15 July 2026, describes a system it calls Model 2, “which is somewhat more capable than Mythos 5”, and offers a “rough qualitative sense” that it is “a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview” [1]. The same passage states its status plainly: “We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities” [1].

The useful reading of that is not drama. Internal-only deployment is a normal stage in the pipeline now, and a model that has not finished predeployment testing is exactly the kind of thing that should not be on sale. The consequence for you is narrower and more practical: the frontier described in coverage is not the frontier you can build on. What you can build on has a price, a version string and a system card. Today that means Claude Opus 5 at $5 per million input tokens and $25 per million output, Claude Sonnet 5 at $2 and $10, and Claude Mythos 5.1 and Fable 5.1 at $10 and $50 [8]. Those you can budget, rate-limit and swap. Capability that exists only inside a lab is not a plan, however impressive it sounds when described.

The most useful sentence in a safety report is the one about its own limits

Skip the headline rating and go looking for the hedges, because that is where the information is. Anthropic’s August 2026 report raised its rating for misalignment in high-stakes settings from “very low” to “low”, and gave as the reason “general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations” rather than any single model failing a specific test [1]. Read alone, that is a small move on an unfamiliar scale.

Read next to a different rating in the same document, it means more. On automated research and development, the report holds at “low” and then undercuts its own instrument in the following sentence: “we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have ‘saturated’—i.e., no longer capture increases in models’ capabilities—and because we are seeing early signs of acceleration” [1]. The report is similarly plain about measuring its own research speedup: “We believe our internal AI R&D efforts are significantly faster than they would be without AI assistance, but not yet by a factor of 2 (though we are uncertain and measurement is difficult)” [1]. NIST makes the same point at framework level, and it applies to every lab, not one: “Some GAI risks are unknown, and are therefore difficult to properly scope or evaluate given the uncertainty about potential GAI scale, complexity, and capabilities” [7].

So the technique is inverted from the obvious one. The rating tells you about the vendor’s candour. The hedges tell you about your risk. A lab that publishes the saturation of its own benchmarks is behaving better than one that quietly retunes its methodology and prints the same word. It also tells you what these ratings are made of: on at least one threat model the measuring instrument has stopped registering the thing it measures, and the published label reads the same either way.

External testing exists, and it is smaller than the word suggests

System cards are the most useful safety document a lab publishes, because they are about one shipped model rather than a policy. The Claude Opus 5 system card covers a model dated 24 July 2026, deployed under the same ASL-3 protections as Claude Opus 4.8, and it names the outside parties involved: the UK AI Security Institute, given access to early checkpoints and assessing the model on three agentic cyber ranges; Dyno Therapeutics on two sequence-to-function evaluations; Trajectory Labs at “roughly 100 hours red-teaming our safeguards”; 10a Labs at “around 16 hours testing a variety of attack techniques”; Grayswan running an automated attacker at 150 attempts per task; Irregular, whose CyScenarioBench tests multi-stage cyber operations; and Mozilla, on an evaluation of exploit development in Firefox 147 [3]. That is more independent scrutiny than almost any software you buy receives. The disclosed effort is also counted in hours, roughly 100 from one firm and around 16 from another, against a system that will sit inside customer workflows for as long as it is deployed, and the lab chose the evaluators, scoped the work and published the summary.

One line on that card should change how you read the capability numbers printed next to it: “As elsewhere in this card, production safety interventions are disabled during evaluation” [3]. The published scores describe the raw model, not the product you talk to, which is the raw model plus filters, classifiers and refusals. The card also warns where its instruments stop, saying its automated evaluations “may not capture the risk posed by improvements in general capabilities supporting biological research productivity” [3]. OpenAI’s framework makes a comparable commitment in comparable terms, saying that “when available and feasible, OpenAI will work with third-parties to independently evaluate models”, and that it will work with third parties to evaluate safeguards where a deployment warrants that and high quality testing exists [4].

When a salesperson tells you a model has been independently tested, the follow-up questions are short. Which model version. Who did it. For how many hours. Where is the result published. If the answer to any of those is a gesture at a framework rather than a document, you have been shown a policy, not a test.

Regulation gives you a floor, not a verdict

The EU AI Act now puts real obligations on providers of general-purpose models with systemic risk, and they took effect on 2 August 2025. Article 55(1) requires a provider to “perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing”; to “assess and mitigate possible systemic risks at Union level”; to report serious incidents to the AI Office “without undue delay”; and to “ensure an adequate level of cybersecurity protection for the general-purpose AI model with systemic risk and the physical infrastructure of the model” [6].

Look closely at what those are. They are process obligations: evaluate, document, report, secure. They are worth having and they are not a certificate. None of them states that a given model is fit for your workflow, and none of them audits the sentence on a vendor’s marketing page. “AI Act compliant” in a proposal means the provider carries a duty, not that anyone checked your use of the model. That distinction matters most when you are the one repeating the claim to a client, because you inherit it. NIST names the position you are in exactly: value chains include third-party components that “might be improperly obtained or not properly vetted, leading to diminished transparency or accountability for downstream users” [7]. You are the downstream user, and the accountability lands with you.

Choose on what you can check

Almost everything that decides whether an AI vendor is a good bet for your business is checkable, and none of it is the safety rating. The exact model version you call, and whether a system card exists for it rather than for the family name. The price, published, in a place you can revisit [8]. The retention and training terms attached to the tier you are actually on, which are frequently different from the ones you read about. Whether the vendor names outside testers and publishes what they found [3]. Whether it publishes its own bad news, like a saturated benchmark, or only its good news [1].

Then run the only evaluation whose scope you control. Take 20 real tasks from your own work, with the real inputs, keep the outputs you consider correct, and re-run the set when you change model or the vendor ships a new version. It is not a safety test and it does not pretend to be. It measures the thing you are actually exposed to, which is whether this model does your work to your standard today, and it catches the quiet regression that no framework will ever tell you about. Put your own numbers into the calculator below before you decide it is too expensive to bother with; at list prices it rarely is, so the reason people skip it is not the money. Price your exit at the same time. If moving your main workflow to a second vendor would take a week, you have not chosen a vendor, you have acquired a dependency.

checklist
Before you depend on a vendor's safety claim
0 of 8 · saved in this browser only
calculator
What your own regression set costs to run
$ / year

Claude Opus 5 list prices, $5 per million input tokens and $25 per million output [8]. Computed in the page; nothing is sent anywhere.

What still goes wrong

You cannot independently evaluate a frontier model, and nothing in this guide changes that. Reading system cards well makes you a better-informed customer, not an auditor. Your 20-task set measures whether a model is good at your work; it says nothing about misuse, misalignment, or the failure modes the labs write reports about. The honest summary is that on the questions those reports address, you are relying on the vendor, and the most you can do is prefer the vendors that tell you where their own instruments are failing.

The reading itself does not scale to the release cadence. Anthropic’s pricing page currently lists five Opus versions available at once, alongside Sonnet, Mythos and Fable models [8], and the August 2026 risk report closes its coverage nine days before the Claude Opus 5 system card is dated [1][3]. Realistically you will read one system card per vendor per year, while the model behind your workflow changes several times in that year. The practical compromise is to read the card once properly when you commit, keep your regression set running for everything after, and re-read only when the vendor changes the model family or you change what the workflow is worth.

This guide is also the wrong document for two groups. If you need assurance rather than judgement — a regulated deployment, a contractual safety commitment, a customer promise you would have to defend — no public safety report currently supplies it, and the correct answer is usually to keep that decision off a model rather than to read harder. And if you are hoping a rating change signals something you should trade, hire or restructure around, it does not. A lab moving one qualitative rating by one notch, on its own scale, with its own instruments, and saying in the same document that it is less confident than last time [1], is a company being careful in public. That is a good sign about the company. It is not a measurement of your exposure.

sources
  1. 01Anthropic — Redacted Risk Report, August 2026www-cdn.anthropic.com
  2. 02Anthropic — Responsible Scaling Policyanthropic.com
  3. 03Anthropic — Claude Opus 5 System Cardwww-cdn.anthropic.com
  4. 04OpenAI — Preparedness Framework, Version 2cdn.openai.com
  5. 05Google DeepMind — Strengthening our Frontier Safety Frameworkdeepmind.google
  6. 06EU AI Act — Article 55, Obligations for providers of general-purpose AI models with systemic riskartificialintelligenceact.eu
  7. 07NIST AI 600-1 — Artificial Intelligence Risk Management Framework: Generative AI Profilenvlpubs.nist.gov
  8. 08Anthropic — Model pricingplatform.claude.com
next guide
How much to trust a brand-new AI feature
9 min · verified 2026-09-04
related guides