How to check an AI vendor's safety claim
A three-pile method for AI vendor claims, so you can tell in 20 minutes which ones you can verify yourself and which are only reputation.
on this page · 0 / 0 checked
Sooner or later a slide or a sales email tells you that the AI product you are about to put your client files through has been “reviewed”, “certified”, “independently red-teamed”, or “aligned with federal guidance”. You are one person, or five. There is no security team down the hall to hand it to, no vendor-risk questionnaire, no lawyer on retainer. There is you, 20 minutes, and a decision about whether the thing goes near real work.
The useful move is not to become an auditor. It is to sort. Almost every safety claim an AI vendor makes falls into one of three piles, and two of those piles can be checked from a browser in a few minutes for free. The third cannot be checked by you at all, and the only mistake that really costs you is treating the third pile like the first two. This guide is for a solo operator or a small team buying its own tools. If you have a procurement function, a security questionnaire and the leverage to demand a report, you are doing a different and better job than this one.
Sort the claim before you try to evaluate it
Pile one is published. The vendor has put the evidence itself on the open web and you can read it today: a system card, a data-use page, a subprocessor list, terms of service, a scaling policy. You do not have to trust a summary of it, because the document is right there and it is dated.
Pile two is attested. An outside party looked at something and wrote a report. You cannot read the report from a search engine, but it exists, it names the auditor, it covers a stated scope and period, and you can usually request a copy under an NDA. SOC 2 reports and ISO certifications live here.
Pile three is asserted. The vendor says a review happened. The criteria are not public, there is no report you can request, and the only thing reaching you is the vendor’s own description of the outcome. “Extensively red-teamed by security experts.” “Trained on ethically sourced data.” “We work with the government on model safety.” Each of these might be entirely true. None of them is evidence in your hands, because there is nothing you or anyone outside the vendor can check.
Sorting takes about a minute per claim and it decides everything that follows. A published claim you verify. An attested claim you ask for. An asserted claim you note, and then price at zero when you are deciding whether your client data goes through the product.
The federal review is an asserted claim by design
The clearest live example of pile three is the United States government’s own frontier-model review, and it is worth understanding because vendors will reference it.
Executive Order 14409, “Promoting Advanced Artificial Intelligence Innovation and Security”, was signed on June 2, 2026, and set a 60-day deadline for a benchmarking process for AI models [1]. On August 4, 2026, officials briefed the finished framework to Meta, Nvidia, Microsoft, OpenAI, Anthropic and a variety of smaller companies; under it, labs have up to 30 days to submit a frontier model to the government before public release [2]. Participation is voluntary, and the order is explicit that “nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models” [1].
The framework’s contents were not published. Fortune reported that “the details will be kept under wraps, only known to a select group of companies that may choose to participate in the process, which is voluntary”, and quoted Chris McGuire, senior fellow for China and emerging technologies at the Council on Foreign Relations, calling the secrecy “baffling”: “We can’t have secret, voluntary rules to regulate the most important tech in the world” [2].
Read that as structure rather than scandal. Whatever the review’s merits, a vendor saying “our model went through federal review” hands you a sentence and nothing else. You cannot see what was tested, whether it touched anything resembling your use case, how a pass differs from a fail, or whether the same company declined to submit a different model. The claim is not false. It is simply not the kind of claim that can do work for you, and any weight you give it is weight you are giving to the vendor’s reputation, which you already had before the sentence.
The documents that are published are free, specific and dated
The compensation for pile three being useless is that pile one is unusually good in this industry right now. Frontier labs publish more about their own testing than most software vendors publish about anything.
OpenAI’s GPT-5.6 system card, published July 9, 2026 on its Deployment Safety Hub, covers three models, Sol, Terra and Luna, and gives all three the same Preparedness Framework designations: “High in Biological and Chemical, High in Cybersecurity, and below High in AI Self-Improvement” [3]. It carries sections on disallowed content, jailbreaks, prompt injection and hallucinations, and it names the outside groups involved, including SecureBio, the UK AISI, METR, Apollo Research and Irregular [3]. Anthropic publishes a Responsible Scaling Policy, at version 3.4 effective July 8, 2026, which sets out the ASL-3 security and deployment standards, requires public Risk Reports to “contain indications of where material was redacted”, and arranges external review so that “all parts of the unredacted report are evaluated by at least one external reviewer” [4].
None of that is independent in the strict sense; the lab chose the evaluations and wrote the summary. It is still worth 10 minutes, for three reasons. It is specific, so you can find the part that maps to your actual exposure, which for most small businesses is prompt injection and hallucination rather than bioweapons. It is dated, so you know which model version it describes and whether the vendor pitching you is describing the current one. And it is falsifiable, which means the vendor took a small risk by writing it down.
Use that as a screening test on smaller vendors too. If a company selling you an AI product has published nothing of this kind, no data-use page, no model documentation, no statement of what it tested, then every safety claim it makes is in pile three regardless of how confident it sounds.
The data policy is the claim that changes what happens to your files
Safety claims are interesting. Data claims are the ones with consequences for you, and they are published, so they are checkable in about 5 minutes.
OpenAI states that “by default, we do not use data from ChatGPT Enterprise, ChatGPT Business, ChatGPT Edu, ChatGPT for Healthcare, ChatGPT for Teachers, or our API platform—including inputs or outputs—for training or improving our models”, and says qualifying organisations can configure how long business data is retained, “including opting for our zero data retention policy in the API platform” [5]. Anthropic states that “by default, we will not use your inputs or outputs from our commercial products (e.g. Claude for Work, Anthropic API, Claude Gov, etc.) to train our models” [6].
The interesting part of both pages is the exception. Anthropic’s page says that if you “explicitly report feedback or bugs to us (e.g. via our thumbs up/down feedback button)”, it may use your chats and coding sessions to train its models, keeping that feedback in its secured back-end for up to 5 years and de-linking it from your user and customer IDs before use [6]. That is a reasonable trade, but it means the thumbs-down button is a data decision and not a mood. Tell whoever else uses the account.
Two habits make this section stick. Read the page for the exact product you are paying for, because consumer and commercial policies differ inside the same company and a screenshot of one proves nothing about the other. And write the date you read it next to the vendor’s name, because these pages change and a policy you checked in 2024 is not a policy you have checked.
A certification tells you a process exists, not that an answer will be right
Pile two is where most buyers overpay in trust. The badges are real and worth something; they just say less than the logo implies.
Anthropic lists ISO 27001:2022, ISO/IEC 42001:2023, SOC 2 Type I and Type II, and a HIPAA-ready configuration with a business associate agreement available, and directs customers to its Trust Portal to request copies of its compliance documentation [7]. OpenAI states that it aligns with SOC 2 Type 2 Trust Services Criteria and holds ISO/IEC 27001, 27017, 27018 and 27701 certifications [5]. That is a genuinely strong showing by the standards of software vendors, and it is exactly the kind of claim you should ask a smaller vendor to match.
What it means is narrower than it sounds. ISO/IEC 42001:2023 is titled “Information technology — Artificial intelligence — Management system” and was published in December 2023 [8]. A management system standard says the organisation has a documented, audited way of governing how it builds and runs AI. It does not say a model will refuse a bad request, protect you from a prompt injection, or avoid inventing a citation in your report. A SOC 2 report is scoped by the criteria selected and the period covered, which is why the only useful follow-up questions are which criteria, which period, and may I have the report.
There is a small joke buried in the paperwork. To read the standard your vendor is certified against, you buy it: the ISO/IEC 42001:2023 PDF is CHF 225 [8]. Almost nobody outside compliance work has read it, which is a large part of why the badge works so well on a slide.
4 questions that separate an answer from a brochure
When a claim matters enough to ask about, send 4 questions in one short email, and judge the shape of the reply rather than its warmth.
Ask what exactly was evaluated, naming the model version and the product tier you are buying, because a review of the flagship model says nothing about the cheap one you were quoted. Ask when, and take any answer older than a year as describing a system that no longer exists. Ask against what criteria, and whether those criteria are published. Ask who else saw the result, by name, and whether you can have the report or the summary.
A good answer is boring: links, dates, an auditor’s name, a PDF behind an NDA, and a straight “that one is confidential, so we do not ask you to rely on it”. A weak answer is adjectives, a logo wall, or a reference to a review whose criteria nobody outside the room has seen. Neither reply tells you the product is unsafe. The first tells you which pile the claim belongs in, which is the whole job.
vendors × minutes, once a year. Computed in the page; nothing is sent anywhere.
What still goes wrong
Sorting claims does not tell you whether a model is safe. It tells you which claims you are entitled to rely on, and that is a smaller thing. A vendor with a perfect published record and four certifications can still ship a model that leaks your prompt into an output or follows an instruction hidden in a web page you asked it to summarise. The published documents are also written by the party being evaluated, and a lab choosing its own evaluations will not often choose the ones it fails.
The asserted pile is not permanently worthless either. A confidential government review may well be doing real work, and the framework’s contents may become public later, or leak into the open through the sheer number of companies that have now seen them. The honest position today is that you cannot check it, so you cannot count it. That is a statement about your position, not about the review’s quality, and it should update the day the criteria are published.
The other thing that goes wrong is time. Everything above has a date on it: policies change, certifications lapse, system cards describe versions that get replaced within months, and the current framework is voluntary until an executive order says otherwise [1]. A check you ran once is a check you ran once. The recurring version of this work is 15 minutes a year per vendor that touches anything you would not want published, and skipping it is the most common way a careful buyer ends up with a stale answer they still believe.
- 01The White House — Promoting Advanced Artificial Intelligence Innovation and Security (EO 14409)whitehouse.gov
- 02Fortune — White House won't publicly release AI model evaluation frameworkfortune.com
- 03OpenAI — GPT-5.6 System Card, Deployment Safety Hubdeploymentsafety.openai.com
- 04Anthropic — Responsible Scaling Policyanthropic.com
- 05OpenAI — Enterprise privacy and business dataopenai.com
- 06Anthropic — Is my data used for model training? (commercial products)privacy.claude.com
- 07Anthropic — What certifications has Anthropic obtained?privacy.claude.com
- 08ISO/IEC 42001:2023 — Information technology, Artificial intelligence, Management systemiso.org