saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

What to do when the model you use gets accused of theft

A decision procedure for solo operators when an AI lab is accused of stealing a rival's model, so you can keep shipping while the story is still unsettled.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

A lab whose model you route work to gets accused of building it by stealing a competitor’s. The accusation arrives as a headline, a quote from a government official, and a threat of sanctions that has not happened yet. Nothing about your morning changes. Your invoices are the same, your prompts still run, the half-finished automation still half works. What changes is that you now have to decide whether to keep the thing running, using roughly the same information as everyone arguing about it online.

Most people answer this as a moral question and answer it in an afternoon. That is the expensive way to do it, in both directions: rip out a working model over a claim that later evaporates, or leave a dependency in place because the rebuttals sounded convincing and you wanted them to be. This guide is for solo operators, freelancers and small teams who buy model access by the token or download open weights and run them. If you have a procurement function, a security team, or a client contract that names your subprocessors, they own this decision and your job is to hand them facts rather than a verdict. If you train or fine-tune models yourself, your exposure is different in kind and this is not enough.

One headline, three different risks

A provenance scandal bundles three risks that have nothing to do with each other, and the mistake is answering all three with one gut feeling.

The first is compliance risk: is there a rule that binds you, today, because you use this model. The second is continuity risk: could your access disappear, and how long would you be down. The third is quality risk: does the model still do the job. Each is settled by a different document, on a different clock, and only one of them is settled by the news.

Continuity risk is the one worth pricing, because it is the one that has a number. Ask how many hours of your week currently pass through this model, what would break if it stopped answering tomorrow morning, and how long it would take you to have the same work running somewhere else. If the honest answer is a day, the scandal is a curiosity. If it is three weeks, you have a dependency problem that existed before the accusation and will still exist after it resolves, whichever way it goes.

Quality risk is the one the headline says least about. A model does not get worse because its origin is disputed. If anything, the accusation cuts the other way: the claim that a model was built by extracting capability from a frontier system is, stripped of its moral framing, a claim that the model is good. Run your own evaluation and you have your answer, independent of who is right about the training data.

The document that binds you is not the one in the news

In July 2026, US Treasury secretary Scott Bessent said that when Chinese firms “conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table” [6]. Read that sentence carefully. It is a statement about the table. As of the reporting, no sanction and no Entity List designation had been issued [6], and until a designation exists there is no rule that reaches a freelancer in Lisbon paying for API calls.

What does reach you is the licence. Kimi K3’s weights ship under Moonshot’s own licence, which covers the model weights, the configuration files, and the inference and training code, and that licence has thresholds in it [3]. A commercial product built on the software with more than 100 million monthly active users, or more than $20 million in monthly revenue, must display “Kimi K3” prominently on the user interface. A licensee running a model-as-a-service business whose aggregate revenue passes $20 million over any consecutive 12 months has to enter a separate agreement with Moonshot. Internal use, defined as any use that does not make the software, its outputs or its underlying capabilities available to third parties, is exempt from those obligations. The copyright notice has to travel with any copy or substantial portion. The software and anything it outputs are supplied as is, with no warranty of any kind.

For almost every reader of this site, that adds up to: the licence asks nothing of you and promises you nothing. Which is exactly the point of reading it. Five minutes with the licence file tells you more about your actual position than five days of following the dispute, and the thresholds are worth knowing because you cross them by growing, not by deciding to.

Your own terms of service are the near risk

The clause most likely to cost you something this quarter is one you already agreed to. OpenAI’s terms, effective 1 January 2026, list among the things you cannot do: use Output “to develop models that compete with OpenAI.” The same terms separately assign you ownership, saying OpenAI “hereby assign to you all our right, title, and interest, if any, in and to Output” [1]. Anthropic’s consumer terms, effective 8 October 2025, prohibit using the services “to develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models or resell the Services” [2]. Owning the output and being allowed to do anything with it are two different things, and the gap between them is where operators walk into trouble.

The realistic failure is small and domestic. You generate a few thousand answers from ChatGPT or Claude to fine-tune a cheaper local model, or to seed a dataset you plan to resell. That is the same technique the sanctions story is about, at a scale nobody will write about, and it is prohibited by the terms either way.

The ordinary uses are not in danger and it is worth saying so plainly. Selling a deliverable that a model helped you write, running client work through an API, keeping outputs in a knowledge base, building a product whose value is the workflow around the model rather than a competing model: none of that trips the clause. The line these terms draw is around building something that replaces the vendor, and the line is easy to stay behind once you have actually read where it is.

Enforcement is not theoretical. In February 2026 Anthropic said DeepSeek, Moonshot AI and MiniMax had generated more than 16 million exchanges with Claude through approximately 24,000 fake accounts, in violation of its terms of service and regional access restrictions, and that it attributed the campaigns with high confidence using IP address correlation, request metadata and infrastructure indicators [8]. Note what the labs do to each other and what they do to you. To each other, press conferences. To you, account termination, retroactively, on evidence you never see.

An accusation is testable on timeline and mechanism

You cannot audit a training run. You can still test a claim against the calendar and against physics, which is what independent researchers did to the Kimi K3 allegation within a day of it landing.

Braden Hancock, a researcher at the Laude Institute, pointed out that Anthropic’s Fable had only been publicly available since 1 July, shortly before K3 shipped: “You can’t distill that much data, train a model, and release it in two weeks.” He said he does not think “you get a model this strong and this quickly on the heels of Fable doing strictly distillation” [7]. Nathan Lambert, an AI researcher at the Allen Institute for AI, argued the technique itself is losing force: “distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]” [7]. Meanwhile the government claims as reported centred on hardware and access, with the White House’s science and technology policy chief Michael Kratsios alleging Moonshot had acquired Nvidia GB300-equipped servers and accessed GB300s in Thailand, rather than on published technical proof of copying [6].

Generalise the procedure, because you will need it again for a different lab. Ask what has been shown versus asserted. Ask whether the mechanism fits the calendar, since a capability that takes months to extract cannot have been extracted in two weeks. Ask who benefits from you believing it, and notice which companies are conspicuously not signing the letters. None of that tells you the truth. It tells you how much weight the claim can carry, which is the only thing you need in order to size your response.

Price the switch before you are forced to make it

The cheapest moment to cost a migration is the moment you are not migrating. Do it on a Tuesday, with a spreadsheet, and the decision stops being emotional.

Kimi K3 is billed at $3.00 per million input tokens on a cache miss, $0.30 per million on a cache hit, and $15.00 per million output tokens [4]. On the other side, Claude Sonnet 5 is $2 per million input and $10 per million output, Claude Opus 5 is $5 and $25, and Claude Haiku 4.5 is $1 and $5 [5]. At that spread, the token bill is rarely what stops you. What stops you is rework: prompts tuned to one model’s habits, a tool-calling format that differs, an output shape your downstream steps assume.

Time the migration too, not just the bill. Pick one job you run often, point it at the fallback, and note the clock: how long to get the prompt behaving, how many outputs you had to fix by hand, what broke downstream. That number, in hours, is the real switching cost, and it is the one you will be spending under pressure if you never spend it calmly. An afternoon spent this way also tells you something the pricing pages cannot, which is whether the cheaper model is cheaper on your work or only on paper.

One detail worth internalising before you compare numbers. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that “produces approximately 30% more tokens for the same text” than the one used by Claude Sonnet 4.6 and earlier [5]. Price per token is not price per job. The only comparison that means anything is your own workload, run twice, priced twice.

calculator
Monthly bill on the fallback model
$ / month

Defaults price 20M input and 4M output tokens at Claude Sonnet 5's $2 / $10 [5]. Compare against the same volumes at your current rates. Computed in the page; nothing is sent anywhere.

The assets worth owning are the ones no vendor holds

A migration is painful in proportion to how much of your system lives inside one vendor’s account. Four things are worth keeping outside it, and none of them are hard.

Keep an evaluation set: twenty to fifty real inputs from your actual work, each with an answer you already know is good. That file is what turns “does the replacement work” from a week of vibes into an afternoon. Keep prompts as plain text files you own, not as strings typed into a Zapier or n8n step where nobody can find them. Keep logs of inputs and outputs for the work that matters, because they are your evidence when a client asks what produced a deliverable. And keep one seam in your code or workflow where the model name is set in a single place, so switching is an edit rather than an archaeology project.

Do that and a provenance scandal becomes an afternoon of testing instead of a crisis. Skip it and you are not really choosing a model, you are choosing whichever one you would find it least painful to keep.

checklist
Before you keep or drop a model under a cloud
0 of 8 · saved in this browser only

What still goes wrong

You cannot verify provenance, and neither can anyone outside the labs involved. Everything you read about who copied whom is either a company’s assertion about a competitor, a government’s assertion, or a researcher reasoning from public timing and cost. Treat all three as evidence about plausibility and none as settled fact, and be honest that this guide is a way of sizing uncertainty, not resolving it.

Speed is the other problem. A designation can land faster than a migration, and if it does, the constraint that reaches you may not even be the sanction itself. It may be your hosting provider, your payment processor or your own client’s policy deciding independently that they would rather not. Downloaded open weights sit on your disk regardless, but the vendor-hosted endpoint you were actually calling can go dark on someone else’s timetable. A fallback you have never once exercised is a plan, not a fallback.

Finally, the thresholds move under you. The Kimi K3 licence obligations that are irrelevant at your current size become real if the thing you are building works [3], and licences get revised between the day you read one and the day you next look. Put a recurring reminder on it. Nobody is going to email you when the terms change.

sources
  1. 01OpenAI — Terms of Useopenai.com
  2. 02Anthropic — Consumer Terms of Serviceanthropic.com
  3. 03Moonshot AI — Kimi K3 Licensehuggingface.co
  4. 04Moonshot AI — Kimi K3 API pricingplatform.kimi.ai
  5. 05Anthropic — Claude API pricingplatform.claude.com
  6. 06TechCrunch — Treasury threatens sanctions after White House claims Moonshot distilled Anthropic's Fabletechcrunch.com
  7. 07TechCrunch — Experts say exploiting Anthropic's Fable isn't how Kimi K3 got so goodtechcrunch.com
  8. 08VentureBeat — Anthropic says DeepSeek, Moonshot and MiniMax used 24,000 fake accounts to distill Claudeventurebeat.com
next guide
Approval prompts stop working before you notice
9 min · verified 2026-09-05
related guides