How to read a model launch that ships without weights
Judge a new model on what you can download and licence today, so a lab's open-weight reputation never quietly becomes part of your plan.
on this page · 0 / 0 checked
A lab you have been watching ships its best model yet. Every model it released before came with weights you could download, and this one arrives as an API endpoint, a price list and a wait. Somewhere in the announcement there is a sentence about weights coming later, or no sentence at all. You have to decide this week whether to move work onto it, and the decision quietly turns into a character judgement about the lab.
That is the wrong instrument. Openness is not a trait a lab has; it is a decision taken again for each release, and it comes attached to a licence with conditions in it. This guide is about what to check instead, in about ten minutes, for any model that launches without a downloadable file. It is written for solo operators and small teams choosing where to send real work. If you have a compliance requirement to run models inside your own infrastructure, the questions below are the opening of a legal review rather than a substitute for one.
Open weights are a decision made per release, not a property of a lab
The tidiest evidence is that several labs ship both kinds at once, from the same page. Mistral’s model list puts Apache 2.0 models such as Mistral Small 4, Mistral Large 3 and the Ministral 3 series alongside Premier models including Codestral, Mistral Embed and its OCR line [6]. Same company, same catalogue, two different answers to the question of whether you can hold the file.
OpenAI released gpt-oss-120b and gpt-oss-20b under Apache 2.0 on 5 August 2025, and was explicit that this sits next to the closed lineup rather than replacing it: “Open models complement our hosted models, giving developers a wider range of tools” [5]. Google publishes Gemma weights, but under its own Gemma Terms of Use rather than a standard open licence [4]. In each case the useful unit of analysis is the release, not the vendor.
So a track record tells you what a lab did before. It does not tell you what the release in front of you is, and it will not be the thing you point at when a project stalls because the file never appeared. Check the release.
Checking it takes one search. Go to the lab’s organisation page on Hugging Face and look for the exact model name you intend to use, not the family name. gpt-oss-120b being Apache 2.0 [5] says nothing about the model behind ChatGPT, and a catalogue that lists Apache 2.0 models next to Premier ones [6] says nothing about which side the entry you want falls on. If the exact string is not there, the weights are not there, whatever the family reputation suggests.
A promised release date is not a file you can download
Moonshot’s Kimi K3 is a clean worked example, including the part where it ends well. The launch post described a 2.8 trillion parameter model, called it “the world’s first open 3T-class model”, and stated plainly that “The full model weights will be released by July 27, 2026” [1]. What was live at launch was the API: 1,048,576 tokens of context, $3.00 per million uncached input tokens, $0.30 per million on a cache hit, $15.00 per million output tokens [2]. The weights did arrive, on 27 July [8], and they are on Hugging Face today with a licence attached [3].
The promise was kept. That is worth noticing precisely because it is not a reason to trust the next one. For the days in between, an operator who had already committed to K3 on the strength of that sentence was running on the API, at Moonshot’s price, on Moonshot’s terms, with no second source. Nothing bad happened. The exposure was still real, and it was taken on the basis of a line in a blog post.
The habit that follows is dull and cheap. Write down what exists. A weights file with a licence exists. A date in an announcement is a plan, and plans held by other people move. If your decision only works in the version of the world where the weights ship on time, you are not making a decision about a model, you are making a forecast about a company.
Weights you cannot run still change what you pay
Kimi K3 is 2.8 trillion parameters [1]. You are not running it in your office, and for most readers of this site the fantasy of self-hosting a frontier model was never the point. Open weights still change your position, just not in the way the phrase suggests.
The change is competition. OpenRouter lists 18 providers serving Kimi K3, with input prices from $2.50 to $6.00 per million tokens and output from $12.75 to $22.50 [7]. Moonshot’s own rate is $3.00 uncached input and $15.00 output [2]. Once the file is public, the model stops being a product one company sells and becomes a thing several companies sell, and none of them can take it away from the others.
That is the practical meaning of open weights for a two-person business: a market, not a server. It gives you a route out of a price rise, a rate limit or a retirement that does not require rewriting your prompts for a different model. An API-only model gives you none of that, however good it is. Both can be reasonable choices. They are not the same choice, and the difference shows up on the day something goes wrong rather than the day you sign up.
The distinction to hold onto is between moving host and moving model. Moving host means the same weights answering the same prompts from a different company, which is mostly a change of base URL and key, though speed, rate limits and the exact serving configuration will differ. Moving model means new prompts, new failure modes and a fresh round of checking outputs. Only the first is available to you when the file is public, and only the second is available when it is not. That is the whole of what you are buying when you prefer an open-weight release, and it is worth being unsentimental about it rather than treating open as a virtue.
The licence is the part that actually binds you
“Open weights” is a category with a wide floor. Kimi K3’s licence is permissive at its core, with the standard requirement that “The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software”, and then adds conditions on scale: a model-as-a-service provider whose aggregate revenue passes $20 million over twelve months must enter into a separate agreement with Moonshot AI before using the software, and products above 100 million monthly active users or $20 million in monthly revenue must display “Kimi K3” prominently on the user interface of that product or service [3]. Almost certainly neither applies to you. Confirming it is still worth more than assuming it.
Google’s Gemma terms carry a different kind of clause. Google claims no rights in the outputs you generate [4], which is the reassuring half. The other half is that “Google reserves the right to restrict (remotely or otherwise) usage of any of the Gemma Services that Google reasonably believes are in violation of this Agreement” [4]. Holding a copy of the weights is not the same as being outside the vendor’s reach. Apache 2.0, as used for gpt-oss [5], is the version with the fewest surprises in it.
The ten-minute version of this check: open the LICENSE file in the model’s repository and search it for four words. Revenue. Users. Display. Restrict. If none of them appear, you are probably looking at a standard open licence. If they do, read those paragraphs properly, because they are the ones written for a case somebody thought about.
Price the dependency before you take it
The per-million price is not the number that decides anything. Your monthly bill is, and so is what happens to you if it doubles. Work it out before you move work over, not after the first invoice.
Two details change the answer more than the headline rate. The first is caching: on Kimi K3 the gap between a cache hit and a cache miss is $0.30 against $3.00 per million input tokens, a factor of ten on the input side alone [2], so a workload that resends the same long context behaves nothing like one that does not. The second is where you buy it. The spread across providers serving the same weights runs from $2.50 to $6.00 on input and $12.75 to $22.50 on output [7], which is a wider range than most people assume exists for an identical model.
If the honest answer to “what if this doubles” is that it would be irritating, take the dependency and stop deliberating. If the answer is that it would end the project, you need either a second provider serving the same file [7] or a model whose weights already exist, and you need it before you commit, not during the incident.
A second provider is only a real second provider if you have sent traffic through it. Create the account, run your ten test tasks against it, keep the key somewhere you can find it, and note the price you were quoted. Half an hour, once. The alternative is discovering during an outage that the fallback needs a billing setup, an approval and a different message format, at which point it was never a fallback, only a name on a list.
Input millions × input price, plus output millions × output price. Defaults use Kimi K3's uncached rates [2]. Computed in the page; nothing is sent anywhere.
What still goes wrong
Published weights do not make the vendor’s own service reliable. Demand for Kimi K3’s hosted version maxed out Moonshot’s compute within days of launch, and the company paused new consumer subscriptions [8]. The weights being downloadable is what made that survivable for developers, since other hosts could serve the same model, but it did nothing for anyone who had built on the vendor’s own endpoint and had no second route configured. Availability and openness are separate properties and they fail separately.
A licence covers the copy you have, not the next one. If a lab ships version four under different terms, your existing download keeps working under the terms it came with, and every new capability arrives under the new ones. Gemma’s restriction clause [4] is the reminder that even a downloaded file can sit inside an agreement with the vendor’s discretion written into it. Read the licence again on each major version instead of assuming it carried over.
Finally, none of this tells you whether the model is any good at your work. Launch-week claims about which model beats which travel faster than the evaluations behind them, and a 2.8 trillion parameter count [1] is a fact about a file, not a prediction about your outputs. Run the model on five or ten tasks you have already done, compare against whatever you use now, and let that decide. The checks in this guide only settle what you are exposed to if the answer turns out to be yes.
- 01Moonshot AI — Kimi K3 tech blogkimi.ai
- 02Moonshot AI — Kimi K3 API pricingplatform.kimi.ai
- 03Moonshot AI — Kimi K3 licence (model weights on Hugging Face)huggingface.co
- 04Google — Gemma Terms of Useai.google.dev
- 05OpenAI — Introducing gpt-ossopenai.com
- 06Mistral AI — Models overviewdocs.mistral.ai
- 07OpenRouter — Kimi K3 providers and pricingopenrouter.ai
- 08Fast Company — What to know about Moonshot AI and its new open-weight model Kimi K3fastcompany.com