saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

Open models in a closed tool's picker

Work out what an open-weight model in your tool's model picker actually changes for your business, and when running one yourself is worth the evening it costs.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You open the model picker in a tool you pay for and there is a name you do not recognise sitting under the ones you do. GitHub’s Copilot documentation groups DeepSeek, Kimi K2.7 Code and Kimi K3 together and calls them open weight models [1]. The phrase carries a suggestion: that this one is somehow yours, that it is free, that picking it means your work stops going to somebody else’s servers. All three of those readings are wrong, and the third one is wrong in a way that can get you in trouble with a client.

This guide is about what an open-weight model actually changes for a business your size, which turns out to be less than the word implies in the moment and more than it implies over a year. It is for a solo operator or a small team paying for tools that have a model picker, not for anyone running their own inference hardware or building a product on model APIs, who have a different set of questions about throughput and evaluation. Every price and policy below was read from the vendor’s own page on 5 September 2026, and any of it can move.

Open weights describe the file, not the plumbing

An open-weight release means the trained model file is published under a licence that lets you download it and run it. OpenAI released gpt-oss-120b and gpt-oss-20b on 5 August 2025 under the Apache 2.0 licence, freely available for download on Hugging Face and served through partners including Azure, AWS, Ollama and llama.cpp [2]. That is a real thing. The file exists, you can hold a copy of it, and nobody can take that copy back.

None of that applies to the session you are in when you pick that model inside a closed tool. The tool is running the model on its own infrastructure. Your prompt travels to the same data centre it always did, under the same account, governed by the same terms of service as the closed models sitting next to it in the list. Choosing the open one changes the model. It does not change the plumbing.

The clearest sign of this is that data terms inside a single tool vary by model for reasons that have nothing to do with licences. GitHub’s documentation states that when Claude Fable 5 or Claude Fable 5.1 is used, Anthropic retains data, including prompts and outputs, by default to operate safety classifiers that detect harmful use, and that other Claude models continue to operate under zero data retention [1]. Those are contracts between two companies. The model’s licence had no say in them, and neither did you.

This matters most in the case people get wrong. If a client contract says their material cannot go to a third-party service, selecting an open-weight model in a hosted picker does not satisfy that clause. Downloading the same weights and running them on a machine you own might. Same model name, two completely different acts.

The licence is the part that genuinely varies

“Open weight” is not one thing, and the difference sits in a short document you can read over a coffee. Take the two Kimi models GitHub files under that single label. Kimi K2.7 Code ships under what Moonshot calls a Modified MIT License, which adds one clause to the standard text: if the software is used for commercial products or services with more than 100 million monthly active users, or more than 20 million US dollars in monthly revenue, you must prominently display “Kimi K2.7 Code” [3].

Kimi K3, sitting beside it in the same group, is not MIT-derived at all. It carries a bespoke Kimi K3 License that repeats the same attribution threshold and adds one MIT has no equivalent for: if you operate a Model as a Service business whose aggregate revenue exceeds 20 million US dollars, you must enter into a separate agreement. That licence also exempts internal use, which it defines as any use that does not make the software available to third parties [4]. gpt-oss, meanwhile, is plain Apache 2.0 with no thresholds in it at all [2].

If you are reading this, none of those thresholds applies to you, and that is the whole reason to look. Five minutes of reading converts a vague worry about legal exposure into a settled fact. The licences that do have teeth tend to have them at a scale you will notice long before you cross it, and the ones that do not still deserve one read, because the clause you are looking for is always short and always near the end.

Three models, one label in one picker, three different documents. When somebody tells you a model is open, that is the beginning of the question rather than the answer to it.

Open models arrive switched off, and they come and go

GitHub lists its open weight group among the models that are disabled by default, regardless of your default availability policy setting [1]. There are editor version floors as well: Kimi K2.7 Code requires Visual Studio Code v1.127 or JetBrains IDEs 1.9.1-251, and Kimi K3 requires Visual Studio Code v1.131 [1]. So the path to trying one runs through a settings toggle and possibly an editor update, and in a team it runs through whoever holds the admin account.

Then look at the tool next door. Cursor’s model documentation lists Composer 2.5, Grok, Claude, Gemini and GPT models, and no open-weight models at all [8]. Two coding tools, both current today, opposite answers.

The default-off setting is easy to read as a warning. It is better read as a handover. The vendor is saying it will run the model for you but is not putting it in front of everyone unasked, which leaves the decision with whoever holds the admin account. In a one-person business that is you, and the toggle costs a minute. In a four-person business it is worth spending that minute in front of the other three, because switching it on quietly is how you end up with two people producing work on a model nobody agreed to and nobody has checked.

A model’s presence in a picker is a commercial decision by the tool vendor, made for reasons you will never see and revisited whenever those reasons change. This is the practical rule that follows: do not write a process, a saved prompt, or a client-facing promise that depends on one specific model being in one specific picker. Name the job the model does. Keep the model itself replaceable, the way you would with any other supplier who has not signed anything.

The durable benefit is a price floor and a way out

Here is what open weights are actually worth to you, and it has almost nothing to do with which name you click today.

Together lists gpt-oss-120B at $0.15 per million input tokens and $0.60 per million output tokens [6]. On OpenAI’s own price list, GPT-6 Astra is $10 per million input and $50 per million output, and the cheaper GPT-5.6 Luna is $0.20 and $1.20 [7]. Those are not equivalent products and nobody should read the gap as a quality verdict. What the gap shows is structural: because the gpt-oss weights are published, several companies can serve the same file and compete on price, availability and terms [2][6]. Where the weights are not published, one company sets the price and one company decides when the model retires.

Before you do the arithmetic, check which of two money situations you are actually in, because they behave differently. Metered billing, where you pay per token or per request, is the one where model choice moves the number, and the calculator below is for that. A flat subscription with a picker in it is not, or not simply: your monthly fee is the same whichever name you click, and what a cheaper model buys you there is headroom against whatever limit the plan meters, not money back. Read how your own plan counts usage before you tell yourself a switch saved anything, because the answer differs by tool and it changes.

That competition works in your favour even if you never use an open model. It is a ceiling on what a closed vendor can charge for comparable work, and it is a place to go if a vendor changes its terms in a way you cannot accept. You are not buying the model. You are buying the existence of a second supplier.

calculator
What routing one job to a cheaper model saves
$ / month

Defaults are GPT-6 Astra output at $50 and gpt-oss-120B output at $0.60. Output tokens only, since that is the side that usually differs most; add the input side separately if your prompts are long. Computed in the page; nothing is sent anywhere.

Running one yourself is a real option and a real job

The small open models are genuinely small enough for the hardware you own. OpenAI says gpt-oss-120b runs efficiently on a single 80 GB GPU, while gpt-oss-20b can run on edge devices with just 16 GB of memory [2]. Ollama lists gpt-oss:20b as a 14GB download with a 128K context window, and gpt-oss:120b at 65GB with the same context window [5]. The 20b model is a laptop-class thing. The 120b model is a server you would have to buy.

The cost of doing this is not the download. It is you. Somebody keeps the model updated, notices when it stops loading after an operating system upgrade, and checks the output quality that a hosted vendor would otherwise be quietly maintaining. Two situations make that worth it. The first is material that genuinely cannot leave your premises, where local execution is the requirement and cost is beside the point. The second is a volume high enough that per-token billing has become a line on your statement you can feel. For everything in between, the hosted version of the same open weights is the better trade, and you keep the option to move later precisely because the file is public.

The middle path is worth naming, because it is the one most small teams land on. You keep paying a host for the open model rather than running it, which costs you nothing in maintenance and still leaves you standing in a market with several suppliers in it. That is not a compromise between the two options. For most of the year it is the correct one, and the local copy is there for the week a client contract makes it necessary.

checklist
Before you route real work to an open model
0 of 8 · saved in this browser only

What still goes wrong

The prices, the pickers and the toggles all move. Every figure here came from a vendor page read on 5 September 2026, and model lists in particular change faster than almost anything else in this business. Treat the specific names as examples of a shape, not as a shopping list, and re-read the picker before you tell a client anything about it.

The most common mistake is the confidentiality one, and it is worth repeating because it is so easy to make in a hurry. “Open weights” is a statement about a file’s licence. It is not a statement about where your prompt goes, who keeps a copy, or for how long. A hosted open model and a hosted closed model are the same risk profile to your client’s lawyer, and the only version of this that changes the answer is the one running on hardware you control, with the maintenance burden that implies.

The last one is quieter. The exit that open weights give you is theoretical until you have used it once. A second supplier you have never tested is not a plan, it is a comforting sentence. If portability is a real reason you like open models, spend one afternoon running your actual work through the same weights at a different host, find out what breaks, and then you have an exit. Otherwise you have the same lock-in as everybody else, with better vocabulary.

sources
  1. 01GitHub Docs — Supported AI models in GitHub Copilotdocs.github.com
  2. 02OpenAI — Introducing gpt-ossopenai.com
  3. 03Hugging Face — moonshotai/Kimi-K2.7-Code LICENSE (Modified MIT)huggingface.co
  4. 04Hugging Face — moonshotai/Kimi-K3 LICENSEhuggingface.co
  5. 05Ollama — gpt-oss model libraryollama.com
  6. 06Together AI — Pricingtogether.ai
  7. 07OpenAI — API pricingdevelopers.openai.com
  8. 08Cursor — Modelscursor.com
next guide
You can buy a faster answer now. Check who is waiting first
9 min · verified 2026-09-05
related guides