friday, september 18, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · running the business

Routing to a Chinese model without betting the business on it

Work out what a US designation would actually change for your setup, what switching would cost you, and what you need to have written down first.

Published 2026-09-04 · Updated 2026-09-04 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

You put the cheap model on the boring half of the work. Classification, first drafts, pulling fields out of invoices, the summaries nobody reads twice. GLM 5.3 Flash runs at $0.075 per million input tokens and $0.25 per million output; DeepSeek V4 Flash is $0.05 and $0.16 [6]. Against Claude Sonnet 5 at $2 and $10, the arithmetic made the decision for you [6]. Then on 21 July 2026 Treasury Secretary Scott Bessent told Fox Business that the administration supports open-source models but not intellectual property theft, and that “if we see, especially, that overseas models are stealing from our great companies, we have the ability to sanction them because of this theft” [8]. You had no way to tell whether that was a headline or a thing that would break your Tuesday.

This guide is for the operator in that position: a few thousand dollars a month of inference at most, no legal department, and clients who occasionally ask which models touch their data. It is not for you if you already hold a federal contract with an AI clause in it, because that is a question for whoever signed it. It is not for a team that already has a compliance function with a written policy on this. What follows is the dull version: what the existing rules actually restrict, what a designation changed the one time it landed on an AI lab, what moving would cost you, and what you should have written down before somebody asks.

The rules that exist restrict exports, not your API bill

The Entity List is what most people mean when they say sanctions in this context. It is run by the Commerce Department’s Bureau of Industry and Security, and its effect is narrower than the word suggests. The Export Administration Regulations impose additional licence requirements, and limit the availability of most licence exceptions, for exports, reexports and transfers (in-country) when a listed entity is a party to the transaction [3]. The direction of travel matters. The list constrains what American companies can send to the listed entity. It does not, by its own terms, say anything about you buying tokens from a reseller that happens to serve that entity’s published model.

The second list worth knowing is the foreign adversary determination at 15 CFR 791.4, which names the People’s Republic of China including Hong Kong and Macau, Cuba, Iran, North Korea, Russia, and the Maduro regime [2]. That is not a sanctions list on its own. It is a definition other rules point at, which is why it turns up in procurement language. One line in it deserves your attention: revisions to the list are effective immediately upon publication in the Federal Register, without prior notice or opportunity for public comment [2]. That is the honest basis for the unease. The definitional layer can move without a warning shot.

So the shape of the risk is not what the July headline implied. Nothing in the export-control text quoted above reaches a private company buying inference. The pressure arrives somewhere less dramatic, which is the next section but one.

Zhipu is the only worked example you have

On 16 January 2025 BIS added 11 entities under the destination of China to the Entity List [3]. Nine of the eleven carry the Zhipu name, including Beijing Zhipu Huazhang Technology Co., Ltd., which the entry also lists under the alias Zhipu AI, and ten of the eleven were added on the stated ground that they advance the People’s Republic of China’s military modernisation through the development and integration of advanced artificial intelligence research [3]. Zhipu AI publishes its weights on Hugging Face under the account zai-org, which describes itself as “Zhipu AI (Z.ai)” and says it builds the ChatGLM family of models [5]. If a designation were going to make an open-weight Chinese model disappear from your stack, this is the case where it would have happened.

It did not. GLM 5.3 Flash sits on that same Hugging Face account today alongside the rest of the GLM 5 releases [5], and it is served on OpenRouter at the prices above, next to GLM 5.3 at $1.15 and $3.50 per million [6]. The designation did real damage to Zhipu’s ability to buy American technology, because a licence is now required for all items subject to the Export Administration Regulations, reviewed under a policy of presumption of denial [3]. It did nothing to the availability of the weights or to the resellers who host them, and that gap has now held for more than a year and a half.

Take the lesson narrowly rather than as reassurance. One designation, one lab, one outcome is not a pattern, and a future action aimed specifically at model distribution rather than at hardware procurement would look different. But it does tell you that “they got sanctioned” and “your endpoint stopped working” are separate events, and that planning for the first as though it automatically causes the second will make you move faster than the evidence warrants.

The clause that reaches you arrives in a contract

Here is where the pressure actually lands on a small business. On 17 June 2026 the General Services Administration published a notice, a set of listening sessions and a draft clause for comment: GSAR 552.239-7001, “Basic Safeguarding of Data Within Large Language Model Artificial Intelligence Systems” [1]. It is a draft, not a final rule. Comments closed on 3 August 2026, and GSA says it is publishing the draft to gather feedback from stakeholders before taking future action, giving deviation and formal rulemaking as the examples of what that action might be [1]. Read it anyway, because it is the clearest statement anyone has published of what a large buyer is about to start asking, and because clauses like this flow down.

Four things in it should shape what you do this month. The contractor must disclose all LLMs used or made available in performance of the contract within 120 days after commencing work, if no earlier date is specified [1]. The contractor must maximise the use of LLMs that are developed, managed and operated by an entity incorporated in the United States and subject to US law and jurisdiction, and each LLM, along with any components performing core model, data storage or processing, output generation, or security functions, is prohibited from being developed, managed or operated by entities subject to the direction, influence or control of adversary foreign governments, with a pointer to 15 CFR 791.4 [1]. The contractor must disclose whether the LLM has been modified or configured to comply with any non-US federal government statutes, regulations or policies, no later than 30 days after award [1]. And if the LLM uses intermediary processing such as reasoning, retrieval or agentic processes, it must summarise the intermediary steps from data input to data output, including at minimum its model routing decisions with accompanying rationale [1].

There is a carve-out worth reading carefully before you panic. Incidental foreign-developed components, and the draft names open-source components and published research as its examples, along with ancillary third-party services and globally operated infrastructure dependencies, are permissible provided they do not introduce security risks or foreign control that would violate the ownership and control criteria, the systems storing or processing government data satisfy applicable federal security requirements, and a risk-based approach is applied focusing on objective criteria such as ownership, control, hosting and security posture [1]. That is a narrow door, not an exemption, and it is a door a subcontractor has to argue their way through in writing.

Your cheap model was aligned under someone else’s content law

The disclosure about modification to comply with a non-US government’s policies is the one most operators cannot answer, and it is answerable. China’s Interim Measures for the Management of Generative AI Services, in force since 15 August 2023, require providers to uphold socialist core values and to avoid generating content that incites subversion of state power, overthrow of the socialist system, or damage to national unity and social stability, among other prohibited categories [4]. Providers of services with public opinion attributes or social mobilisation capacity must complete a security assessment and file their algorithm under the existing registration rules [4].

Two qualifications keep this accurate. The measures apply to services that provide generated content to the public within the People’s Republic of China, so a copy of open weights running on a server in Virginia is not itself a regulated service under them [4]. And every frontier lab shapes model behaviour to some rule set; the American ones just do it under a different one. The point is not that the model is compromised. The point is that when a client or a prime contractor asks whether your system has been configured to comply with a non-US government’s policy, “I have never looked into it” is a worse answer than a short, sourced paragraph you wrote once and reuse.

Price the switch before you need it

Most of the fear here is really a fear of an unpriced bill. Price it and the fear usually shrinks. GPT-5 Nano, a US-hosted option on OpenRouter’s list, runs at $0.05 per million input tokens and $0.40 output, which matches DeepSeek V4 Flash on input and costs two and a half times as much on output [6]. GPT-5.6 Luna is $0.20 and $1.20 [6]. Claude Haiku 4.5 is $1 and $5 [6]. Only at the top of the range does the gap get wide, and at the top of the range the Chinese options are not cheap either: Kimi K3 is $2.50 and $14, and Qwen3.8 Max is $2 and $6 [6].

Read those headline figures for what they are. OpenRouter shows the average price customers actually pay next to the rates providers post, and an open-weight model is served by a crowd of providers at different rates: 24 of them serve GLM 5.3 Flash, posting from $0.10 and $0.35 up to $0.15 and $0.50, and 18 serve Kimi K3 with output rates running from $12.75 to $22.50 per million [6]. A first-party model like GPT-5 Nano, served by two endpoints, does not move that way [6]. If your cheap lane is an open-weight model, the number you budget against depends on which providers you let it route to.

For the classification-and-drafting workloads that most cheap routing actually carries, the monthly difference between a Chinese flash model and a US one is usually a two-figure number, not a business decision. Run your own numbers in the calculator below before you conclude otherwise. If the answer is small, you have just bought yourself the right to stop thinking about this, because you can move whenever you want to and absorb it. If the answer is large, that is worth knowing too, because it means you have a genuine dependency and should treat the rest of this guide as work rather than reading.

Make the swap a configuration change, not a project

Three different exposures hide behind the phrase “we use a Chinese model”, and they fail in different ways. If you call the model through a US aggregator, your dependency is that aggregator’s willingness and ability to keep serving it, which is a vendor problem. If you call a Chinese lab’s own API directly, you have added a cross-border data path and a payment relationship on top. If you have downloaded open weights and run them yourself, the model keeps working no matter what happens upstream, and what you lose is future versions and support. Write down which of the three you have for each model. It takes 10 minutes and it is the single most useful artefact in this whole exercise.

Then make the swap cheap. On OpenRouter, the provider object on a request accepts order, only, ignore, allow_fallbacks, data_collection and zdr fields, which between them let you name the providers a request may use, the order to try them in, whether to fall back when the primary is unavailable, and whether to restrict routing to zero-data-retention endpoints [7]. That is enough machinery to express “prefer the cheap model, fall back to this US one, never route through providers that may store data” as configuration rather than code. Whatever tool you use, the requirement is the same: the model name lives in one place, you can change it without a deploy, and you have an eval set of 20 to 50 real examples you can run against the replacement so the switch is a measurement rather than a leap. If you rely on open weights, keep your own copy of the weights and of the licence text they shipped under. A copy on your own disk does not stop working because the news changed.

checklist
Before the next headline
0 of 8 · saved in this browser only
calculator
What switching would cost you per month
$ / month

Defaults are GLM 5.3 Flash against GPT-5.6 Luna at OpenRouter's listed prices. Computed in the page; nothing is sent anywhere.

What still goes wrong

The calculator prices tokens, and tokens are the cheap part. What it cannot price is the prompt work. A prompt tuned against one model’s habits often loses a few points of accuracy on another, and getting those points back is a day of somebody’s attention, not a line in a config file. That is the real switching cost for most small teams, and it is why the eval set matters more than the fallback declaration. If you have never run your prompts against a second model, you do not actually know what your switching cost is.

The second limit is that this guide describes a legal position, not a client’s mood. A prime contractor can decline to renew you over a model in your stack without any rule requiring them to, and they do not have to explain. The GSA draft is the visible edge of that expectation, and expectations of this kind tend to spread through supply chains faster than the rules they anticipate [1]. If you sell into government, defence, health or finance, treat the disclosure questions in that draft as things you will be asked in the next 12 months regardless of whether the clause is ever finalised.

The third is timing. The foreign adversary determination can be revised effective immediately on publication, without notice or comment [2], and a future action aimed at model distribution rather than at an entity’s purchasing would not resemble the Zhipu case at all. Nothing in this guide should be read as a prediction that the current position holds. It is an argument for building the switch, pricing it, and then getting back to work, rather than either ignoring the question or rebuilding your stack on the strength of a television interview.

sources
  1. 01GSA — General Services Acquisition Regulation; Acquisition of Information and Communication Technology; Notice of Listening Sessions and Request for Comments (91 FR 36559)federalregister.gov
  2. 0215 CFR 791.4 — Determination of foreign adversariesecfr.gov
  3. 03BIS — Addition of Entities to and Revision of Entry on the Entity List, effective 16 January 2025federalregister.gov
  4. 04Cyberspace Administration of China — Interim Measures for the Management of Generative AI Servicescac.gov.cn
  5. 05Hugging Face — Zhipu AI (Z.ai) organisation, zai-orghuggingface.co
  6. 06OpenRouter — Models and pricingopenrouter.ai
  7. 07OpenRouter — Provider routing and provider selectionopenrouter.ai
  8. 08TechCrunch — US threatens sanctions against Chinese AI models over IP thefttechcrunch.com
next guide
Plan for losing your model vendor
9 min · verified 2026-09-04
related guides