Cloudflare doubles context for Kimi K2.6
Archive item — written before sources were shown.
Cloudflare cut serving costs for Kimi K2.6 and GLM 5.2 with KV cache quantization and INT4 weight compression, roughly doubling context and cutting size 40%.
Cloudflare detailed three inference optimizations for serving Moonshot’s Kimi and Z.ai’s GLM model families on its Workers AI platform, published August 3. FP8 KV-cache quantization roughly doubled usable context on Kimi K2.6, from about 686,000 to 1.37 million tokens, while lifting peak throughput 41% at 64 concurrent requests, with negligible accuracy loss. INT4 weight compression cut GLM 5.2’s checkpoint size 40% (705GB to 421GB) with decode-throughput gains from 16% to 55% depending on concurrency, and a new KV-cache integrity check guards against corruption across concurrent requests for under 1% overhead. The stack runs on the open-source SGLang framework across disaggregated prefill/decode pools on H200 GPUs.
- 01Smaller, faster, safer: running Kimi and GLM at scaleblog.cloudflare.com · primary, technical
