How Mixture of Experts Routing Works in High-Scale Systems
Discover how Mixture of Experts (MoE) routing architectures optimize LLM inference speed and efficiency in production systems with concrete benchmarks.
Tag archive
Discover how Mixture of Experts (MoE) routing architectures optimize LLM inference speed and efficiency in production systems with concrete benchmarks.

Most language models are trained once and frozen forever. mini-AGI is a small byte-level model that keeps learning continuously on a single 8GB consumer GPU, using a mixture-of-experts architecture that pages experts to disk and a simple learning-rate trick to avoid catastrophic forgetting.
Alibaba has not publicly released Qwen4, but its Qwen3.8-Flash-Next repository previews the planned architecture and its T-Head unit has announced the 144 GB Zhenwu M890 accelerator.
Xiaomi released two one-million-context multimodal sparse agent models and a 9B Qwen distill, but its efficient Flash model still requires a roughly 178 GB weight download.
The InternLM team released Intern-S2-397B under Apache 2.0 on 11-13 September 2026, a multimodal mixture-of-experts model aimed at scientific reasoning and long agent tasks; the FP8 weights are a 406.3 GB download and the recommended setup is a node
DeepSeek released V4.1 Flash under an MIT licence on 10 September 2026: a 552-billion-parameter backbone plus a separate 196-billion-parameter memory module, split into an encoder and a decoder so that reading text costs half as much compute as writi
A project called Deltafin runs the full uncompressed Kimi K3 model — 2.8 trillion parameters, with 1.45 TB of expert weights — on a single MacBook Pro by streaming experts from four external SSDs on demand, sustaining almost exactly one token per sec
I bounded an ambiguity in a config file at 2.4% and called it immaterial. The model's own header settled it months later, and correcting the assumption pushed a gate from 1.1% off to 3.4% off, outside the tolerance it had been passing.
A paper posted September 1, 2026 ran the first compute-matched test of looped mixture-of-experts transformers and found that re-running the middle half of the layers a second time saves compute at the frontier, with savings growing as budgets grow.
slotstream, a single Swift binary released as a Show HN on September 1, 2026, runs the 104 GB Qwen3.8-Flash-Next mixture-of-experts model on Macs with a fraction of that memory by keeping a small trunk resident and reading expert weights off the SSD
Z.ai released GLM-5.3-Flash under an MIT licence and confirmed it is the anonymous \u201cOx Alpha\u201d model that topped OpenRouter for a week -- served, the company says, entirely on a cluster of Chinese AI accelerators at per-token cost comparable
Alibaba's Qwen released Qwen3.8-Flash-Next, a preview of the architecture behind Qwen4, whose headline idea is scaling parameters through a 20-million-entry table of word pairs and triples that can be offloaded off the GPU -- 51 billion parameters th