DeepSeek and Xiaomi’s open models use sparse experts to cut AI compute costs
DeepSeek V4.1 Flash and Xiaomi MiMo-V2.6 Pro, released 12 days apart in September, use Mixture-of-Experts architecture: each token activates only a fraction of the model’s parameters. The DEV Community article says V4.1 Flash has 552 billion parameters, with about 16 billion active per output token, while MiMo-V2.6 Pro has 1.02 trillion parameters and 42 billion active; both are described as MIT-licensed.
Bottom line — V4.1 Flash activates about 16 billion of its 552 billion parameters per output token, but sparse computation does not reduce the memory needed to hold the full model.
Go deeper 6
-
The DEV Community article compares the router to a triage nurse: it sends each token to a small number of relevant expert networks.
-
According to DeepInfra, sparse activation lowers per-token computation, but all expert weights must still be available in GPU memory.
-
The DEV Community article says DeepSeek-V3’s 671 billion parameters require about 671 GB of weights in FP8, despite computing like a 37-billion-parameter model.
-
DeepInfra notes that routing tokens between GPUs adds communication overhead, which can offset some compute savings at high serving volumes.
-
The DEV Community article describes router collapse: a few experts can attract most tokens while others receive too little training; load-balancing losses are one proposed fix.
-
Sivaro traces the original mixture-of-experts idea to a 1991 paper by Jacobs, Jordan, Nowlan and Hinton, and the scalable sparse-gating approach to Shazeer and colleagues in 2017.