Mixture-of-Experts
5 pieces on Mixture-of-Experts.
News & Analysis
FreeToken Runs 284B MoE Models on a Single Consumer GPU
FreeToken's q* policy and semantic anchor checkpointing deliver 3–4× faster decode and 6–30× faster prefill for frontier MoE models on RTX consumer hardware.
GLM-5.3-Flash and Qwen3.8-Flash-Next Independently Hit the Same 3:1 Attention Ratio
Two Chinese AI labs independently converged on a 3:1 linear-to-full attention ratio, a 2048-token sparse budget, four gated residual streams, and Muon.
Qwen3.8-Flash-Next: 125B MoE Runs at 6B Active Params, Previews Qwen4
Alibaba's Qwen team releases a 180B-on-disk multimodal MoE that activates only 6B parameters per token, trained at one-ninth the cost of Qwen3.7-Plus.

NVIDIA Nemotron 3.5 Lightning: 30B MoE with 3B Active Parameters
NVIDIA ships Nemotron 3.5 Lightning, a 30B MoE with 3B active parameters and 1M-token context, plus NeMo Switchyard for per-step agent routing.
GLM-5.2 Raises the Bar for Text-Only Open-Weights LLMs
Z.ai's GLM-5.2 arrives as a 753B-parameter open-weights text model with a 1M-token context window, strong benchmark results, and aggressive API pricing.