Inference Efficiency
5 pieces on Inference Efficiency.
News & Analysis
DeepSeek-V4.1-Flash: 890 Bytes Per Token, 437x Smaller KV Cache
DeepSeek's 552B MoE model cuts global KV cache to 890 bytes per token — 437x below V1 — using CED, CSA2, and FP4 quantization.
150M-Parameter BDH-CQ Scores 29.2% on ARC-AGI-1 at $0.0007 per Task
Pathway's 150M-parameter BDH-CQ model scores 29.2% pass@2 on ARC-AGI-1 at $0.0007 per task, trained on SageMaker HyperPod with H200 GPUs.
GLM-5.3-Flash and Qwen3.8-Flash-Next Independently Hit the Same 3:1 Attention Ratio
Two Chinese AI labs independently converged on a 3:1 linear-to-full attention ratio, a 2048-token sparse budget, four gated residual streams, and Muon.
Qwen3.8-Flash-Next: 125B MoE Runs at 6B Active Params, Previews Qwen4
Alibaba's Qwen team releases a 180B-on-disk multimodal MoE that activates only 6B parameters per token, trained at one-ninth the cost of Qwen3.7-Plus.
ALTK-Evolve: Agent Memory Gains Depend on Model Tier, Not Just Size
IBM Research tests memory injection across 8 models on AppWorld: weaker models gain +16.1pp at +5% token cost via retrieval; strong models need the full set.