Inference Efficiency

5 pieces on Inference Efficiency.

News & Analysis

news
Shan2026-09-10
DeepSeekMixture of ExpertsLong ContextOpen WeightsInference Efficiency

DeepSeek-V4.1-Flash: 890 Bytes Per Token, 437x Smaller KV Cache

DeepSeek's 552B MoE model cuts global KV cache to 890 bytes per token — 437x below V1 — using CED, CSA2, and FP4 quantization.

Read more
news
Shan2026-09-09
ArchitectureInference EfficiencyAmazon SageMakerOpen WeightsBenchmarks

150M-Parameter BDH-CQ Scores 29.2% on ARC-AGI-1 at $0.0007 per Task

Pathway's 150M-parameter BDH-CQ model scores 29.2% pass@2 on ARC-AGI-1 at $0.0007 per task, trained on SageMaker HyperPod with H200 GPUs.

Read more
news
Shan2026-08-28
Open WeightsArchitectureMixture of ExpertsInference EfficiencyChinese AI Labs

GLM-5.3-Flash and Qwen3.8-Flash-Next Independently Hit the Same 3:1 Attention Ratio

Two Chinese AI labs independently converged on a 3:1 linear-to-full attention ratio, a 2048-token sparse budget, four gated residual streams, and Muon.

Read more
news
Shan2026-08-26
Open WeightsMixture of ExpertsAlibabaMultimodalInference Efficiency

Qwen3.8-Flash-Next: 125B MoE Runs at 6B Active Params, Previews Qwen4

Alibaba's Qwen team releases a 180B-on-disk multimodal MoE that activates only 6B parameters per token, trained at one-ninth the cost of Qwen3.7-Plus.

Read more
news
Shan2026-08-20
Agentic AIIBM ResearchInference EfficiencyBenchmarksOpen Weights

ALTK-Evolve: Agent Memory Gains Depend on Model Tier, Not Just Size

IBM Research tests memory injection across 8 models on AppWorld: weaker models gain +16.1pp at +5% token cost via retrieval; strong models need the full set.

Read more