Inference Optimization
5 pieces on Inference Optimization.
News & Analysis
LinkedIn's 0.6B Student Model Trains 8x Faster via Multi-Teacher Distillation
LinkedIn stacks five infrastructure optimizations to achieve an 8x training speedup for a 0.6B-parameter job-search ranking model supervised by two teacher LLMs.
Anthropic's Apache 2.0 Commerce Agents Blueprint: Skills Over Subagents
Anthropic's open-source commerce-agents repo ships shopping and merchant agents, four verticals, and a gate layer—all under Apache 2.0.
NVIDIA MPS Cuts ASR GPU Count from 16 to 4 on EC2
NVIDIA CUDA MPS on EC2 g7e.4xlarge delivers 92.1 RPS at 352 ms mean latency, reducing Heidi Health's GPU fleet 75% while holding sub-second SLAs.
LFM2.5-DSpark Hits 3.18x GPU Speedup With Zero Output Change
Liquid AI's DSpark draft checkpoints deliver up to 3.18x throughput on H100 and 2.87x on M4 Max MacBook Pro, with bit-identical greedy output.

Software Extraction Beats Hardware Acquisition at the AI Frontier
GLM-5.3, Kog, and Needle 2 show that post-training, bandwidth recovery, and compression outperform new compute spend in August 2026.