Inference Optimization

5 pieces on Inference Optimization.

News & Analysis

news
Shan2026-09-12
Knowledge DistillationLLM ServingRanking SystemsSGLangLinkedInInference Optimization

LinkedIn's 0.6B Student Model Trains 8x Faster via Multi-Teacher Distillation

LinkedIn stacks five infrastructure optimizations to achieve an 8x training speedup for a 0.6B-parameter job-search ranking model supervised by two teacher LLMs.

Read more
news
Shan2026-09-03
AnthropicAI AgentsOpen SourceClaudeInference Optimization

Anthropic's Apache 2.0 Commerce Agents Blueprint: Skills Over Subagents

Anthropic's open-source commerce-agents repo ships shopping and merchant agents, four verticals, and a gate layer—all under Apache 2.0.

Read more
news
Shan2026-08-27
NVIDIAAmazon EC2Inference OptimizationASRCUDA MPS

NVIDIA MPS Cuts ASR GPU Count from 16 to 4 on EC2

NVIDIA CUDA MPS on EC2 g7e.4xlarge delivers 92.1 RPS at 352 ms mean latency, reducing Heidi Health's GPU fleet 75% while holding sub-second SLAs.

Read more
news
Shan2026-08-20
Liquid AISpeculative DecodingInference OptimizationSmall Language ModelsOpen Weights

LFM2.5-DSpark Hits 3.18x GPU Speedup With Zero Output Change

Liquid AI's DSpark draft checkpoints deliver up to 3.18x throughput on H100 and 2.87x on M4 Max MacBook Pro, with bit-identical greedy output.

Read more
Software Extraction Beats Hardware Acquisition at the AI Frontier
articles
Shan2026-08-15
inference optimizationpost-trainingon-device AIagentic AIAI infrastructure

Software Extraction Beats Hardware Acquisition at the AI Frontier

GLM-5.3, Kog, and Needle 2 show that post-training, bandwidth recovery, and compression outperform new compute spend in August 2026.

Read more