Systems Engineering
6 pieces on Systems Engineering.
News & Analysis

Infrastructure Rewrites, Not Model Weights, Drive 2026 AI Gains
Four independent engineering efforts show runtime routing, service rewrites, and kernel fusion outperforming model compression on production AI systems.
cuDNN Graph API: Fusion, Autotuning, and Plan Reuse Explained
NVIDIA's cuDNN Frontend Graph API makes kernel fusion, engine selection, and compilation timing explicit — here's how each technique works.
Perplexity's Ivy, Tulip and ROSE Serve pplx-embed Without a Separate Engine
Perplexity reuses its LLM prefill/decode kernels for embedding inference, with 512 tokens saturating a sub-1B model on Hopper and Blackwell hardware.
Chunked Prefill Beats Disaggregation Below 1,000 GPUs
Every major inference framework now ships prefill-decode disaggregation — but for most teams it's the wrong default, and chunked prefill is the safer fix.
Risk-Scored Routing Cuts Human Review to High-Signal Queries Only
A text-to-SQL team replaced blanket approval gates with a four-signal risk router, sending only genuinely ambiguous actions to human reviewers.

Four Stack-Layer Gains Prove Systems Engineering Now Rivals Scaling
Four advances in 48 hours — custom silicon, 4-bit quantization, speculative decoding, and a new transport protocol — show the inference stack delivering discontinuous gains.