Systems Engineering

6 pieces on Systems Engineering.

News & Analysis

Infrastructure Rewrites, Not Model Weights, Drive 2026 AI Gains
articles
Shan2026-09-16
inference optimizationsystems engineeringLLM infrastructureproduction AINVIDIAOpenAI

Infrastructure Rewrites, Not Model Weights, Drive 2026 AI Gains

Four independent engineering efforts show runtime routing, service rewrites, and kernel fusion outperforming model compression on production AI systems.

Read more
news
Shan2026-09-15
NVIDIAcuDNNGPU InferenceInference OptimizationSystems Engineering

cuDNN Graph API: Fusion, Autotuning, and Plan Reuse Explained

NVIDIA's cuDNN Frontend Graph API makes kernel fusion, engine selection, and compilation timing explicit — here's how each technique works.

Read more
news
Shan2026-09-06
EmbeddingsInference InfrastructureGPU ServingPerplexitySystems Engineering

Perplexity's Ivy, Tulip and ROSE Serve pplx-embed Without a Separate Engine

Perplexity reuses its LLM prefill/decode kernels for embedding inference, with 512 tokens saturating a sub-1B model on Hopper and Blackwell hardware.

Read more
news
Shan2026-09-04
LLM InferenceGPU InfrastructureMLOpsSystems Engineering

Chunked Prefill Beats Disaggregation Below 1,000 GPUs

Every major inference framework now ships prefill-decode disaggregation — but for most teams it's the wrong default, and chunked prefill is the safer fix.

Read more
news
Shan2026-08-29
Agentic AIHuman-in-the-LoopLLM ApplicationsSystems EngineeringLLM Evaluation

Risk-Scored Routing Cuts Human Review to High-Signal Queries Only

A text-to-SQL team replaced blanket approval gates with a four-signal risk router, sending only genuinely ambiguous actions to human reviewers.

Read more
Four Stack-Layer Gains Prove Systems Engineering Now Rivals Scaling
articles
Shan2026-08-26
inferencesystems engineeringquantizationAI hardware

Four Stack-Layer Gains Prove Systems Engineering Now Rivals Scaling

Four advances in 48 hours — custom silicon, 4-bit quantization, speculative decoding, and a new transport protocol — show the inference stack delivering discontinuous gains.

Read more