LLM Inference
5 pieces on LLM Inference, including 1 step-by-step guide.
Guides
News & Analysis
Chunked Prefill Beats Disaggregation Below 1,000 GPUs
Every major inference framework now ships prefill-decode disaggregation — but for most teams it's the wrong default, and chunked prefill is the safer fix.
Read more →
Shopify's Gisting Cuts 6,000-Token Prompts to 1,500, Drops Latency 38%
Shopify compressed Sidekick's system prompt from 6,000 to 1,500 tokens using learned gist embeddings, cutting end-to-end latency from 6.8s to 4.2s in production.
Read more →
AWS Cuts RAG Token Costs 33% With a Two-Call Compression Pattern
A post-retrieval Lambda function routes chunks through Claude Haiku before Claude Sonnet, cutting tokens by 8.6× and costs by 33% across 500K documents.
Read more →
PagedAttention vs RadixAttention: How LLMs Tame the KV Cache
Two architectures attack LLM KV cache inefficiency from opposite angles — memory fragmentation and redundant prefill. Here is how each works.
Read more →
