Inference
11 pieces on Inference.
News & Analysis
Four-Layer Local AI Stack Runs SLMs With No Cloud Dependency
A practical framework for serving, IDE integration, terminal automation, and retrieval — assembling open-weight 1B–14B models into a real development workflow.
Salesforce Gets Multi-AZ HA on SageMaker With SchedulingConfig
Salesforce's 50+ production models needed 2-AZ compliance. SageMaker's new SchedulingConfig parameter closes the gap without surrendering GPU cost efficiency.

Four Stack-Layer Gains Prove Systems Engineering Now Rivals Scaling
Four advances in 48 hours — custom silicon, 4-bit quantization, speculative decoding, and a new transport protocol — show the inference stack delivering discontinuous gains.
OpenAI's Jalapeño Beats Blackwell on Inference Efficiency at Hot Chips
OpenAI's Jalapeño chip outperforms Nvidia Blackwell on tokens-per-user and throughput-per-kilowatt in SemiAnalysis' InferenceX benchmark.
Etched Raises $700M at $21B Valuation After Jane Street Deploys Its Hardware
Etched's valuation doubled from $10.3B to $21B in a month after quant firm Jane Street tested its inference chips and deployed a rack in its own datacenter.
NVIDIA TRTMC: Hugging Face to C++ TensorRT in Two Commands, No ONNX
NVIDIA's TensorRT Model Connect converts supported checkpoints to native C++ inference in two CLI commands, no ONNX export, across 76 model families.
Groq Raises $350M at $3.5B Valuation in Neocloud Pivot
Groq closes $350M at $3.5B — down from $6.9B — as it abandons custom LPU chips and operates Nvidia GPUs across 13 data centers.

NVIDIA Nemotron 3.5 Lightning: 30B Parameters, 3B Active
NVIDIA's Nemotron 3.5 Lightning activates only 3B of 30B parameters per token, targeting the execution layer of multi-model agent stacks.

Needle 2: 45M-Parameter Tool-Calling Model in a 14MB Binary
Cactus Compute's Needle 2 runs a full inference session in 28MB of RAM, hits 500 tokens/sec on a Raspberry Pi 5, and needs no GPU.

GPT-5.6 Sol Ultrafast Mode Delivers 750 Tokens/sec via Cerebras
OpenAI's new Ultrafast API tier runs GPT-5.6 Sol at up to 750 output tokens/sec — 14× Standard speed — powered by Cerebras silicon.

OpenAI Ultrafast Mode Hits 750 Tokens/sec on GPT-5.6 Sol
OpenAI's Ultrafast preview mode delivers up to 750 output tokens per second on GPT-5.6 Sol — 14x standard speed — via a Cerebras chip partnership.