Inference

11 pieces on Inference.

News & Analysis

news
Shan2026-08-29
Small Language ModelsLocal AIOpen WeightsDeveloper ToolsInference

Four-Layer Local AI Stack Runs SLMs With No Cloud Dependency

A practical framework for serving, IDE integration, terminal automation, and retrieval — assembling open-weight 1B–14B models into a real development workflow.

Read more
news
Shan2026-08-28
Amazon SageMakerHigh AvailabilityInferenceSalesforceMLOps

Salesforce Gets Multi-AZ HA on SageMaker With SchedulingConfig

Salesforce's 50+ production models needed 2-AZ compliance. SageMaker's new SchedulingConfig parameter closes the gap without surrendering GPU cost efficiency.

Read more
Four Stack-Layer Gains Prove Systems Engineering Now Rivals Scaling
articles
Shan2026-08-26
inferencesystems engineeringquantizationAI hardware

Four Stack-Layer Gains Prove Systems Engineering Now Rivals Scaling

Four advances in 48 hours — custom silicon, 4-bit quantization, speculative decoding, and a new transport protocol — show the inference stack delivering discontinuous gains.

Read more
news
Shan2026-08-25
OpenAICustom SiliconInferenceBroadcomBenchmarks

OpenAI's Jalapeño Beats Blackwell on Inference Efficiency at Hot Chips

OpenAI's Jalapeño chip outperforms Nvidia Blackwell on tokens-per-user and throughput-per-kilowatt in SemiAnalysis' InferenceX benchmark.

Read more
news
Shan2026-08-19
AI HardwareFundraisingInferenceStartupsNVIDIA

Etched Raises $700M at $21B Valuation After Jane Street Deploys Its Hardware

Etched's valuation doubled from $10.3B to $21B in a month after quant firm Jane Street tested its inference chips and deployed a rack in its own datacenter.

Read more
news
Shan2026-08-18
NVIDIATensorRTInferenceOpen SourceEdge AI

NVIDIA TRTMC: Hugging Face to C++ TensorRT in Two Commands, No ONNX

NVIDIA's TensorRT Model Connect converts supported checkpoints to native C++ inference in two CLI commands, no ONNX export, across 76 model families.

Read more
news
Shan2026-08-17
GroqNeocloudNvidiaInferenceFundraisingData Centers

Groq Raises $350M at $3.5B Valuation in Neocloud Pivot

Groq closes $350M at $3.5B — down from $6.9B — as it abandons custom LPU chips and operates Nvidia GPUs across 13 data centers.

Read more
NVIDIA Nemotron 3.5 Lightning: 30B Parameters, 3B Active
news
Shan2026-08-16
NVIDIAAI AgentsOpen WeightsLarge Language ModelsInference

NVIDIA Nemotron 3.5 Lightning: 30B Parameters, 3B Active

NVIDIA's Nemotron 3.5 Lightning activates only 3B of 30B parameters per token, targeting the execution layer of multi-model agent stacks.

Read more
Needle 2: 45M-Parameter Tool-Calling Model in a 14MB Binary
news
Shan2026-08-14
Small Language ModelsEdge AIOpen WeightsTool CallingInference

Needle 2: 45M-Parameter Tool-Calling Model in a 14MB Binary

Cactus Compute's Needle 2 runs a full inference session in 28MB of RAM, hits 500 tokens/sec on a Raspberry Pi 5, and needs no GPU.

Read more
GPT-5.6 Sol Ultrafast Mode Delivers 750 Tokens/sec via Cerebras
news
Shan2026-08-13
OpenAIInferenceCerebrasAPIGPT-5.6

GPT-5.6 Sol Ultrafast Mode Delivers 750 Tokens/sec via Cerebras

OpenAI's new Ultrafast API tier runs GPT-5.6 Sol at up to 750 output tokens/sec — 14× Standard speed — powered by Cerebras silicon.

Read more
OpenAI Ultrafast Mode Hits 750 Tokens/sec on GPT-5.6 Sol
news
Shan2026-08-13
OpenAIGPT-5.6 SolInferenceCerebrasAgentic AI

OpenAI Ultrafast Mode Hits 750 Tokens/sec on GPT-5.6 Sol

OpenAI's Ultrafast preview mode delivers up to 750 output tokens per second on GPT-5.6 Sol — 14x standard speed — via a Cerebras chip partnership.

Read more