AI Inference
7 pieces on AI Inference.
News & Analysis
OpenAI's Jalapeño ASIC Beats Nvidia GB200/GB300 on Latency and Throughput
OpenAI's Jalapeño chip delivers 1.5–1.9× more AI work per watt and 1.7–3.6× lower latency than Nvidia GB200/GB300 across three models.
Jalapeño Beats GB200/GB300 by 1.9× Efficiency, 3.6× Latency
OpenAI's custom Jalapeño chip outperforms NVIDIA GB200 and GB300 on InferenceX across three open-weight models up to 1T parameters.

Kog Bets Deep GPU Engineering Can Deliver 10x LLM Speed
French startup Kog hit 3,000 TPS with a 2B-parameter model. Now it needs to prove the same approach works on production LLMs by September.

Kog Targets 10x LLM Speed With Assembly-Level GPU Tuning
French startup Kog hit 3,000 TPS on Laneformer 2B and is targeting a first major LLM at 10x speed by September 2026.
Kog Bets Software Can Unlock 10x Faster LLM Inference on Existing GPUs
French startup Kog hit 3,000 TPS on a 2B-parameter model. Now it must prove the same approach scales to production LLMs by September 2026.
Kog Targets 10x LLM Speed by Exploiting GPU Memory Bandwidth
French startup Kog hit 3,000 TPS on its 2B-parameter Laneformer model and aims to prove 10x speed on a major LLM by September.

Kog Bets Software Can Unlock Stranded GPU Bandwidth for Inference
French startup Kog hit 3,000 per-request TPS on AMD MI300X and Nvidia H200 GPUs with a 2B-param model. A full LLM target follows in September.