AI Inference

7 pieces on AI Inference.

News & Analysis

news
Shan2026-08-26
OpenAICustom SiliconAI InferenceNvidiaBenchmarks

OpenAI's Jalapeño ASIC Beats Nvidia GB200/GB300 on Latency and Throughput

OpenAI's Jalapeño chip delivers 1.5–1.9× more AI work per watt and 1.7–3.6× lower latency than Nvidia GB200/GB300 across three models.

Read more
news
Shan2026-08-25
Custom SiliconAI InferenceOpenAIBenchmarksHardware

Jalapeño Beats GB200/GB300 by 1.9× Efficiency, 3.6× Latency

OpenAI's custom Jalapeño chip outperforms NVIDIA GB200 and GB300 on InferenceX across three open-weight models up to 1T parameters.

Read more
Kog Bets Deep GPU Engineering Can Deliver 10x LLM Speed
news
Shan2026-08-17
AI InferenceStartupsGPUOpen WeightsEurope

Kog Bets Deep GPU Engineering Can Deliver 10x LLM Speed

French startup Kog hit 3,000 TPS with a 2B-parameter model. Now it needs to prove the same approach works on production LLMs by September.

Read more
Kog Targets 10x LLM Speed With Assembly-Level GPU Tuning
news
Shan2026-08-15
AI InferenceGPU OptimizationOpen WeightsStartupsEurope

Kog Targets 10x LLM Speed With Assembly-Level GPU Tuning

French startup Kog hit 3,000 TPS on Laneformer 2B and is targeting a first major LLM at 10x speed by September 2026.

Read more
news
Shan2026-08-15
AI InferenceGPU OptimizationOpen WeightsEuropean AIStartups

Kog Bets Software Can Unlock 10x Faster LLM Inference on Existing GPUs

French startup Kog hit 3,000 TPS on a 2B-parameter model. Now it must prove the same approach scales to production LLMs by September 2026.

Read more
news
Shan2026-08-14
AI InferenceOpen WeightsStartupsEuropeGPU

Kog Targets 10x LLM Speed by Exploiting GPU Memory Bandwidth

French startup Kog hit 3,000 TPS on its 2B-parameter Laneformer model and aims to prove 10x speed on a major LLM by September.

Read more
Kog Bets Software Can Unlock Stranded GPU Bandwidth for Inference
news
Shan2026-08-14
AI InferenceOpen WeightsStartupsGPUFrance

Kog Bets Software Can Unlock Stranded GPU Bandwidth for Inference

French startup Kog hit 3,000 per-request TPS on AMD MI300X and Nvidia H200 GPUs with a 2B-param model. A full LLM target follows in September.

Read more