Benchmarks
27 pieces on Benchmarks.
News & Analysis
EnvHarness Wraps Static Benchmarks, Lifts ALFWorld OOD Score 9 Points
Google Cloud AI Research's EnvHarness reshapes frozen benchmarks via a plug-in layer, gaining 9.0 OOD points on ALFWorld and 9.8% fewer SWE-bench steps.
Gemini 3.5 Transcribe: 2.6% WER, 85+ Languages, Two API Surfaces
Google's Gemini 3.5 Transcribe posts 2.6% non-streaming WER across 85+ languages, with a dual-endpoint design that forces hard architectural choices.
Google DeepMind's Double-Blind AI Evals Use Cryptographic Isolation
Google DeepMind pilots the world's first double-blind frontier AI evaluation, using Confidential Space to cryptographically protect both model weights and test prompts.
OpenAI's Jalapeño ASIC Beats Nvidia GB200/GB300 on Latency and Throughput
OpenAI's Jalapeño chip delivers 1.5–1.9× more AI work per watt and 1.7–3.6× lower latency than Nvidia GB200/GB300 across three models.
Jalapeño Beats GB200/GB300 by 1.9× Efficiency, 3.6× Latency
OpenAI's custom Jalapeño chip outperforms NVIDIA GB200 and GB300 on InferenceX across three open-weight models up to 1T parameters.
OpenAI's Jalapeño Beats Blackwell on Inference Efficiency at Hot Chips
OpenAI's Jalapeño chip outperforms Nvidia Blackwell on tokens-per-user and throughput-per-kilowatt in SemiAnalysis' InferenceX benchmark.
Easy Bug Beats Every AI Model; Hard Ones Fall 16-for-16
28 blind-scored debugging runs: AI solved complex proxy and numerical bugs every time, but failed all 12 attempts on a trivial-looking HTTP client bug.
ASR Benchmarks Are Gameable: 6 of 11 Top Models Reproduce Audio Errors
Hume AI tested 11 open-source ASR models and found six reproduce VoxPopuli's transcript errors even when audio contradicts them — exposing WER as a gameable metric.
Nvidia's AVO Harness Takes Claude Opus 5 from 30% to 100% on ARC-AGI-3
Nvidia's custom AVO harness lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3 — without changing the model at all.
ALTK-Evolve: Agent Memory Gains Depend on Model Tier, Not Just Size
IBM Research tests memory injection across 8 models on AppWorld: weaker models gain +16.1pp at +5% token cost via retrieval; strong models need the full set.
Harper 5.2 Beats Vercel Stack Up to 14× on Live Personalized Reads
Harper's benchmark across 474 load tests shows up to 14× latency advantage over a Vercel/Neon/Upstash/Ably stack on live, personalized-data paths.
Kimi K3's 1M-Token Window Costs 16× More Than RAG on 12 Questions
A controlled 12-question blind trial pits Kimi K3's full 127K-token prompt against a tuned RAG pipeline. Long-context wins on completeness, loses on cost and latency.
Sonic-3.6 Hits 1,283 Elo and Leads Both Artificial Analysis Speech Arenas
Cartesia's Sonic-3.6 tops both Artificial Analysis speech leaderboards with 1,283 Elo (Provider Voice) and 1,123 Elo (Controlled Voice), at $49/1M characters.
Flat Recovery Across All Densities: Edge Utilization Is What Moves
A 50-run benchmark shows information recovery stays between 0.924–0.976 across all densities. Edge utilization tells the real story.
Z.ai GLM-5.3: Benchmark Gains From Post-Training Alone
GLM-5.3 reuses the 743B GLM-5.2 base model unchanged. Every benchmark gain comes from scaled post-training environments and longer training runs.

Z.ai GLM-5.3: Big Benchmark Gains from Post-Training Alone
Z.ai's GLM-5.3 reuses the 743B GLM-5.2 base model unchanged, delivering major gains on long-horizon coding and cybersecurity benchmarks through post-training scale alone.

Z.ai GLM-5.3: Frontier Gains From Post-Training Alone
GLM-5.3 reuses GLM-5.2's 743B base model unchanged. Terminal-Bench 3.0 jumps from 4.6 to 28.3; CyberGym hits 84.5%, edging past GPT-5.6 Sol.

Z.ai GLM-5.3: Post-Training Gains on a Fixed 743B Base Model
GLM-5.3 reuses GLM-5.2's 743B base unchanged. Terminal-Bench 3.0 jumps from 4.6 to 28.3; CyberGym hits 84.5%, edging closed frontier models.

Z.ai GLM-5.3: Frozen 743B Base, All Gains from Post-Training
Z.ai's GLM-5.3 reuses the frozen GLM-5.2 743B base model, extracting every benchmark gain through scaled post-training alone.

Gemini 3.7 Flash: Coding and Agent Model at $0.75/1M Input Tokens
Google ships Gemini 3.7 Flash three weeks after 3.6 Flash — FrontierCode 43.6%, DeepSWE 65.3%, at $0.75/1M input until Dec 31 2026.

Grok 4.6: 500K-Context Post-Training Upgrade for Agents
SpaceXAI ships Grok 4.6 with a 500K-token context window, xhigh reasoning effort, and agentic RL — priced at $2/$6 per 1M with a 200K billing cliff.

Liquid AI LFM2.5-VL-3B: 3.1B On-Device Vision-Language Model
Liquid AI's 3.1B-parameter LFM2.5-VL-3B scores 69.4 across 28 vision benchmarks, matches 4.7B rivals, and adds tool calling for on-device agents.

PROVE: Xiaomi's Perception-Aligned Video Removal Metrics RC-S and RC-T
Xiaomi MiLM Plus releases PROVE, two reference-free metrics and a real-world benchmark that outperform PSNR, ReMOVE, and CFD on video object removal.
NOOA: NVIDIA's Object-Oriented Agent Framework Explained
NVIDIA open-sources NOOA, a Python framework that collapses prompt templates, tool schemas, and workflow graphs into one class—with 82.2% on SWE-bench Verified.
Harness-1 Shows Smaller Open Models Can Beat Frontier AI at Search
Harness-1 is a 20B open-source search agent that beats GPT-5.4 on recall by moving search memory out of the model and into a structured environment.
DeepSWE Reshuffles the AI Coding Leaderboard and Puts GPT-5.5 on Top
A tougher coding benchmark shows wider gaps between frontier AI models, with GPT-5.5 leading and verifier quality becoming the real story.
OpenAI Makes GPT-5.5 Instant the New Default for ChatGPT
OpenAI has rolled out GPT-5.5 Instant as ChatGPT's new default model, promising faster responses, fewer hallucinations in high-stakes domains, stronger reasoning scores, and broader memory-source visibility.