Benchmarks

27 pieces on Benchmarks.

News & Analysis

news
Shan2026-08-30
Reinforcement LearningAgent TrainingGoogle Cloud AIBenchmarksOpen Source

EnvHarness Wraps Static Benchmarks, Lifts ALFWorld OOD Score 9 Points

Google Cloud AI Research's EnvHarness reshapes frozen benchmarks via a plug-in layer, gaining 9.0 OOD points on ALFWorld and 9.8% fewer SWE-bench steps.

Read more
news
Shan2026-08-28
GoogleSpeech-to-TextAPIBenchmarksMachine Learning

Gemini 3.5 Transcribe: 2.6% WER, 85+ Languages, Two API Surfaces

Google's Gemini 3.5 Transcribe posts 2.6% non-streaming WER across 85+ languages, with a dual-endpoint design that forces hard architectural choices.

Read more
news
Shan2026-08-27
Google DeepMindAI SafetyBenchmarksConfidential ComputingFrontier AI

Google DeepMind's Double-Blind AI Evals Use Cryptographic Isolation

Google DeepMind pilots the world's first double-blind frontier AI evaluation, using Confidential Space to cryptographically protect both model weights and test prompts.

Read more
news
Shan2026-08-26
OpenAICustom SiliconAI InferenceNvidiaBenchmarks

OpenAI's Jalapeño ASIC Beats Nvidia GB200/GB300 on Latency and Throughput

OpenAI's Jalapeño chip delivers 1.5–1.9× more AI work per watt and 1.7–3.6× lower latency than Nvidia GB200/GB300 across three models.

Read more
news
Shan2026-08-25
Custom SiliconAI InferenceOpenAIBenchmarksHardware

Jalapeño Beats GB200/GB300 by 1.9× Efficiency, 3.6× Latency

OpenAI's custom Jalapeño chip outperforms NVIDIA GB200 and GB300 on InferenceX across three open-weight models up to 1T parameters.

Read more
news
Shan2026-08-25
OpenAICustom SiliconInferenceBroadcomBenchmarks

OpenAI's Jalapeño Beats Blackwell on Inference Efficiency at Hot Chips

OpenAI's Jalapeño chip outperforms Nvidia Blackwell on tokens-per-user and throughput-per-kilowatt in SemiAnalysis' InferenceX benchmark.

Read more
news
Shan2026-08-23
Agentic AICode GenerationBenchmarksOpen SourceAI Safety

Easy Bug Beats Every AI Model; Hard Ones Fall 16-for-16

28 blind-scored debugging runs: AI solved complex proxy and numerical bugs every time, but failed all 12 attempts on a trivial-looking HTTP client bug.

Read more
news
Shan2026-08-22
Speech RecognitionBenchmarksOpen WeightsEvaluationHugging Face

ASR Benchmarks Are Gameable: 6 of 11 Top Models Reproduce Audio Errors

Hume AI tested 11 open-source ASR models and found six reproduce VoxPopuli's transcript errors even when audio contradicts them — exposing WER as a gameable metric.

Read more
news
Shan2026-08-22
NvidiaAgentic AIBenchmarksOpen WeightsLLM Infrastructure

Nvidia's AVO Harness Takes Claude Opus 5 from 30% to 100% on ARC-AGI-3

Nvidia's custom AVO harness lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3 — without changing the model at all.

Read more
news
Shan2026-08-20
Agentic AIIBM ResearchInference EfficiencyBenchmarksOpen Weights

ALTK-Evolve: Agent Memory Gains Depend on Model Tier, Not Just Size

IBM Research tests memory injection across 8 models on AppWorld: weaker models gain +16.1pp at +5% token cost via retrieval; strong models need the full set.

Read more
news
Shan2026-08-20
DatabaseServerlessBenchmarksDistributed SystemsArchitecture

Harper 5.2 Beats Vercel Stack Up to 14× on Live Personalized Reads

Harper's benchmark across 474 load tests shows up to 14× latency advantage over a Vercel/Neon/Upstash/Ably stack on live, personalized-data paths.

Read more
news
Shan2026-08-19
Large Language ModelsRAGKimi K3BenchmarksInference Cost

Kimi K3's 1M-Token Window Costs 16× More Than RAG on 12 Questions

A controlled 12-question blind trial pits Kimi K3's full 127K-token prompt against a tuned RAG pipeline. Long-context wins on completeness, loses on cost and latency.

Read more
news
Shan2026-08-18
Text-to-SpeechCartesiaVoice AIBenchmarksState-Space Models

Sonic-3.6 Hits 1,283 Elo and Leads Both Artificial Analysis Speech Arenas

Cartesia's Sonic-3.6 tops both Artificial Analysis speech leaderboards with 1,283 Elo (Provider Voice) and 1,123 Elo (Controlled Voice), at $49/1M characters.

Read more
news
Shan2026-08-18
Agent ArchitectureMulti-Agent SystemsGraph AlgorithmsAI AgentsBenchmarks

Flat Recovery Across All Densities: Edge Utilization Is What Moves

A 50-run benchmark shows information recovery stays between 0.924–0.976 across all densities. Edge utilization tells the real story.

Read more
news
Shan2026-08-17
Open WeightsBenchmarksCybersecurityCode GenerationLarge Language Models

Z.ai GLM-5.3: Benchmark Gains From Post-Training Alone

GLM-5.3 reuses the 743B GLM-5.2 base model unchanged. Every benchmark gain comes from scaled post-training environments and longer training runs.

Read more
Z.ai GLM-5.3: Big Benchmark Gains from Post-Training Alone
news
Shan2026-08-16
Open WeightsBenchmarksCybersecurityCodingPost-Training

Z.ai GLM-5.3: Big Benchmark Gains from Post-Training Alone

Z.ai's GLM-5.3 reuses the 743B GLM-5.2 base model unchanged, delivering major gains on long-horizon coding and cybersecurity benchmarks through post-training scale alone.

Read more
Z.ai GLM-5.3: Frontier Gains From Post-Training Alone
news
Shan2026-08-16
Open WeightsCoding ModelsCybersecurityPost-TrainingBenchmarks

Z.ai GLM-5.3: Frontier Gains From Post-Training Alone

GLM-5.3 reuses GLM-5.2's 743B base model unchanged. Terminal-Bench 3.0 jumps from 4.6 to 28.3; CyberGym hits 84.5%, edging past GPT-5.6 Sol.

Read more
Z.ai GLM-5.3: Post-Training Gains on a Fixed 743B Base Model
news
Shan2026-08-16
Open WeightsCoding AgentsCybersecurityBenchmarksPost-Training

Z.ai GLM-5.3: Post-Training Gains on a Fixed 743B Base Model

GLM-5.3 reuses GLM-5.2's 743B base unchanged. Terminal-Bench 3.0 jumps from 4.6 to 28.3; CyberGym hits 84.5%, edging closed frontier models.

Read more
Z.ai GLM-5.3: Frozen 743B Base, All Gains from Post-Training
news
Shan2026-08-15
Open WeightsCoding ModelsCybersecurityBenchmarksPost-Training

Z.ai GLM-5.3: Frozen 743B Base, All Gains from Post-Training

Z.ai's GLM-5.3 reuses the frozen GLM-5.2 743B base model, extracting every benchmark gain through scaled post-training alone.

Read more
Gemini 3.7 Flash: Coding and Agent Model at $0.75/1M Input Tokens
news
Shan2026-08-13
Google DeepMindLarge Language ModelsAI AgentsBenchmarksPricing

Gemini 3.7 Flash: Coding and Agent Model at $0.75/1M Input Tokens

Google ships Gemini 3.7 Flash three weeks after 3.6 Flash — FrontierCode 43.6%, DeepSWE 65.3%, at $0.75/1M input until Dec 31 2026.

Read more
Grok 4.6: 500K-Context Post-Training Upgrade for Agents
news
Shan2026-08-13
SpaceXAILarge Language ModelsAI AgentsBenchmarksDeveloper Tools

Grok 4.6: 500K-Context Post-Training Upgrade for Agents

SpaceXAI ships Grok 4.6 with a 500K-token context window, xhigh reasoning effort, and agentic RL — priced at $2/$6 per 1M with a 200K billing cliff.

Read more
Liquid AI LFM2.5-VL-3B: 3.1B On-Device Vision-Language Model
news
Shan2026-08-13
Vision-Language ModelsOn-Device AISmall Language ModelsLiquid AIBenchmarks

Liquid AI LFM2.5-VL-3B: 3.1B On-Device Vision-Language Model

Liquid AI's 3.1B-parameter LFM2.5-VL-3B scores 69.4 across 28 vision benchmarks, matches 4.7B rivals, and adds tool calling for on-device agents.

Read more
PROVE: Xiaomi's Perception-Aligned Video Removal Metrics RC-S and RC-T
news
Shan2026-08-12
Computer VisionBenchmarksOpen WeightsVideo EditingXiaomi

PROVE: Xiaomi's Perception-Aligned Video Removal Metrics RC-S and RC-T

Xiaomi MiLM Plus releases PROVE, two reference-free metrics and a real-world benchmark that outperform PSNR, ReMOVE, and CFD on video object removal.

Read more
news
Shan2026-08-09
AI AgentsNVIDIAPythonOpen SourceBenchmarksLLM Infrastructure

NOOA: NVIDIA's Object-Oriented Agent Framework Explained

NVIDIA open-sources NOOA, a Python framework that collapses prompt templates, tool schemas, and workflow graphs into one class—with 82.2% on SWE-bench Verified.

Read more
news
Shan2026-06-11
AI AgentsOpen SourceRAGSearchBenchmarks

Harness-1 Shows Smaller Open Models Can Beat Frontier AI at Search

Harness-1 is a 20B open-source search agent that beats GPT-5.4 on recall by moving search memory out of the model and into a structured environment.

Read more
news
Shan2026-05-28
AICoding AgentsBenchmarksGPT-5.5

DeepSWE Reshuffles the AI Coding Leaderboard and Puts GPT-5.5 on Top

A tougher coding benchmark shows wider gaps between frontier AI models, with GPT-5.5 leading and verifier quality becoming the real story.

Read more
News
Shan2026-05-06T09:30:00+08:00
OpenAIGPT-5.5ChatGPTModelsMemoryBenchmarks

OpenAI Makes GPT-5.5 Instant the New Default for ChatGPT

OpenAI has rolled out GPT-5.5 Instant as ChatGPT's new default model, promising faster responses, fewer hallucinations in high-stakes domains, stronger reasoning scores, and broader memory-source visibility.

Read more