Frontier Benchmark Matrix

Source-backed July 2026 benchmark results covering the June–July release wave (Claude Fable 5, Claude Sonnet 5, GPT-5.6 Sol, Grok 4.5) alongside provider evaluation sheets and linked public leaderboards.

Claude Fable 5 (June 2026) sets new highs among tracked models on SWE-Bench Pro (80.3%) and OSWorld-Verified (85.0%).

GPT-5.6 Sol edges out Claude Fable 5 on Terminal-Bench 2.1 (88.8% vs 88.0%); Grok 4.5 lands near-frontier coding at $2/$6 per 1M tokens.

GPT-5.5 still holds ARC-AGI-2 (85.0%), GDPval-AA (1773 Elo), and MRCR v2 128k (94.8%) pending fuller GPT-5.6 disclosures.

BenchmarkAreaClaude Fable 5Claude Sonnet 5Claude Opus 4.7GPT-5.6 SolGPT-5.5Grok 4.5Gemini 3.5 FlashGemini 3.1 Pro
Terminal-Bench 2.1%Source
Coding88.0%80.4%66.1%88.8%78.2%83.3%76.2%70.3%
SWE-Bench Pro (Public)%Provider eval sheet
Coding80.3%63.2%64.3%64.6%58.6%64.7%53.9%54.2%
MCP Atlas%Source
Agentic--79.1%-75.3%-83.6%78.2%
Toolathon%Provider eval sheet
Agentic----55.6%-56.5%-
OSWorld-Verified%Provider eval sheet
UI Control85.0%81.2%78.0%-78.7%-78.4%76.2%
Finance Agent v2%Source
Expert Tasks--51.5%-51.8%-57.9%43.0%
GDPval-AAEloSource
Expert Tasks176016181753-1773-16561314
CharXiv Reasoning%Provider eval sheet
Multimodal--82.1%-84.1%-84.2%83.3%
MMMU-Pro%Provider eval sheet
Multimodal--75.2%-81.2%-83.6%80.5%
Blueprint-Bench 2%Source
Multimodal38.6%-24.5%-36.2%-33.6%26.5%
MRCR v2 (128k avg)%Provider eval sheet
Long Context--59.3%-94.8%-77.3%84.9%
MRCR v2 (1M pointwise)%Provider eval sheet
Long Context--Not supported-Not supportedNot supported26.6%26.3%
Humanity's Last Exam%Provider eval sheet
Reasoning59.0%-46.9%-41.4%-40.2%44.4%
ARC-AGI-2%Source
Reasoning--75.8%-85.0%-72.1%77.1%

Methodology Notes

This page intentionally stays separate from the main LLM Leaderboard. The main leaderboard aggregates only the sources it explicitly claims to use. This matrix is a benchmark-by-benchmark reference built from provider evaluation sheets (Google's Gemini 3.5 Flash sheet, plus the June–July 2026 disclosures for Claude Fable 5, Claude Sonnet 5, GPT-5.6, and Grok 4.5) and direct links to the public benchmark pages they cite. A dash means no published score for that model yet; we never mix with-tools and without-tools runs in the same row.