Frontier Benchmark Matrix
Source-backed July 2026 benchmark results covering the June–July release wave (Claude Fable 5, Claude Sonnet 5, GPT-5.6 Sol, Grok 4.5) alongside provider evaluation sheets and linked public leaderboards.
Claude Fable 5 (June 2026) sets new highs among tracked models on SWE-Bench Pro (80.3%) and OSWorld-Verified (85.0%).
GPT-5.6 Sol edges out Claude Fable 5 on Terminal-Bench 2.1 (88.8% vs 88.0%); Grok 4.5 lands near-frontier coding at $2/$6 per 1M tokens.
GPT-5.5 still holds ARC-AGI-2 (85.0%), GDPval-AA (1773 Elo), and MRCR v2 128k (94.8%) pending fuller GPT-5.6 disclosures.
| Benchmark | Area | Claude Fable 5 | Claude Sonnet 5 | Claude Opus 4.7 | GPT-5.6 Sol | GPT-5.5 | Grok 4.5 | Gemini 3.5 Flash | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|---|---|---|
| Coding | 88.0% | 80.4% | 66.1% | 88.8% | 78.2% | 83.3% | 76.2% | 70.3% | |
SWE-Bench Pro (Public)%Provider eval sheet | Coding | 80.3% | 63.2% | 64.3% | 64.6% | 58.6% | 64.7% | 53.9% | 54.2% |
| Agentic | - | - | 79.1% | - | 75.3% | - | 83.6% | 78.2% | |
Toolathon%Provider eval sheet | Agentic | - | - | - | - | 55.6% | - | 56.5% | - |
OSWorld-Verified%Provider eval sheet | UI Control | 85.0% | 81.2% | 78.0% | - | 78.7% | - | 78.4% | 76.2% |
| Expert Tasks | - | - | 51.5% | - | 51.8% | - | 57.9% | 43.0% | |
| Expert Tasks | 1760 | 1618 | 1753 | - | 1773 | - | 1656 | 1314 | |
CharXiv Reasoning%Provider eval sheet | Multimodal | - | - | 82.1% | - | 84.1% | - | 84.2% | 83.3% |
MMMU-Pro%Provider eval sheet | Multimodal | - | - | 75.2% | - | 81.2% | - | 83.6% | 80.5% |
| Multimodal | 38.6% | - | 24.5% | - | 36.2% | - | 33.6% | 26.5% | |
MRCR v2 (128k avg)%Provider eval sheet | Long Context | - | - | 59.3% | - | 94.8% | - | 77.3% | 84.9% |
MRCR v2 (1M pointwise)%Provider eval sheet | Long Context | - | - | Not supported | - | Not supported | Not supported | 26.6% | 26.3% |
Humanity's Last Exam%Provider eval sheet | Reasoning | 59.0% | - | 46.9% | - | 41.4% | - | 40.2% | 44.4% |
| Reasoning | - | - | 75.8% | - | 85.0% | - | 72.1% | 77.1% |
Methodology Notes
This page intentionally stays separate from the main LLM Leaderboard. The main leaderboard aggregates only the sources it explicitly claims to use. This matrix is a benchmark-by-benchmark reference built from provider evaluation sheets (Google's Gemini 3.5 Flash sheet, plus the June–July 2026 disclosures for Claude Fable 5, Claude Sonnet 5, GPT-5.6, and Grok 4.5) and direct links to the public benchmark pages they cite. A dash means no published score for that model yet; we never mix with-tools and without-tools runs in the same row.