LLM Leaderboard
Aggregated rankings from the world's most trusted AI evaluation labs.
Loading interactive rankings...
Top 15 models by AI Mastery Index (2026-09-14)
| # | Model | Provider | Weights | AI Mastery Index | Arena Elo | MMLU-Pro | Input $/1M |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | Closed | 96.9 | 1504 | 95.9% | $10.00 |
| 2 | GPT-6 Astra | OpenAI | Closed | 96.7 | 1496 | 95.8% | $10.00 |
| 3 | Claude Opus 5 | Anthropic | Closed | 96.6 | 1493 | 95.7% | $5.00 |
| 4 | Claude Fable 5 | Anthropic | Closed | 96.4 | 1507 | 95.6% | $10.00 |
| 5 | GPT-5.6 Sol | OpenAI | Closed | 95.7 | 1483 | 95.1% | $4.00 |
| 6 | Claude Sonnet 5 | Anthropic | Closed | 94.8 | 1472 | 94.6% | $2.00 |
| 7 | Claude Opus 4.8 | Anthropic | Closed | 94.6 | 1463 | 94.5% | $5.00 |
| 8 | GPT-5.5 Pro | OpenAI | Closed | 94.2 | 1468 | 94.2% | $30.00 |
| 9 | GPT-5.6 Terra | OpenAI | Closed | 93.9 | 1470 | 93.5% | $2.00 |
| 10 | Grok 4.6 | xAI | Closed | 93.8 | 1470 | 93.1% | $2.00 |
| 11 | Claude 4.7 Opus | Anthropic | Closed | 93.8 | 1494 | 93.8% | $5.00 |
| 12 | Gemini 3.8 Flash | Closed | 93.5 | 1494 | 93.2% | $0.75 | |
| 13 | Grok 4.5 | xAI | Closed | 93.4 | 1465 | 92.8% | $2.00 |
| 14 | Gemini 3.1 Ultra | Closed | 92.5 | 1480 | 92.5% | $20.00 | |
| 15 | Muse Spark 1.1 | Meta | Closed | 92.3 | 1492 | 91.9% | $1.25 |
Coding Performance
Measured via HumanEval++ and LiveCodeBench. Reflects ability to handle complex system-level refactoring and library integration.
Agentic Reasoning
Our proprietary AMSE-2026 benchmark. Measures success rates in 10-turn planning loops with self-correction and tool use.
Intelligence ROI
Calculated as (Average Score / Log10(Cost per 1M tokens)). Higher is better value.
Arena Elo and per-token pricing are read from the public Arena board and each vendor's own pricing page. Where a model has no published MMLU-Pro or Hugging Face result yet, the remaining columns carry a peer-relative estimate rather than a measured score. Figures as of 2026-09-14.