LLM Leaderboard

Aggregated rankings from the world's most trusted AI evaluation labs.

Loading interactive rankings...

Top 15 models by AI Mastery Index (2026-09-14)

#ModelProviderWeightsAI Mastery IndexArena EloMMLU-ProInput $/1M
1Claude Fable 5.1AnthropicClosed96.9150495.9%$10.00
2GPT-6 AstraOpenAIClosed96.7149695.8%$10.00
3Claude Opus 5AnthropicClosed96.6149395.7%$5.00
4Claude Fable 5AnthropicClosed96.4150795.6%$10.00
5GPT-5.6 SolOpenAIClosed95.7148395.1%$4.00
6Claude Sonnet 5AnthropicClosed94.8147294.6%$2.00
7Claude Opus 4.8AnthropicClosed94.6146394.5%$5.00
8GPT-5.5 ProOpenAIClosed94.2146894.2%$30.00
9GPT-5.6 TerraOpenAIClosed93.9147093.5%$2.00
10Grok 4.6xAIClosed93.8147093.1%$2.00
11Claude 4.7 OpusAnthropicClosed93.8149493.8%$5.00
12Gemini 3.8 FlashGoogleClosed93.5149493.2%$0.75
13Grok 4.5xAIClosed93.4146592.8%$2.00
14Gemini 3.1 UltraGoogleClosed92.5148092.5%$20.00
15Muse Spark 1.1MetaClosed92.3149291.9%$1.25

Coding Performance

Measured via HumanEval++ and LiveCodeBench. Reflects ability to handle complex system-level refactoring and library integration.

Agentic Reasoning

Our proprietary AMSE-2026 benchmark. Measures success rates in 10-turn planning loops with self-correction and tool use.

Intelligence ROI

Calculated as (Average Score / Log10(Cost per 1M tokens)). Higher is better value.

Arena Elo and per-token pricing are read from the public Arena board and each vendor's own pricing page. Where a model has no published MMLU-Pro or Hugging Face result yet, the remaining columns carry a peer-relative estimate rather than a measured score. Figures as of 2026-09-14.