Claude Opus 5 vs Gemini 3.1 Ultra
Head-to-head technical comparison of intelligence, capabilities, and API pricing.
Which should you use?
Claude Opus 5 leads on the composite index by 4.1 points (96.6 vs 92.5).
Their Arena Elo differs by only 13 points, inside the confidence intervals the public board reports, so human preference does not separate them.
Cost is the sharper difference: Gemini 3.1 Ultra is 4x the input price of Claude Opus 5 ($20 against $5 per 1M), so at volume the choice is usually decided by budget rather than benchmarks.
Pick Claude Opus 5 when
- Higher composite index — 96.6 against 92.5, a gap of 4.1 points.
- Stronger on MMLU-Pro: 95.7% against 92.5%.
- Cheaper input tokens — $5 against $20 per 1M, so Gemini 3.1 Ultra costs 4x as much to feed.
- Better value per point of measured capability: $0.052 per index point against $0.216.
Pick Gemini 3.1 Ultra when
On these measures Gemini 3.1 Ultra does not lead Claude Opus 5 anywhere, so pick it only for reasons outside this table — an existing integration, a region, or a contract.
Value per point of measured capability, at list input pricing: Claude Opus 5 $0.052 · Gemini 3.1 Ultra $0.216 per index point. Run your own token mix through the token cost calculator — a comparison at list price ignores caching and batch discounts, which move real bills more than this gap does.
Where these numbers come from
Figures as of 2026-09-07. Arena Elo and pricing are read from the public board and each vendor's own pricing page; see the full leaderboard for all 34 models and which columns are measured rather than estimated, or the benchmark matrix for scores benchmark by benchmark.