Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model6059545151504444443824
Output tokens per second · Higher is better
gpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning model31919719615611911911169645947
Weighted average cost (USD) per Intelligence Index task · Lower is better
DeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$0.04$0.06$0.12$0.24$0.26$0.31$0.35$0.37$0.59$1.04$2.75
New article published · 10 Jul
Muse Spark 1.1: Meta gains 8 Intelligence Index points in three months
New language model evaluation · 10 Jul
GLM-5.2 (Non-reasoning)GLM-5.2 (Non-reasoning)
New language model evaluation · 10 Jul
Muse Spark 1.1 (xhigh)Muse Spark 1.1 (xhigh)
New article published · 9 Jul
GPT-5.6 benchmarks across Intelligence, Speed and Cost
New language model evaluation · 9 Jul
JT-4.1 Flash 236B A21BJT-4.1 Flash 236B A21B
New language model evaluation · 9 Jul
GPT-5.6 Sol (xhigh)GPT-5.6 Sol (xhigh)
New language model evaluation · 9 Jul
GPT-5.6 Sol (max)GPT-5.6 Sol (max)
New language model evaluation · 9 Jul
GPT-5.6 Sol (high)GPT-5.6 Sol (high)
New language model evaluation · 9 Jul
GPT-5.6 Sol (medium)GPT-5.6 Sol (medium)
New language model evaluation · 9 Jul
GPT-5.6 Sol (low)GPT-5.6 Sol (low)
New language model evaluation · 9 Jul
GPT-5.6 Terra (max)GPT-5.6 Terra (max)
New language model evaluation · 9 Jul
GPT-5.6 Terra (xhigh)GPT-5.6 Terra (xhigh)
New language model evaluation · 9 Jul
GPT-5.6 Terra (high)GPT-5.6 Terra (high)
New language model evaluation · 9 Jul
GPT-5.6 Terra (medium)GPT-5.6 Terra (medium)
New language model evaluation · 9 Jul
GPT-5.6 Terra (low)GPT-5.6 Terra (low)
New language model evaluation · 9 Jul
GPT-5.6 Luna (max)GPT-5.6 Luna (max)
New language model evaluation · 9 Jul
GPT-5.6 Luna (xhigh)GPT-5.6 Luna (xhigh)
New language model evaluation · 9 Jul
GPT-5.6 Luna (high)GPT-5.6 Luna (high)
New language model evaluation · 9 Jul
GPT-5.6 Luna (medium)GPT-5.6 Luna (medium)
New language model evaluation · 9 Jul
GPT-5.6 Luna (low)GPT-5.6 Luna (low)See more

IntelligenceUpdated

Intelligence of leading AI models based on our independent evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model605956555554535151515046464444444240383834303029241714
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model605956555554535151515046464444444240383834303029241714
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better
DeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$0.24$0.17$0.38$0.58$0.42$0.63$0.20$0.22$0.17$0.87$0.87$0.24$0.31$0.48$0.21$0.37$0.30$0.56$0.81$1.18$0.02$0.03$0.04$0.06$0.12$0.14$0.21$0.24$0.24$0.26$0.29$0.31$0.33$0.35$0.37$0.55$0.59$0.86$1.04$1.06$1.08$1.53$1.80$2.75$0.21$0.37
Reasoning models are indicated by a lightbulb icon

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Nov 22May 23Nov 23May 24Oct 24Apr 25Oct 25Mar 26Jul 26Release Date015304560Artificial Analysis Intelligence Index
Reasoning models are indicated by a lightbulb icon.

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Performance, cost, and execution time for leading coding agents on end-to-end software engineering tasks

Explore Artificial Analysis Coding Agent Index

Artificial Analysis Coding Agent Index

Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better
Codex - GPT-5.6 Sol (max)Label for Codex - GPT-5.6 Sol (max)CodexGPT-5.6 Sol(max)Codex - GPT-5.6 Terra (max)Label for Codex - GPT-5.6 Terra (max)CodexGPT-5.6 Terra(max)Claude Code - Fable 5 (max) (with fallback)Label for Claude Code - Fable 5 (max) (with fallback)Claude CodeFable 5(max)(with fallback)Codex - GPT-5.5 (xhigh)Label for Codex - GPT-5.5 (xhigh)CodexGPT-5.5(xhigh)Grok Build - Grok 4.5 (high)Label for Grok Build - Grok 4.5 (high)Grok BuildGrok 4.5(high)Codex - GPT-5.6 Luna (max)Label for Codex - GPT-5.6 Luna (max)CodexGPT-5.6 Luna(max)Claude Code - Opus 4.8 (max)Label for Claude Code - Opus 4.8 (max)Claude CodeOpus 4.8(max)Codex - GPT-5.5 (medium)Label for Codex - GPT-5.5 (medium)CodexGPT-5.5(medium)Opencode - Muse Spark 1.1 (xhigh)Label for Opencode - Muse Spark 1.1 (xhigh)OpencodeMuse Spark1.1(xhigh)Claude Code - Opus 4.8 (medium)Label for Claude Code - Opus 4.8 (medium)Claude CodeOpus 4.8(medium)Claude Code - GLM-5.2Label for Claude Code - GLM-5.2Claude CodeGLM-5.2Cursor CLI - Composer 2.5 FastLabel for Cursor CLI - Composer 2.5 FastCursor CLIComposer 2.5FastClaude Code - DeepSeek V4 Pro (high)Label for Claude Code - DeepSeek V4 Pro (high)Claude CodeDeepSeek V4Pro(high)Gemini CLI - Gemini 3.1 Pro (high)Label for Gemini CLI - Gemini 3.1 Pro (high)Gemini CLIGemini 3.1Pro(high)8077777676757371696758524743

Image & Video Leaderboards

Top models from our Image Arena and Video Arena leaderboards, with 95% confidence intervals

Text to Image Leaderboard

Elo scores from blind preference votes in our Image Arena. See the full leaderboard here.
133712781268126312611259125212211216120612061200119711951191GPT Image 2 (high)Logo of GPT Image 2 (high)Reve 2.0Logo of Reve 2.0MAI-Image-2.5Logo of MAI-Image-2.5HiDream-O1-Image-1.5Logo of HiDream-O1-Image-1.5Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Logo of Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)GPT Image 1.5 (high)Logo of GPT Image 1.5 (high)Nano Banana 2 (Gemini 3.1 Flash Image Preview)Logo of Nano Banana 2 (Gemini 3.1 Flash Image Preview)Cosmos3-Super-Text2Image (agentic)Logo of Cosmos3-Super-Text2Image (agentic)Nano Banana Pro (Gemini 3 Pro Image)Logo of Nano Banana Pro (Gemini 3 Pro Image)MAI-Image-2.5-FlashLogo of MAI-Image-2.5-FlashRecraft V4.1 Utility ProLogo of Recraft V4.1 Utility Progrok-imagine-image-qualityLogo of grok-imagine-image-qualityKrea 2 MediumLogo of Krea 2 MediumRecraft V4.1 UtilityLogo of Recraft V4.1 UtilityFLUX.2 [max]Logo of FLUX.2 [max]

Speech Leaderboards

Top models from our Text to Speech Arena, Speech to Text and Speech to Speech evaluations

Text to Speech Arena Leaderboard

Elo scores from blind preference votes in our Text to Speech Arena · See the full leaderboard here.
123412141207120512051204118911851182117511741152114911471144Simba 3.2Logo of Simba 3.2Gemini 3.1 Flash TTSLogo of Gemini 3.1 Flash TTSSonic 3.5Logo of Sonic 3.5Fun-Realtime-TTSLogo of Fun-Realtime-TTSRealtime TTS-2 - Research PreviewLogo of Realtime TTS-2 - Research PreviewRealtime TTS 1.5 MaxLogo of Realtime TTS 1.5 MaxSpaceXAI TTSLogo of SpaceXAI TTSSpeech 2.8 HDLogo of Speech 2.8 HDAsync Flash v1.5Logo of Async Flash v1.5StepAudio 2.5 TTSLogo of StepAudio 2.5 TTSEleven v3Logo of Eleven v3Lightning V3.1 Pro TTS (Jun 2026)Logo of Lightning V3.1 Pro TTS (Jun 2026)Speech 2.8 TurboLogo of Speech 2.8 TurboAsync Pro v1.0Logo of Async Pro v1.0Step TTS 2 (Mar 2026)Logo of Step TTS 2 (Mar 2026)

Relative Elo score of the models as determined by responses from users in Artificial Analysis' Speech Arena. Some models may not be shown due to not yet having enough votes.

Measures the performance of models on specific capabilities and industries

Artificial Analysis Agentic Index

Measures performance in agentic workflows, focusing on behaviors like tool use, planning, autonomy, and complex problem solving.
GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning model5453474747464645433837363531313029272421201916141332
Reasoning models are indicated by a lightbulb icon

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better

Agentic real-world work tasks, (Elo-500)/2000

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model63%62%55%55%55%55%52%51%50%45%44%42%40%39%38%35%34%33%29%23%23%21%20%15%15%0%0%

Agentic tool use

GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning model33%33%32%31%28%28%27%27%27%26%25%25%23%21%16%15%14%14%13%13%12%12%11%9%9%8%5%

Agentic coding & terminal use

GPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model88%88%85%85%84%82%81%81%79%78%78%75%74%66%65%65%64%62%54%51%51%44%43%40%26%15%12%

Coding

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model60%59%58%56%56%54%54%54%53%53%53%53%50%50%50%49%47%45%45%43%43%42%40%40%39%33%25%

Reasoning & knowledge

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelK2 Think V2Logo of K2 Think V2Reasoning model53%47%46%45%45%44%42%41%40%40%40%38%37%37%36%36%35%34%32%27%27%23%18%13%10%10%9%

Scientific reasoning

Gemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning model94%94%94%93%93%93%93%92%92%92%91%91%91%90%90%89%89%89%89%87%87%86%78%75%72%71%67%

Physics reasoning

GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 Pro (xhigh)Logo of GPT-5.5 Pro (xhigh)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning model32%31%30%29%27%21%21%21%18%17%15%15%13%13%13%8%8%7%4%4%3%2%1%1%0%0%0%0%
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning model61%59%57%55%52%52%47%46%43%42%41%38%37%35%33%31%30%25%25%23%22%22%20%18%17%16%15%
MiniMax-M3Logo of MiniMax-M3Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning model84%77%75%75%74%72%71%64%63%62%61%50%46%45%41%39%18%18%15%14%12%11%11%10%9%6%4%

Long context reasoning

GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model74%74%74%74%74%73%73%71%71%70%70%70%69%69%68%68%67%66%66%64%63%63%62%61%53%51%27%

Agentic knowledge work, Elo

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model15831495138913541326126011581110932908873870866863831811752603546506445364128500

Agentic SaaS workflows

Grok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning model51%51%49%49%46%43%43%42%42%39%38%28%26%20%19%17%17%16%14%12%10%8%6%

Legal agentic work, task all-pass rate

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model14%13%8%8%8%7%5%5%4%3%3%3%2%2%1%0%0%0%0%0%0%0%0%0%

Agentic business operations

Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model51%50%47%45%45%44%43%42%41%40%40%39%34%32%31%29%29%28%26%26%

Instruction following

MiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning model83%81%81%81%80%79%79%77%76%76%76%76%76%73%73%71%71%69%69%63%63%62%54%

Long-horizon agentic tasks

Gemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning model47%38%34%32%28%24%17%15%3%2%

Kubernetes incident root-cause analysis

GPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model56%51%46%43%42%40%40%38%38%37%34%33%32%31%30%27%6%

Visual reasoning

Gemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning model84%83%82%81%80%80%79%79%79%78%77%77%73%65%59%
Reasoning models are indicated by a lightbulb icon.

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

AA-Briefcase is a frontier agentic evaluation for long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos

AA-Briefcase Elo

AA-Briefcase is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better
15831495138913541326126011581110932908873870866863831811752603546506445364128500Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning model
Reasoning models are indicated by a lightbulb icon

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

AA-Omniscience is a knowledge and hallucination benchmark that rewards accuracy, punishes bad guesses and provides a comprehensive view of which models produce factually reliable outputs across different domains

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning model403327262322201818151464410−1−4−10−11−23−30−34−36−45−50−54
Reasoning models are indicated by a lightbulb icon

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

GDPval-AA v2 evaluates AI models on real-world, economically valuable tasks across a wide range of occupations

GDPval-AA v2 Leaderboard

Elo rating for performance on real-world work tasks · Anchored to a human baseline of 1,000 · Higher is better
Human Baseline (1,000)
1760174816081600159315921539151414941395137613491307127312651190118911641085962962929907804799493370Claude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning model
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Openness Index assesses how 'open' models are on the basis of their availability and transparency across different components.

Artificial Analysis Openness Index: Components

Openness Index underlying score contribution by components, up to a maximum of 18 (higher is more open)
K2 Think V2Logo of K2 Think V2Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning model6.06.06.06.06.06.06.06.06.04.04.05.02.06.06.03.03.02.01.01.01.01.02.02.01.02.01.516.015.09.09.08.07.07.07.07.06.06.06.02.02.01.5
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index

Most attractive quadrant
1520253035404550556065Artificial Analysis Intelligence Index0102030405060708090100Artificial Analysis Openness Index

Output Tokens

Output tokens of leading AI models based on our independent evaluations

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index
Gemma 4 31BLogo of Gemma 4 31BReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning model8k10k8k10k10k12k11k14k14k15k15k18k15k18k12k18k18k25k28k30k33k35k32k37k35k34k56k12k13k14k14k15k16k17k19k19k20k22k22k23k24k24k25k28k33k36k37k38k39k41k43k45k48k69k6k4k5k6k5k5k5k6k5k8k6k12k7k10k8k8k7k5k8k6k10k14k13k
Reasoning models are indicated by a lightbulb icon

The number of tokens required per Intelligence Index task. This is calculated by multiplying the output tokens per eval by the relative weights of each benchmark in the Intelligence Index, then dividing by task count (excluding repeats).

Price and Cost

Price and real-world costs of leading AI models based on our independent evaluations

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better
DeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$0.24$0.17$0.38$0.58$0.42$0.63$0.20$0.22$0.17$0.87$0.87$0.24$0.31$0.48$0.21$0.37$0.30$0.56$0.81$1.18$0.02$0.03$0.04$0.06$0.12$0.14$0.21$0.24$0.24$0.26$0.29$0.31$0.33$0.35$0.37$0.55$0.59$0.86$1.04$1.06$1.08$1.53$1.80$2.75$0.21$0.37
Reasoning models are indicated by a lightbulb icon

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index
DeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model$470$434$680$743$681$821$540$402$360$622$586$636$715$564$630$724$74$96$98$176$204$288$443$528$539$548$601$815$820$852$870$1,041$1,395$1,631$1,754$2,630$2,824$3,753$4,010$5,631$508
Reasoning models are indicated by a lightbulb icon

The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model's input, cache hit, cache write, reasoning, and answer token prices and the number of tokens used across evaluations (excluding repeats).

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)
DeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning model<0.010.15<0.01<0.010.060.250.20.160.150.260.10.10.50.150.250.150.20.20.250.50.50.510.140.150.440.440.30.681.250.60.951.251.41121.52.51.5222.5555100.280.60.870.871.22.682.53.644.254.45667.57.5910121525303050
Reasoning models are indicated by a lightbulb icon

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

The blended cache price shown here uses cache hit price only. Other caching costs differ by provider:

  • Anthropic: charges a separate cache write fee, with different rates for 5-minute and 1-hour TTLs (1-hour TTL is more expensive).
  • Google (Vertex/Gemini): charges a per-hour cache storage fee in addition to cache hit pricing. Some providers also use tiered pricing for prompts above 200K tokens.
  • OpenAI, DeepSeek, others: typically charge only cache hit pricing with no write or storage fee.

See Prompt Caching for the full breakdown.

Price per token generated by the model (received from the API), represented as USD per million Tokens.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Speed & Latency

Comparison of first-party API performance

Output Speed

Output tokens per second · Higher is better
gpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelGPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning model31921619719719615614113412611911911110910210179736964595956514735
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better
GPT-5.6 Luna (max)Logo of GPT-5.6 Luna (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.7 MaxLogo of Qwen3.7 MaxReasoning modelgpt-oss-120b (high)Logo of gpt-oss-120b (high)Reasoning modelNemotron 3 UltraLogo of Nemotron 3 UltraReasoning modelGrok 4.5 (high)Logo of Grok 4.5 (high)Reasoning modelGrok 4.3 (high)Logo of Grok 4.3 (high)Reasoning modelGPT-5.6 Terra (max)Logo of GPT-5.6 Terra (max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGemini 3.5 FlashLogo of Gemini 3.5 FlashReasoning modelMistral Medium 3.5Logo of Mistral Medium 3.5Reasoning modelMuse Spark 1.1 (xhigh)Logo of Muse Spark 1.1 (xhigh)Reasoning modelGLM-5.2 (max)Logo of GLM-5.2 (max)Reasoning modelGPT-5.6 Sol (max)Logo of GPT-5.6 Sol (max)Reasoning modelMiniMax-M3Logo of MiniMax-M3Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelClaude Fable 5 (with fallback)Logo of Claude Fable 5 (with fallback)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (max)Logo of DeepSeek V4 Flash (max)Reasoning modelClaude Opus 4.8 (max)Logo of Claude Opus 4.8 (max)Reasoning modelDeepSeek V4 Pro (max)Logo of DeepSeek V4 Pro (max)Reasoning modelClaude Sonnet 5 (max)Logo of Claude Sonnet 5 (max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning model1.51.61.71.91.92.02.22.32.82.82.83.03.33.43.53.74.25.45.56.66.97.28.58.613.0
Reasoning models are indicated by a lightbulb icon

The weighted average time (seconds) per Artificial Analysis Intelligence Index task. This is calculated by dividing output tokens per task by output speed, weighted by the relative weights of each benchmark in the Intelligence Index.

API Provider Performance

Output Speed vs. Price: gpt-oss-120b (high)

Output tokens per second · USD per 1M tokens (blended) · 10,000 input tokens
Most attractive quadrant
$0$0.05$0.1$0.15$0.2$0.25$0.3$0.35$0.4$0.45Price (USD per M Tokens)02004006008001.00k1.20k1.40k1.60k1.80k2.00kOutput Speed (Output Tokens per Second)
Reasoning models are indicated by a lightbulb icon.

Smaller, emerging providers are offering high output speed and at competitive prices.

Price per token, shown in USD per million tokens. Price is a blend of cache hit, input, and output token prices using the selected ratio (default 7:2:1 cache-input-output).

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent median (P50) measurement over the past 72 hours to reflect sustained changes in performance.

Pricing (Cache Hit, Input, and Output): gpt-oss-120b (high)

Price (USD per M Tokens) · Lower is better · 10,000 input tokens
CoreWeaveLogo of CoreWeaveDeepInfraLogo of DeepInfraNovitaLogo of NovitaGoogle VertexLogo of Google VertexBasetenLogo of BasetenAmazonLogo of AmazonAzureLogo of AzureDatabricksLogo of DatabricksDeepInfra (Turbo)Logo of DeepInfra (Turbo)FireworksLogo of FireworksGroqLogo of GroqMakoraLogo of MakoraNebius BaseLogo of Nebius BaseTogether AILogo of Together AISambaNovaLogo of SambaNovaParasailLogo of ParasailScalewayLogo of ScalewayCerebrasLogo of CerebrasCloudflareLogo of Cloudflare0.080.060.040.040.050.090.10.150.150.150.150.150.150.150.150.150.220.10.170.350.350.140.170.250.360.50.60.60.60.60.60.60.60.60.60.590.750.70.750.75

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

The blended cache price shown here uses cache hit price only. Other caching costs differ by provider:

  • Anthropic: charges a separate cache write fee, with different rates for 5-minute and 1-hour TTLs (1-hour TTL is more expensive).
  • Google (Vertex/Gemini): charges a per-hour cache storage fee in addition to cache hit pricing. Some providers also use tiered pricing for prompts above 200K tokens.
  • OpenAI, DeepSeek, others: typically charge only cache hit pricing with no write or storage fee.

See Prompt Caching for the full breakdown.

Price per token generated by the model (received from the API), represented as USD per million Tokens.

Output Speed: gpt-oss-120b (high)

Output speed: output tokens per second · 10,000 input tokens
CerebrasLogo of CerebrasSambaNovaLogo of SambaNovaFireworksLogo of FireworksTogether AILogo of Together AIGroqLogo of GroqGoogle VertexLogo of Google VertexAzureLogo of AzureNebius BaseLogo of Nebius BaseDatabricksLogo of DatabricksDeepInfra (Turbo)Logo of DeepInfra (Turbo)BasetenLogo of BasetenParasailLogo of ParasailScalewayLogo of ScalewayMakoraLogo of MakoraCloudflareLogo of CloudflareAmazonLogo of AmazonNovitaLogo of NovitaCoreWeaveLogo of CoreWeaveDeepInfraLogo of DeepInfra1765695673570478421362358341273269227182179177120965243

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).