Independent analysis of AI

Understand the AI landscape to choose the best model and provider for your use case

Highlights

Intelligence
Artificial Analysis Intelligence Index · Higher is better
Intelligence: Artificial Analysis Intelligence Index · Higher is better
  1. GPT-5.5 (xhigh): 60
  2. Claude Opus 4.7 (max): 57
  3. Gemini 3.1 Pro Preview: 57
  4. GPT-5.4 (xhigh): 57
  5. Kimi K2.6: 54
  6. MiMo-V2.5-Pro: 54
  7. Grok 4.3: 53
  8. Muse Spark: 52
  9. DeepSeek V4 Pro (Max): 52
  10. NVIDIA Nemotron 3 Super: 36
  11. gpt-oss-120B (high): 33
Speed
Output tokens per second · Higher is better
Speed: Output tokens per second · Higher is better
  1. gpt-oss-120B (high): 229
  2. NVIDIA Nemotron 3 Super: 183
  3. Gemini 3.1 Pro Preview: 136
  4. Grok 4.3: 103
  5. GPT-5.4 (xhigh): 86
  6. GPT-5.5 (xhigh): 72
  7. MiMo-V2.5-Pro: 63
  8. Claude Opus 4.7 (max): 49
  9. Kimi K2.6: 34
  10. DeepSeek V4 Pro (Max): 34
Price
USD per 1M tokens (3:1 input-output ratio) · Lower is better
Price: USD per 1M tokens (3:1 input-output ratio) · Lower is better
  1. gpt-oss-120B (high): 0.3
  2. NVIDIA Nemotron 3 Super: 0.4
  3. MiMo-V2.5-Pro: 1.5
  4. Grok 4.3: 1.6
  5. Kimi K2.6: 1.7
  6. DeepSeek V4 Pro (Max): 2.2
  7. Gemini 3.1 Pro Preview: 4.5
  8. GPT-5.4 (xhigh): 5.6
  9. Claude Opus 4.7 (max): 10.9
  10. GPT-5.5 (xhigh): 11.3
Get personalized recommendations based on your priorities for intelligence, speed, and cost.
Personalized model recommendation
Compare AI agents across capabilities, pricing, and platform support.
Explore agents for general work, coding, customer support, and more
What is the Artificial Analysis Intelligence Index?
Learn about the Artificial Analysis Intelligence Index and how it is calculated

Changelog

New language model evaluation · 4 May
Nemotron 3 Nano Omni 30B A3B ReasoningNemotron 3 Nano Omni 30B A3B Reasoning
New feature launched · 2 May
Artificial AnalysisCache pricing now available in language model pricing
New article published · 30 Apr
Recent open weights model launches
New article published · 30 Apr
xAI launches Grok 4.3 with improved agentic performance and lower pricing
New language model evaluation · 30 Apr
Hy3-preview (Non-reasoning)Hy3-preview (Non-reasoning)
New language model evaluation · 30 Apr
Grok 4.3Grok 4.3
New language model evaluation · 30 Apr
Mistral Medium 3.5Mistral Medium 3.5
New language model evaluation · 30 Apr
GPT-5.5 Pro (xhigh)GPT-5.5 Pro (xhigh)
New language model evaluation · 29 Apr
MiMo-V2.5-Pro (Non-reasoning)MiMo-V2.5-Pro (Non-reasoning)
New language model evaluation · 29 Apr
DeepSeek V4 Flash (Non-reasoning)DeepSeek V4 Flash (Non-reasoning)
New language model evaluation · 29 Apr
DeepSeek V4 Pro (Non-reasoning)DeepSeek V4 Pro (Non-reasoning)
New language model evaluation · 29 Apr
Kimi K2.6 (Non-reasoning)Kimi K2.6 (Non-reasoning)
New language model evaluation · 29 Apr
Granite 4.1 3BGranite 4.1 3B
New language model evaluation · 29 Apr
Granite 4.1 30BGranite 4.1 30B
New language model evaluation · 29 Apr
Granite 4.1 8BGranite 4.1 8B
New language model evaluation · 27 Apr
Hy3-preview (Reasoning)Hy3-preview (Reasoning)
New language model evaluation · 27 Apr
EXAONE 4.5 33B (Non-reasoning)EXAONE 4.5 33B (Non-reasoning)
New language model evaluation · 27 Apr
EXAONE 4.5 33BEXAONE 4.5 33B
New language model evaluation · 27 Apr
MiMo-V2.5MiMo-V2.5
New article published · 24 Apr
DeepSeek is back among the leading open weights models with V4 Pro and V4 FlashSee more

Intelligence

Intelligence of leading AI models based on our independent evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelMuse SparkLogo of Muse SparkReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning model605757575454535252525251504947464542393736363328262424
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.0 includes: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelMuse SparkLogo of Muse SparkReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning model605757575454535252525251504947464542393736363328262424
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Intelligence Index v4.0 includes: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).

Intelligence vs. Cost to Run Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index · Cost to run Intelligence Index
Most attractive quadrant
Alibaba
Amazon
Anthropic
DeepSeek
Google
Kimi
MiniMax
Mistral
NVIDIA
OpenAI
xAI
Xiaomi
Z AI
32641282565121.02k2.05k4.10k8.19kCost to Run Intelligence Index (USD, Log Scale)20253035404550556065Artificial Analysis Intelligence Indexgpt-oss-20B (high)Mistral Small 4gpt-oss-120B (high)DeepSeek V3.2DeepSeek V4 Flash (Max)NVIDIA Nemotron 3 SuperMiniMax-M2.7Gemini 3 FlashGrok 4.3Qwen3.5 397B A17BMiMo-V2.5-ProNova 2.0 Pro Preview (medium)GLM-5.1Claude 4.5 HaikuQwen3.6 Max PreviewGemini 3.1 Pro PreviewKimi K2.6DeepSeek V4 Pro (Max)GPT-5.4 mini (xhigh)GPT-5.4 (xhigh)GPT-5.5 (xhigh)Claude Sonnet 4.6 (max)Claude Opus 4.7 (max)
Reasoning models are indicated by a lightbulb icon.

The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model's input and output token pricing and the number of tokens used across evaluations (excluding repeats).

Artificial Analysis Intelligence Index v4.0 includes: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt
Alibaba
Anthropic
DeepSeek
Google
Kimi
MBZUAI Institute of Foundation Models
Meta
MiniMax
Mistral
OpenAI
Upstage
xAI
Xiaomi
Z AI
Nov ’22Jan ’23Mar ’23May ’23Jul ’23Sep ’23Nov ’23Jan ’24Mar ’24May ’24Jul ’24Sep ’24Nov ’24Jan ’25Mar ’25May ’25Jul ’25Sep ’25Nov ’25Jan ’26Mar ’26May ’26Jul ’26Release Date0510152025303540455055606570Artificial Analysis Intelligence Index
Reasoning models are indicated by a lightbulb icon.

Artificial Analysis Intelligence Index v4.0 includes: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Image & Video Leaderboards

Top models from our Image Arena and Video Arena leaderboards, with 95% confidence intervals

Text to Image Leaderboard

Elo scores from blind preference votes in our Image Arena. See the full leaderboard here.
133812731261121912011201119511891188118311811181117311691164GPT Image 2 (high)Logo of GPT Image 2 (high)GPT Image 1.5 (high)Logo of GPT Image 1.5 (high)Nano Banana 2 (Gemini 3.1 Flash Image Preview)Logo of Nano Banana 2 (Gemini 3.1 Flash Image Preview)Nano Banana Pro (Gemini 3 Pro Image)Logo of Nano Banana Pro (Gemini 3 Pro Image)Seedream 4.0Logo of Seedream 4.0FLUX.2 [max]Logo of FLUX.2 [max]MAI-Image-2Logo of MAI-Image-2Peanut (Open Weights Coming Soon)Logo of Peanut (Open Weights Coming Soon)FLUX.2 [pro]Logo of FLUX.2 [pro]FLUX.2 [flex]Logo of FLUX.2 [flex]grok-imagine-imageLogo of grok-imagine-imageImagineArt 2.0Logo of ImagineArt 2.0Imagen 4 UltraLogo of Imagen 4 UltraSeedream 4.5Logo of Seedream 4.5FLUX.2 [dev] TurboLogo of FLUX.2 [dev] Turbo

Intelligence Breakdown

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Results claimed by AI Lab (not yet independently verified)
GDPval-AA
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning model64%63%59%59%54%53%52%50%50%50%49%47%46%44%41%35%35%35%34%31%25%24%22%18%9%8%5%
Terminal-Bench Hard
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning model61%58%54%53%52%52%46%46%44%44%43%43%41%39%39%38%36%36%36%29%27%24%24%17%11%8%7%
𝜏²-Bench Telecom
Grok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning model98%98%96%96%96%96%96%95%94%94%93%92%91%89%87%86%85%83%80%76%68%66%60%60%55%41%25%
AA-LCR
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning model74%74%73%73%71%70%70%70%70%70%69%69%66%66%66%65%64%63%62%62%60%54%53%51%45%31%27%
AA-Omniscience Accuracy
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning model57%55%54%50%46%45%43%40%38%37%37%35%33%33%31%26%24%24%23%22%22%22%20%18%17%16%16%
AA-Omniscience Non-Hallucination Rate
MiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning model75%75%74%71%66%64%61%56%54%50%41%33%27%18%18%14%13%12%11%11%10%10%9%8%6%6%4%
Humanity's Last Exam
Gemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning model45%44%42%40%40%36%36%35%35%34%32%30%29%28%28%27%27%23%22%19%19%10%10%10%10%10%9%
GPQA Diamond
Gemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning model94%94%92%91%91%90%90%89%89%89%89%88%88%88%87%87%87%86%84%80%79%78%77%72%71%69%67%
SciCode
Gemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning model59%57%56%55%54%52%51%50%50%50%47%47%47%47%45%44%43%43%43%42%39%39%38%36%34%33%25%
IFBench
Grok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning model81%80%79%79%79%78%77%77%77%76%76%76%76%76%76%74%73%72%71%69%65%63%61%59%57%54%48%
CritPt
GPT-5.5 Pro (xhigh)Logo of GPT-5.5 Pro (xhigh) which relates to the data aboveGPT-5.5 Pro (xhigh)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4 Pro(Max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4 Flash(Max)Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 Max PreviewReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning model31%27%23%18%13%12%11%10%9%8%8%7%5%4%4%3%3%3%2%1%1%1%1%0%0%0%0%0%
APEX-Agents-AA
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeek V3.2Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B (high)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIA Nemotron 3SuperReasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B (high)Reasoning model38%33%32%28%28%28%15%14%11%3%2%1%
MMMU-Pro
Gemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1 ProPreviewReasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5 (xhigh)Reasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3 FlashReasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus 4.7(max)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4 (xhigh)Reasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397B A17BReasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaude Sonnet 4.6(max)Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview (medium)Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5 HaikuReasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistral Small 4Reasoning model82%81%80%80%79%79%78%78%77%73%73%73%65%59%57%
Reasoning models are indicated by a lightbulb icon.

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

Artificial Analysis Intelligence Index v4.0 includes: GDPval-AA, 𝜏²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

AA-Omniscience is a knowledge and hallucination benchmark that rewards accuracy, punishes bad guesses and provides a comprehensive view of which models produce factually reliable outputs across different domains

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
Gemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelMuse SparkLogo of Muse SparkReasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning model33262018121210664421−4−10−19−21−23−30−30−34−42−45−48−50−54−64
Reasoning models are indicated by a lightbulb icon

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

GDPval-AA evaluates AI models on real-world, economically valuable tasks across a wide range of occupations

GDPval-AA Leaderboard

Elo scores for agentic performance on real-world work tasks using web and shell access via Stirrup, an open-source harness developed by Artificial Analysis
GPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh) which relates to the data aboveGPT-5.5(xhigh)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max) which relates to the data aboveClaude Opus4.7 (max)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max) which relates to the data aboveClaudeSonnet 4.6(max) Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh) which relates to the data aboveGPT-5.4(xhigh)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-Pro which relates to the data aboveMiMo-V2.5-ProReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max) which relates to the data aboveDeepSeek V4Pro (Max)Reasoning modelGLM-5.1Logo of GLM-5.1 which relates to the data aboveGLM-5.1Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7 which relates to the data aboveMiniMax-M2.7Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max Preview which relates to the data aboveQwen3.6 MaxPreviewReasoning modelGrok 4.3Logo of Grok 4.3 which relates to the data aboveGrok 4.3Reasoning modelKimi K2.6Logo of Kimi K2.6 which relates to the data aboveKimi K2.6Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh) which relates to the data aboveGPT-5.4 mini(xhigh)Reasoning modelMuse SparkLogo of Muse Spark which relates to the data aboveMuse SparkReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max) which relates to the data aboveDeepSeek V4Flash (Max)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro Preview which relates to the data aboveGemini 3.1Pro PreviewReasoning modelGemini 3 FlashLogo of Gemini 3 Flash which relates to the data aboveGemini 3FlashReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2 which relates to the data aboveDeepSeekV3.2Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17B which relates to the data aboveQwen3.5 397BA17BReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 Haiku which relates to the data aboveClaude 4.5HaikuReasoning modelGemma 4 31BLogo of Gemma 4 31B which relates to the data aboveGemma 4 31BReasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 Super which relates to the data aboveNVIDIANemotron 3Super Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium) which relates to the data aboveNova 2.0 ProPreview(medium) Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high) which relates to the data abovegpt-oss-120B(high)Reasoning modelMistral Small 4Logo of Mistral Small 4 which relates to the data aboveMistralSmall 4Reasoning modelSolar Pro 3Logo of Solar Pro 3 which relates to the data aboveSolar Pro 3Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high) which relates to the data abovegpt-oss-20B(high)Reasoning modelK2 Think V2Logo of K2 Think V2 which relates to the data aboveK2 Think V2Reasoning model177417531675167415731554153515081506150014831437142213881314120511991193117411151005974947862676652608

Artificial Analysis Openness Index assesses how 'open' models are on the basis of their availability and transparency across different components.

Artificial Analysis Openness Index: Components

Openness Index underlying score contribution by components, up to a maximum of 18 (higher is more open)
K2 Think V2Logo of K2 Think V2Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning model6.06.06.06.06.06.06.06.06.06.06.04.03.02.06.06.03.03.02.01.01.01.01.01.01.02.01.02.01.516.015.09.09.08.07.07.07.07.07.07.06.04.02.02.01.5
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index

Artificial Analysis Openness Index · Artificial Analysis Intelligence Index
Most attractive quadrant
Alibaba
Anthropic
DeepSeek
Google
Kimi
MBZUAI Institute of Foundation Models
MiniMax
Mistral
NVIDIA
OpenAI
Xiaomi
Z AI
20253035404550556065Artificial Analysis Intelligence Index0102030405060708090100Artificial Analysis Openness IndexClaude 4.5 Haikugpt-oss-20B (high)Mistral Small 4gpt-oss-120B (high)MiniMax-M2.7Gemma 4 31BQwen3.5 397B A17BKimi K2.6MiMo-V2.5-ProGLM-5.1DeepSeek V4 Flash (Max)DeepSeek V4 Pro (Max)NVIDIA Nemotron 3 SuperK2 Think V2

Output Tokens

Output tokens of leading AI models based on our independent evaluations

Output Tokens Used to Run Artificial Analysis Intelligence Index

Tokens used to run all evaluations in the Artificial Analysis Intelligence Index
DeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelSolar Pro 3Logo of Solar Pro 3Reasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelK2 Think V2Logo of K2 Think V2Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelDeepSeek V3.2Logo of DeepSeek V3.2Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning modelMuse SparkLogo of Muse SparkReasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning model110M110M93M95M82M80M79M79M80M73M68M65M68M57M58M53M53M49M32M33M240M240M200M190M170M120M120M110M110M110M99M92M88M87M87M86M78M75M74M72M61M61M58M57M53M39M36M13M17M19M
Reasoning models are indicated by a lightbulb icon

The number of tokens required to run all evaluations in the Artificial Analysis Intelligence Index (excluding repeats).

Cost Efficiency

Cost of leading AI models based on our independent evaluations

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index
Claude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelgpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning model$479$1,717$614$626$636$509$396$485$5,117$4,206$3,357$2,851$1,354$1,071$948$892$861$620$544$467$462$418$395$278$176$145$113$76$67$48$19$1,101$420
Reasoning models are indicated by a lightbulb icon

The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model's input and output token pricing and the number of tokens used across evaluations (excluding repeats).

Speed & Latency

Comparison of first-party API performance

Output Speed

Output tokens per second · Higher is better
gpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelGemma 4 31BLogo of Gemma 4 31BReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning model2752291831771741661621361039986767263605953534936353434
Reasoning models are indicated by a lightbulb icon

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

Price

Price of leading AI models based on our independent evaluations

Pricing now includes a “Cache Hit Price” alongside Input and Output pricing, with new blend ratios.

Pricing: Cache Hit, Input, and Output

Price: USD per 1M tokens
gpt-oss-20B (high)Logo of gpt-oss-20B (high)Reasoning modelDeepSeek V4 Flash (Max)Logo of DeepSeek V4 Flash (Max)Reasoning modelgpt-oss-120B (high)Logo of gpt-oss-120B (high)Reasoning modelMistral Small 4Logo of Mistral Small 4Reasoning modelDeepSeek V3.2Logo of DeepSeek V3.2Reasoning modelNVIDIA Nemotron 3 SuperLogo of NVIDIA Nemotron 3 SuperReasoning modelMiniMax-M2.7Logo of MiniMax-M2.7Reasoning modelGemini 3 FlashLogo of Gemini 3 FlashReasoning modelGrok 4.3Logo of Grok 4.3Reasoning modelMiMo-V2.5-ProLogo of MiMo-V2.5-ProReasoning modelQwen3.5 397B A17BLogo of Qwen3.5 397B A17BReasoning modelKimi K2.6Logo of Kimi K2.6Reasoning modelDeepSeek V4 Pro (Max)Logo of DeepSeek V4 Pro (Max)Reasoning modelGPT-5.4 mini (xhigh)Logo of GPT-5.4 mini (xhigh)Reasoning modelGLM-5.1Logo of GLM-5.1Reasoning modelClaude 4.5 HaikuLogo of Claude 4.5 HaikuReasoning modelQwen3.6 Max PreviewLogo of Qwen3.6 Max PreviewReasoning modelNova 2.0 Pro Preview (medium)Logo of Nova 2.0 Pro Preview (medium)Reasoning modelGemini 3.1 Pro PreviewLogo of Gemini 3.1 Pro PreviewReasoning modelGPT-5.4 (xhigh)Logo of GPT-5.4 (xhigh)Reasoning modelClaude Sonnet 4.6 (max)Logo of Claude Sonnet 4.6 (max)Reasoning modelClaude Opus 4.7 (max)Logo of Claude Opus 4.7 (max)Reasoning modelGPT-5.5 (xhigh)Logo of GPT-5.5 (xhigh)Reasoning model<0.010.150.140.20.060.050.20.20.160.010.080.260.10.130.310.20.250.30.50.50.050.140.150.150.30.30.30.51.2510.60.951.740.751.41.251.31.2522.53.756.2550.20.280.60.60.450.751.232.533.643.484.54.457.8101215152530
Reasoning models are indicated by a lightbulb icon

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

The blended bar shown here uses cache hit price only. Other caching costs differ by provider:

  • Anthropic: charges a separate cache write fee, with different rates for 5-minute and 1-hour TTLs (1-hour TTL is more expensive). Blended price charts use Anthropic cache write price for the input leg.
  • Google (Vertex/Gemini): charges a per-hour cache storage fee in addition to cache hit pricing. Some providers also use tiered pricing for prompts above 200K tokens.
  • OpenAI, DeepSeek, others: typically charge only cache hit pricing with no write or storage fee.

See Prompt Caching for the full breakdown.

Price per token generated by the model (received from the API), represented as USD per million Tokens.

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).

API Provider Performance

Output Speed vs. Price: gpt-oss-120B (high)

Output tokens per second · USD per 1M tokens · 10,000 Input Tokens
Most attractive quadrant
Amazon
Azure
Baseten
Cerebras
Clarifai
Cloudflare
Databricks
DeepInfra
DeepInfra (Turbo)
Eigen AI
Fireworks
Google Vertex
Groq
Lightning AI
Nebius Base
Nebius Fast
Novita
Parasail
SambaNova
Scaleway
Together.ai
Weights & Biases
$0.05$0.10$0.15$0.20$0.25$0.30$0.35$0.40$0.45$0.50Price (USD per M Tokens)02004006008001.00k1.20k1.40k1.60k1.80k2.00k2.20kOutput Speed (Output Tokens per Second)ParasailDeepInfraNovitaTogether.aiFireworksWeights & BiasesCloudflareDeepInfra (Turbo)ScalewayLightning AINebius BaseBasetenAmazon BedrockDatabricksMicrosoft AzureGoogle VertexGroqClarifaiEigen AINebius FastSambaNovaCerebras
Reasoning models are indicated by a lightbulb icon.

Smaller, emerging providers are offering high output speed and at competitive prices.

Price per token, represented as USD per million Tokens. Price is a blend of cache hit, input, and output token prices using the selected ratio.

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent median (P50) measurement over the past 72 hours to reflect sustained changes in performance.

Pricing (Cache Hit, Input, and Output): gpt-oss-120B (high)

USD per 1M tokens · Lower is better · 10,000 Input Tokens
DeepInfraLogo of DeepInfraNovitaLogo of NovitaGoogle VertexLogo of Google VertexClarifaiLogo of ClarifaiLightning AILogo of Lightning AIBasetenLogo of BasetenNebius FastLogo of Nebius FastEigen AILogo of Eigen AIDatabricksLogo of DatabricksAzureLogo of AzureAmazonLogo of AmazonTogether.aiLogo of Together.aiDeepInfra (Turbo)Logo of DeepInfra (Turbo)Weights & BiasesLogo of Weights & BiasesFireworksLogo of FireworksNebius BaseLogo of Nebius BaseGroqLogo of GroqSambaNovaLogo of SambaNovaParasailLogo of ParasailScalewayLogo of ScalewayCerebrasLogo of CerebrasCloudflareLogo of Cloudflare0.080.060.040.050.090.090.10.10.10.10.150.150.150.150.150.150.150.150.150.220.10.170.350.350.190.250.360.360.40.50.50.50.60.60.60.60.60.60.60.60.60.590.750.70.750.75

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Price per token included in the request/message sent to the API, represented as USD per million Tokens.

The blended bar shown here uses cache hit price only. Other caching costs differ by provider:

  • Anthropic: charges a separate cache write fee, with different rates for 5-minute and 1-hour TTLs (1-hour TTL is more expensive). Blended price charts use Anthropic cache write price for the input leg.
  • Google (Vertex/Gemini): charges a per-hour cache storage fee in addition to cache hit pricing. Some providers also use tiered pricing for prompts above 200K tokens.
  • OpenAI, DeepSeek, others: typically charge only cache hit pricing with no write or storage fee.

See Prompt Caching for the full breakdown.

Price per token generated by the model (received from the API), represented as USD per million Tokens.

Output Speed: gpt-oss-120B (high)

Output speed: output tokens per second · 10,000 Input Tokens
CerebrasLogo of CerebrasSambaNovaLogo of SambaNovaNebius FastLogo of Nebius FastEigen AILogo of Eigen AIClarifaiLogo of ClarifaiGroqLogo of GroqGoogle VertexLogo of Google VertexAzureLogo of AzureDatabricksLogo of DatabricksAmazonLogo of AmazonBasetenLogo of BasetenNebius BaseLogo of Nebius BaseLightning AILogo of Lightning AIScalewayLogo of ScalewayDeepInfra (Turbo)Logo of DeepInfra (Turbo)CloudflareLogo of CloudflareWeights & BiasesLogo of Weights & BiasesFireworksLogo of FireworksTogether.aiLogo of Together.aiNovitaLogo of NovitaDeepInfraLogo of DeepInfraParasailLogo of Parasail18996956616145044754442662662562392312131841761231177167433622

Tokens per second received while the model is generating tokens (ie. after first chunk has been received from the API for models which support streaming).

Figures represent performance of the model's first-party API (e.g. OpenAI for o1) or the median across providers where a first-party API is not available (e.g. Meta's Llama models).