Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Jun 19, 2026
773,605 sessions
28 models
Model
1
11
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
14.00%±1.53%
16.27%±2.58%29.86%±5.65%12.24%±2.92%10.08%±1.78%1.56%±0.15%16,082
2
27
Anthropic
Anthropic · Proprietary
8.89%±1.18%
10.65%±2.11%15.23%±4.20%9.11%±2.19%9.00%±1.30%0.44%±0.76%27,769
3
29
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.04%±1.41%
5.13%±2.62%14.82%±5.22%3.73%±2.76%14.96%±1.26%1.56%±0.15%13,300
4
29
Anthropic
Anthropic · Proprietary
7.98%±1.23%
3.53%±2.58%11.79%±4.46%9.26%±2.31%13.88%±0.82%1.46%±0.17%28,493
5
29
OpenAI · Proprietary
7.96%±0.89%
6.33%±1.66%10.87%±3.14%7.61%±1.69%13.42%±1.10%1.56%±0.15%38,401
6
29
Anthropic
Anthropic · Proprietary
7.83%±1.26%
5.00%±2.52%10.73%±4.55%9.46%±2.31%12.47%±1.14%1.50%±0.16%28,471
7
210
Anthropic
Anthropic · Proprietary
7.03%±1.22%
5.11%±2.50%9.75%±4.14%7.27%±2.25%11.49%±1.59%1.55%±0.15%28,327
8
310
OpenAI · Proprietary
6.80%±0.86%
4.97%±1.62%8.43%±3.07%7.45%±1.60%11.59%±1.14%1.56%±0.15%38,793
9
310
OpenAI · Proprietary
6.58%±0.88%
5.61%±1.71%5.28%±3.11%8.11%±1.73%12.32%±0.92%1.56%±0.15%38,390
10
713
Z.ai · MIT · SiliconFlow
4.40%±1.77%
9.96%±3.23%12.69%±6.49%5.80%±3.38%3.57%±1.98%1.56%±0.15%12,237
11
1013
Anthropic
Anthropic · Proprietary
4.09%±1.38%
5.60%±2.48%11.95%±4.48%6.98%±2.43%8.88%±1.26%12.96%±2.92%25,069
12
1013
Anthropic
Anthropic · Proprietary
3.05%±1.12%
1.46%±2.55%3.46%±3.53%3.86%±2.16%11.84%±1.70%1.54%±0.15%28,435
13
1013
Z.ai · MIT · SiliconFlow
2.01%±0.96%
3.30%±1.98%0.53%±3.31%0.39%±1.96%5.08%±1.05%1.56%±0.15%31,007
14
1419
Google · Proprietary
0.03%±0.84%
1.06%±1.85%1.71%±2.74%1.19%±1.60%2.62%±1.23%1.17%±0.20%32,864
15
1420
Google · Proprietary
0.47%±0.79%
0.26%±1.73%0.98%±2.51%2.35%±1.47%5.47%±1.39%1.50%±0.16%38,750
16
1420
DeepSeek · MIT · SiliconFlow
0.81%±1.25%
0.75%±2.76%1.84%±4.32%3.54%±2.56%2.26%±1.14%0.18%±0.33%26,289
17
1420
Moonshot · Modified MIT · Fireworks
1.04%±0.88%
0.43%±1.88%2.97%±2.91%3.47%±1.77%0.11%±1.34%1.56%±0.15%36,268
18
1420
Kimi K2.7 Code
Moonshot · Modified MIT · Fireworks
1.28%±1.55%
3.22%±2.83%0.45%±5.13%7.88%±3.08%2.86%±3.20%1.56%±0.15%15,957
19
1420
DeepSeek · MIT · SiliconFlow
1.70%±1.09%
4.35%±2.10%1.60%±3.76%7.70%±2.25%3.03%±1.55%0.54%±0.39%30,471
20
1521
MiniMax · Proprietary · Fireworks
2.18%±1.23%
1.17%±2.77%6.97%±4.07%7.87%±2.63%3.55%±1.15%1.56%±0.15%15,100
21
2022
Alibaba · Proprietary · Fireworks
4.20%±0.97%
0.56%±1.97%6.90%±3.21%9.74%±2.05%1.60%±1.36%2.21%±0.53%32,582
22
2225
Grok Build 0.1
xAI · Proprietary
6.35%±0.95%
6.85%±2.22%11.34%±2.99%9.51%±2.00%2.09%±1.42%1.94%±0.35%29,766
23
2226
Grok 4.3 (High)
xAI · Proprietary
7.08%±1.05%
8.97%±2.41%15.20%±2.87%6.28%±1.93%4.61%±2.58%0.33%±0.42%18,337
24
2126
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
7.53%±3.74%
5.54%±6.82%2.30%±13.04%20.07%±7.28%10.39%±6.73%0.63%±1.07%4,610
25
2226
MiniMax · Modified MIT · Fireworks
7.94%±0.87%
12.51%±2.09%15.42%±2.57%9.64%±1.71%3.61%±1.56%1.46%±0.18%32,712
26
2326
Google · Proprietary
8.39%±0.85%
11.32%±1.81%13.09%±2.18%4.65%±1.49%13.86%±2.45%0.98%±0.38%38,781
27
2727
Google · Apache 2.0
13.01%±1.71%
5.55%±2.04%7.94%±2.97%6.17%±1.89%27.80%±5.64%17.59%±4.84%28,265
28
2828
xAI · Proprietary
17.72%±1.26%
12.59%±1.85%14.98%±2.25%5.33%±1.49%56.46%±5.17%0.73%±0.31%38,079
Signal Leaders
  1. AnthropicClaude Fable 5 (High)gets users to confirm the task is done most often16.27%±2.58%
  2. AnthropicClaude Fable 5 (High)draws the most positive responses relative to negative ones29.86%±5.65%
  3. AnthropicClaude Fable 5 (High)lands user corrections best12.24%±2.92%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.96%±1.26%
  5. Kimi K2.7 Codeleast likely to hallucinate tools it doesn't have1.56%±0.15%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. AnthropicClaude Fable 5 (High)16.27%
  2. AnthropicClaude Opus 4.8 (Thinking)10.65%
  3. GLM 5.2 (Max)9.96%
  4. GPT 5.5 (High)6.33%
  5. GPT 5.4 (High)5.61%
  6. AnthropicClaude Opus 4.85.60%
  7. GPT 5.5 (xHigh)5.13%
  8. AnthropicClaude Opus 4.65.11%
  9. AnthropicClaude Opus 4.75.00%
  10. GPT 5.54.97%
294,569 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. AnthropicClaude Fable 5 (High)29.86%
  2. AnthropicClaude Opus 4.8 (Thinking)15.23%
  3. GPT 5.5 (xHigh)14.82%
  4. GLM 5.2 (Max)12.69%
  5. AnthropicClaude Opus 4.811.95%
  6. AnthropicClaude Opus 4.7 (Thinking)11.79%
  7. GPT 5.5 (High)10.87%
  8. AnthropicClaude Opus 4.710.73%
  9. AnthropicClaude Opus 4.69.75%
  10. GPT 5.58.43%
102,234 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. AnthropicClaude Fable 5 (High)12.24%
  2. AnthropicClaude Opus 4.79.46%
  3. AnthropicClaude Opus 4.7 (Thinking)9.26%
  4. AnthropicClaude Opus 4.8 (Thinking)9.11%
  5. GPT 5.4 (High)8.11%
  6. GPT 5.5 (High)7.61%
  7. GPT 5.57.45%
  8. AnthropicClaude Opus 4.67.27%
  9. AnthropicClaude Opus 4.86.98%
  10. AnthropicClaude Sonnet 4.63.86%
174,327 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. GPT 5.5 (xHigh)14.96%
  2. AnthropicClaude Opus 4.7 (Thinking)13.88%
  3. GPT 5.5 (High)13.42%
  4. AnthropicClaude Opus 4.712.47%
  5. GPT 5.4 (High)12.32%
  6. AnthropicClaude Sonnet 4.611.84%
  7. GPT 5.511.59%
  8. AnthropicClaude Opus 4.611.49%
  9. AnthropicClaude Fable 5 (High)10.08%
  10. AnthropicClaude Opus 4.8 (Thinking)9.00%
166,868 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. Kimi K2.7 Code1.56%
  2. Kimi K2.61.56%
  3. GLM 5.11.56%
  4. AnthropicClaude Fable 5 (High)1.56%
  5. GLM 5.2 (Max)1.56%
  6. GPT 5.5 (xHigh)1.56%
  7. Minimax M31.56%
  8. GPT 5.4 (High)1.56%
  9. GPT 5.5 (High)1.56%
  10. GPT 5.51.56%
641,583 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology