• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent HighView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Jul 21, 2026
1,242,857 sessions
38 models
Rank by
Model
1
14
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
12.72%±2.00%
10.67%±3.84%23.94%±7.42%14.62%±3.80%12.97%±1.30%1.39%±0.17%23,549
2
18
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
10.12%±1.69%
7.25%±3.29%23.53%±6.57%9.71%±2.78%8.74%±1.30%1.39%±0.17%15,991
3
19
Anthropic
Claude Opus 4.8 (Thinking)
Anthropic · Proprietary
9.75%±1.39%
8.90%±2.62%19.42%±5.05%9.78%±2.58%10.43%±1.07%0.22%±1.11%34,147
4
19
Kimi K3
Moonshot · Proprietary
9.71%±1.52%
14.00%±2.92%20.30%±5.45%6.52%±3.14%6.33%±1.27%1.39%±0.17%11,490
5
212
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
8.66%±1.89%
8.14%±3.67%16.88%±7.14%6.20%±3.66%10.81%±0.90%1.25%±0.18%24,359
6
210
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.41%±0.87%
6.65%±1.78%11.08%±3.13%8.18%±1.65%14.77%±0.80%1.39%±0.17%40,667
7
212
Anthropic
Claude Opus 4.7 (Thinking)
Anthropic · Proprietary
7.94%±1.24%
5.67%±2.55%11.55%±4.36%8.62%±2.36%12.57%±1.13%1.28%±0.19%35,151
8
212
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
7.67%±1.25%
4.97%±2.57%12.48%±4.38%8.95%±2.32%10.62%±1.53%1.33%±0.17%35,672
9
312
GPT 5.5 (High)
OpenAI · Proprietary
7.61%±0.81%
6.20%±1.59%9.80%±2.89%8.77%±1.44%11.90%±1.07%1.39%±0.17%65,859
10
614
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
6.50%±1.00%
8.65%±1.97%12.94%±3.63%4.71%±1.79%4.78%±1.15%1.39%±0.17%38,221
11
515
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
6.42%±1.24%
3.12%±2.63%9.94%±4.21%6.53%±2.28%11.14%±1.35%1.39%±0.17%34,862
12
1015
GPT 5.5
OpenAI · Proprietary
5.65%±0.76%
3.92%±1.58%5.67%±2.65%6.08%±1.39%11.22%±0.90%1.39%±0.17%66,796
13
1015
GPT 5.4 (High)
OpenAI · Proprietary
5.64%±0.77%
6.23%±1.59%3.13%±2.70%7.75%±1.46%9.72%±0.90%1.39%±0.17%66,142
14
615
Grok 4.5
SpaceXAI · Proprietary
5.56%±1.33%
3.86%±2.88%8.17%±4.91%3.80%±2.34%10.56%±1.14%1.39%±0.17%21,424
15
1117
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
3.56%±1.65%
7.10%±2.73%11.63%±4.80%8.25%±2.62%9.82%±1.40%18.98%±4.59%32,216
16
1517
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
2.84%±1.15%
0.62%±2.62%0.65%±3.77%1.35%±2.18%11.45%±1.47%1.35%±0.17%35,646
17
1520
GLM 5.1
Z.ai · MIT · SiliconFlow
1.43%±0.78%
1.12%±1.74%0.99%±2.69%0.15%±1.53%3.79%±0.89%1.39%±0.17%57,532
18
1724
Meta
Muse Spark 1.1
Meta · Proprietary
0.67%±0.89%
4.39%±2.03%4.30%±2.73%4.50%±1.70%6.40%±1.64%1.36%±0.17%28,128
19
1726
Qwen3.7 Max
Alibaba · Proprietary
0.09%±1.07%
1.84%±2.60%5.73%±3.50%0.02%±2.01%7.20%±1.42%0.83%±0.29%15,992
20
1826
Gemini 3.1 Pro Preview
Google · Proprietary
0.47%±0.68%
2.05%±1.49%0.56%±2.22%1.99%±1.22%8.29%±1.11%1.32%±0.18%67,658
21
1826
Qwen3.7 Plus
Alibaba · Proprietary
0.76%±1.25%
1.74%±3.08%6.50%±3.88%1.41%±2.55%5.58%±1.85%0.30%±0.51%12,816
22
1729
Kimi K2.7 Code
Moonshot · Modified MIT
1.02%±1.69%
3.79%±3.49%0.95%±6.00%8.37%±3.29%2.86%±2.78%1.39%±0.17%10,082
23
1926
Gemini 3.5 Flash (High)
Google · Proprietary
1.03%±0.80%
2.89%±1.77%3.95%±2.48%0.68%±1.46%1.92%±1.55%1.48%±0.39%45,992
24
1827
DeepSeek V4 Pro
DeepSeek · MIT
1.19%±1.06%
4.80%±2.70%5.65%±3.42%2.11%±2.05%5.76%±1.02%0.87%±0.26%16,514
25
1829
Tencent
Hy3
Tencent · Apache 2.0
2.23%±2.87%
4.65%±6.11%2.87%±9.99%7.10%±5.95%2.56%±3.98%0.89%±0.82%3,530
26
1929
Kimi K2.6
Moonshot · Modified MIT
2.57%±1.75%
1.67%±3.51%3.04%±5.45%6.72%±3.26%6.17%±3.87%1.39%±0.17%10,139
27
2329
Minimax M3
MiniMax · MiniMax Community License
3.10%±1.05%
7.49%±2.73%9.67%±3.33%5.43%±2.14%6.15%±0.95%0.93%±0.38%16,030
28
2429
Mimo V2.5 Pro
Xiaomi · MIT
3.39%±1.11%
5.88%±2.73%10.33%±3.35%2.94%±2.14%1.69%±1.91%0.49%±0.34%16,479
29
2429
DeepSeek V4 Flash
DeepSeek · MIT
3.49%±1.06%
6.55%±2.78%9.75%±3.31%4.16%±2.06%3.46%±1.15%0.46%±0.40%16,015
30
3033
Thinking Machines
Inkling
Thinky · Apache 2.0
6.41%±1.31%
7.19%±3.50%19.01%±3.70%11.60%±3.00%6.12%±1.58%0.40%±0.49%10,678
31
3034
Gemini 3.5 Flash (Medium)
Google · Proprietary
6.80%±1.69%
13.18%±4.10%8.24%±4.98%10.20%±3.24%3.28%±3.38%0.91%±0.52%8,641
32
3034
Grok Build 0.1
SpaceXAI · Proprietary
8.01%±0.81%
4.60%±1.76%11.93%±2.39%12.26%±1.58%12.02%±1.80%0.78%±0.17%59,109
33
3034
Grok 4.3 (High)
SpaceXAI · Proprietary
8.25%±0.81%
8.72%±1.72%14.91%±2.00%7.31%±1.31%11.37%±2.42%1.08%±0.18%47,866
34
3134
Gemini 3 Flash
Google · Proprietary
8.65%±0.76%
8.74%±1.58%12.32%±1.90%5.33%±1.22%16.88%±2.00%0.03%±1.18%68,372
35
3537
Minimax M2.7
MiniMax · Modified MIT
12.47%±1.34%
17.13%±3.13%15.66%±3.70%17.46%±2.45%13.35%±3.31%1.23%±0.20%16,212
36
3538
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
13.50%±2.38%
15.08%±5.11%12.20%±7.28%21.37%±4.80%18.77%±5.33%0.09%±0.67%10,263
37
3538
Gemma 4 31B
Google · Apache 2.0
14.51%±1.60%
2.31%±1.74%4.49%±2.62%6.87%±1.53%33.53%±5.14%25.33%±5.11%54,817
38
3638
Grok 4.3
SpaceXAI · Proprietary
15.04%±1.03%
10.92%±1.61%16.20%±1.85%7.79%±1.23%41.51%±4.26%1.21%±0.18%67,800
Signal Leaders
  1. Kimi K3gets users to confirm the task is done most often14.00%±2.92%
  2. AnthropicClaude Fable 5 (High)draws the most positive responses relative to negative ones23.94%±7.42%
  3. AnthropicClaude Fable 5 (High)lands user corrections best14.62%±3.80%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.77%±0.80%
  5. Kimi K3least likely to hallucinate tools it doesn't have1.39%±0.17%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1Kimi K314.00%
    1Kimi K314.00%
  2. 2AnthropicClaude Fable 5 (High)10.67%
    2AnthropicClaude Fable 5 (High)10.67%
  3. 3AnthropicClaude Opus 4.8 (Thinking)8.90%
    3AnthropicClaude Opus 4.8 (Thinking)8.90%
  4. 4GLM 5.2 (Max)8.65%
    4GLM 5.2 (Max)8.65%
  5. 5AnthropicClaude Sonnet 5 (High)8.14%
    5AnthropicClaude Sonnet 5 (High)8.14%
  6. 6GPT 5.6 Sol (xHigh)7.25%
    6GPT 5.6 Sol (xHigh)7.25%
  7. 7AnthropicClaude Opus 4.87.10%
    7AnthropicClaude Opus 4.87.10%
  8. 8GPT 5.5 (xHigh)6.65%
    8GPT 5.5 (xHigh)6.65%
  9. 9GPT 5.4 (High)6.23%
    9GPT 5.4 (High)6.23%
  10. 10GPT 5.5 (High)6.20%
    10GPT 5.5 (High)6.20%
679,671 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1AnthropicClaude Fable 5 (High)23.94%
    1AnthropicClaude Fable 5 (High)23.94%
  2. 2GPT 5.6 Sol (xHigh)23.53%
    2GPT 5.6 Sol (xHigh)23.53%
  3. 3Kimi K320.30%
    3Kimi K320.30%
  4. 4AnthropicClaude Opus 4.8 (Thinking)19.42%
    4AnthropicClaude Opus 4.8 (Thinking)19.42%
  5. 5AnthropicClaude Sonnet 5 (High)16.88%
    5AnthropicClaude Sonnet 5 (High)16.88%
  6. 6GLM 5.2 (Max)12.94%
    6GLM 5.2 (Max)12.94%
  7. 7AnthropicClaude Opus 4.712.48%
    7AnthropicClaude Opus 4.712.48%
  8. 8AnthropicClaude Opus 4.811.63%
    8AnthropicClaude Opus 4.811.63%
  9. 9AnthropicClaude Opus 4.7 (Thinking)11.55%
    9AnthropicClaude Opus 4.7 (Thinking)11.55%
  10. 10GPT 5.5 (xHigh)11.08%
    10GPT 5.5 (xHigh)11.08%
259,622 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Fable 5 (High)14.62%
    1AnthropicClaude Fable 5 (High)14.62%
  2. 2AnthropicClaude Opus 4.8 (Thinking)9.78%
    2AnthropicClaude Opus 4.8 (Thinking)9.78%
  3. 3GPT 5.6 Sol (xHigh)9.71%
    3GPT 5.6 Sol (xHigh)9.71%
  4. 4AnthropicClaude Opus 4.78.95%
    4AnthropicClaude Opus 4.78.95%
  5. 5GPT 5.5 (High)8.77%
    5GPT 5.5 (High)8.77%
  6. 6AnthropicClaude Opus 4.7 (Thinking)8.62%
    6AnthropicClaude Opus 4.7 (Thinking)8.62%
  7. 7AnthropicClaude Opus 4.88.25%
    7AnthropicClaude Opus 4.88.25%
  8. 8GPT 5.5 (xHigh)8.18%
    8GPT 5.5 (xHigh)8.18%
  9. 9GPT 5.4 (High)7.75%
    9GPT 5.4 (High)7.75%
  10. 10AnthropicClaude Opus 4.66.53%
    10AnthropicClaude Opus 4.66.53%
436,382 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1GPT 5.5 (xHigh)14.77%
    1GPT 5.5 (xHigh)14.77%
  2. 2AnthropicClaude Fable 5 (High)12.97%
    2AnthropicClaude Fable 5 (High)12.97%
  3. 3AnthropicClaude Opus 4.7 (Thinking)12.57%
    3AnthropicClaude Opus 4.7 (Thinking)12.57%
  4. 4GPT 5.5 (High)11.90%
    4GPT 5.5 (High)11.90%
  5. 5AnthropicClaude Sonnet 4.611.45%
    5AnthropicClaude Sonnet 4.611.45%
  6. 6GPT 5.511.22%
    6GPT 5.511.22%
  7. 7AnthropicClaude Opus 4.611.14%
    7AnthropicClaude Opus 4.611.14%
  8. 8AnthropicClaude Sonnet 5 (High)10.81%
    8AnthropicClaude Sonnet 5 (High)10.81%
  9. 9AnthropicClaude Opus 4.710.62%
    9AnthropicClaude Opus 4.710.62%
  10. 10Grok 4.510.56%
    10Grok 4.510.56%
415,584 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1Kimi K31.39%
    1Kimi K31.39%
  2. 2GPT 5.5 (High)1.39%
    2GPT 5.5 (High)1.39%
  3. 3GPT 5.51.39%
    3GPT 5.51.39%
  4. 4GPT 5.5 (xHigh)1.39%
    4GPT 5.5 (xHigh)1.39%
  5. 5AnthropicClaude Fable 5 (High)1.39%
    5AnthropicClaude Fable 5 (High)1.39%
  6. 6Kimi K2.61.39%
    6Kimi K2.61.39%
  7. 7Grok 4.51.39%
    7Grok 4.51.39%
  8. 8GPT 5.4 (High)1.39%
    8GPT 5.4 (High)1.39%
  9. 9GLM 5.2 (Max)1.39%
    9GLM 5.2 (Max)1.39%
  10. 10Kimi K2.7 Code1.39%
    10Kimi K2.7 Code1.39%
1,512,983 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© High Intelligence 2026