Skip to main content
TokenCost logoTokenCost

LLM Leaderboard

Live rankings by quality, speed, and value. Data from Artificial Analysis benchmarks.

Composite intelligence score from Artificial Analysis benchmarks

#2
Claude Sonnet 4.6 Adaptive
Anthropic
47
idx
#1
GPT-5.4
OpenAI
51
idx
#3
Gemini 3.1 Pro
Google
47
idx

Quality vs Price

Higher and left = better value. Hover for details.

$0.1$0.5$1$2$5$101020304050Input Price / 1M tokens (log)Quality Index
OpenAI
Anthropic
Google
xAI
Meta
Mistral
DeepSeek
Amazon
NVIDIA
Cohere
Perplexity
Moonshot
Zhipu
MiniMax
#ModelQuality Index
1GPT-5.4OpenAI
51idx
2Claude Sonnet 4.6 AdaptiveAnthropic
47idx
3Gemini 3.1 ProGoogle
47idx
4GPT-5.3 CodexOpenAI
44idx
5Claude Opus 4.6 AdaptiveAnthropic
44idx
6GPT-5.2OpenAI
42idx
7Gemini 3 ProGoogle
40idx
8GLM-5Zhipu
40idx
9Claude Opus 4.6Anthropic
38idx
10Gemini 3 Flash ReasoningGoogle
38idx
11Grok 4.20xAI
37idx
12GPT-5.1OpenAI
37idx
13Claude Sonnet 4.6Anthropic
36idx
14Kimi K2.5Moonshot
35idx
15GPT-5OpenAI
35idx
16Claude Opus 4.5Anthropic
35idx
17Claude Sonnet 4Anthropic
34idx
18GPT-5 MediumOpenAI
34idx
19MiniMax M2.5MiniMax
34idx
20Grok 4xAI
33idx
21Grok 4.1 Fast ReasoningxAI
31idx
22o3OpenAI
30idx
23Claude 4.5 Haiku ReasoningAnthropic
30idx
24Gemini 3 FlashGoogle
27idx
25Gemini 2.5 ProGoogle
26idx
26o4 MiniOpenAI
26idx
27Nemotron 3 Super 120BNVIDIA
25idx
28GPT-5 MiniOpenAI
25idx
29Gemini 3.1 Flash-LiteGoogle
25idx
30DeepSeek V3.2 (Chat)DeepSeek
25idx
31Claude Haiku 4.5Anthropic
24idx
32o1OpenAI
23idx
33Nova 2.0 Pro ReasoningAmazon
22idx
34DeepSeek R1DeepSeek
20idx
35GPT-5 NanoOpenAI
20idx
36GPT-4.1OpenAI
19idx
37o3 MiniOpenAI
19idx
38Grok 4.1 FastxAI
17idx
39Mistral Large 3Mistral
16idx
40GPT-4.1 MiniOpenAI
15idx
41Llama 4 MaverickMeta
14idx
42Gemini 2.5 FlashGoogle
14idx
43Claude Haiku 3.5Anthropic
12idx
44Gemini 2.0 FlashGoogle
12idx
45Nova 2.0 LiteAmazon
12idx
46GPT-4oOpenAI
11idx
47Llama 4 ScoutMeta
10idx
48GPT-4.1 NanoOpenAI
10idx
49Sonar ProPerplexity
9idx
50Gemini 2.0 Flash-LiteGoogle
9idx
51Command ACohere
8idx
52GPT-4o MiniOpenAI
7idx
53Gemini 2.5 Flash-LiteGoogle
7idx
54Mistral Small 3.2Mistral
5idx

Rankings based on live benchmark data. Quality = composite intelligence index. Value = quality index / input cost per 1M tokens. Latency = time to first token.

Data by Artificial Analysis

How to Use the LLM Leaderboard

  1. 1

    Choose a ranking metric

    Switch between Quality, Speed, and Value tabs to rank models by the metric that matters most to your use case.

  2. 2

    Filter by provider

    Use the provider buttons to focus on specific vendors. Compare only OpenAI models, or pit Anthropic against Google.

  3. 3

    Explore the scatter chart

    The interactive quality-vs-price chart plots every model so you can visually identify the best value picks.

Why Use This Leaderboard

  • Three ranking modes (Quality, Speed, and Value) for different decision criteria
  • Benchmark data from Artificial Analysis, refreshed every 6 hours
  • Interactive scatter chart plotting quality against cost for visual comparison
  • Provider filtering to narrow the field to vendors you're evaluating
  • Includes output speed (tokens/sec) and time-to-first-token for latency planning

Common Use Cases

Model selection

Find the highest-quality model within your budget by sorting on the Value tab.

Latency optimization

Sort by Speed to find the fastest models for real-time applications like chat or autocomplete.

Benchmark tracking

Check back regularly to see how new model releases stack up against existing options.

Stakeholder reporting

Use the scatter chart to show leadership why a specific model offers the best quality-to-cost ratio.

Related Tools

Frequently Asked Questions

Common questions about the LLM leaderboard