Skip to main content
TokenCost logoTokenCost

LLM Leaderboard

Live rankings by quality, speed, and value. Data from Artificial Analysis benchmarks.

Composite intelligence score from Artificial Analysis benchmarks

#2
Gemini 3.1 Pro
Google
48
idx
#1
GPT-5.4
OpenAI
53
idx
#3
GPT-5.3 Codex
OpenAI
46
idx

Quality vs Price

Higher and left = better value. Hover for details.

$0.1$0.5$1$2$5$101020304050Input Price / 1M tokens (log)Quality Index
OpenAI
Anthropic
Google
xAI
Meta
Mistral
Amazon
NVIDIA
Cohere
Perplexity
Moonshot
Zhipu
MiniMax
#ModelQuality Index
1GPT-5.4OpenAI
53idx
2Gemini 3.1 ProGoogle
48idx
3GPT-5.3 CodexOpenAI
46idx
4GPT-5.2OpenAI
43idx
5GLM-5Zhipu
41idx
6Claude Opus 4.6Anthropic
39idx
7Grok 4.20xAI
38idx
8GPT-5.1OpenAI
38idx
9Claude Sonnet 4.6Anthropic
37idx
10Kimi K2.5Moonshot
36idx
11Claude Opus 4.5Anthropic
36idx
12GPT-5OpenAI
35idx
13MiniMax M2.5MiniMax
35idx
14o3OpenAI
31idx
15Gemini 3 FlashGoogle
28idx
16o4 MiniOpenAI
26idx
17Gemini 2.5 ProGoogle
26idx
18GPT-5 MiniOpenAI
26idx
19Nemotron 3 Super 120BNVIDIA
26idx
20Gemini 3.1 Flash-LiteGoogle
26idx
21Claude Haiku 4.5Anthropic
24idx
22o1OpenAI
24idx
23Nova 2.0 Pro ReasoningAmazon
22idx
24GPT-5 NanoOpenAI
20idx
25GPT-4.1OpenAI
20idx
26o3 MiniOpenAI
19idx
27Mistral Large 3Mistral
16idx
28GPT-4.1 MiniOpenAI
15idx
29Llama 4 MaverickMeta
15idx
30Gemini 2.5 FlashGoogle
14idx
31Claude Haiku 3.5Anthropic
12idx
32Nova 2.0 LiteAmazon
12idx
33GPT-4oOpenAI
11idx
34Llama 4 ScoutMeta
10idx
35GPT-4.1 NanoOpenAI
10idx
36Sonar ProPerplexity
9idx
37Command ACohere
8idx
38GPT-4o MiniOpenAI
7idx
39Gemini 2.5 Flash-LiteGoogle
7idx

Rankings based on live benchmark data. Quality = composite intelligence index. Value = quality index / input cost per 1M tokens. Latency = time to first token.

Data by Artificial Analysis

How to Use the LLM Leaderboard

  1. 1

    Choose a ranking metric

    Switch between Quality, Speed, and Value tabs to rank models by the metric that matters most to your use case.

  2. 2

    Filter by provider

    Use the provider buttons to focus on specific vendors. Compare only OpenAI models, or pit Anthropic against Google.

  3. 3

    Explore the scatter chart

    The interactive quality-vs-price chart plots every model so you can visually identify the best value picks.

Why Use This Leaderboard

  • Three ranking modes (Quality, Speed, and Value) for different decision criteria
  • Benchmark data from Artificial Analysis, refreshed every 6 hours
  • Interactive scatter chart plotting quality against cost for visual comparison
  • Provider filtering to narrow the field to vendors you're evaluating
  • Includes output speed (tokens/sec) and time-to-first-token for latency planning

Common Use Cases

Model selection

Find the highest-quality model within your budget by sorting on the Value tab.

Latency optimization

Sort by Speed to find the fastest models for real-time applications like chat or autocomplete.

Benchmark tracking

Check back regularly to see how new model releases stack up against existing options.

Stakeholder reporting

Use the scatter chart to show leadership why a specific model offers the best quality-to-cost ratio.

Related Tools

Frequently Asked Questions

Common questions about the LLM leaderboard