Model comparisons

Compare skills, cost, and fit.

Compare coding, reasoning, math, instruction-following, speed, and context. Then check price and value before choosing what to test.

Show company

Capability signals

Compare skills and behavior

Compare by

Building, editing, and debugging software

RankModelRelative signalScore
1

Claude Opus 5

Anthropic

97
2

GPT-5.6 Sol

OpenAI

96
3

Claude Opus 4.8

Anthropic

94
4

Claude Sonnet 5

Anthropic

94
5

Grok Build 0.1

xAI

92
6

GPT-5.6 Terra

OpenAI

91
7

Claude Sonnet 4.6

Anthropic

91
8

Gemini 2.5 Pro

Google

91
9

Grok 4.5

xAI

89
10

Grok 4.3

xAI

86
11

GPT-5.6 Luna

OpenAI

84
12

Claude Haiku 4.5

Anthropic

84
13

Gemini 3.6 Flash

Google

84
14

Gemini 3.1 Flash-Lite

Google

78

These are directional AI Stack composite signals on a normalized 0–100 scale, not official provider benchmark scores. Use them to narrow the field, then test finalists on your own work.

Price and capacity

Compare cost, value, and working memory

Rank by

ModelGood forInput / 1MOutput / 1MTypical costOutput / $1Memory
1

Gemini 3.1 Flash-Lite

Google

Low cost · High volume

$0.25$1.50$0.56667K1M
2

Grok Build 0.1

xAI

Coding · Implementation

$1.00$2.00$1.25500K256K
3

Grok 4.3

xAI

Value · Large context

$1.25$2.50$1.56400K1M
4

Claude Haiku 4.5

Anthropic

Speed · High volume

$1.00$5.00$2.00200K200K
5

GPT-5.6 Luna

OpenAI

High volume · Fast iteration

$1.00$6.00$2.25167K1.05M
6

Gemini 3.6 Flash

Google

Speed · Multimodal

$1.50$7.50$3.00133K1M
7

Grok 4.5

xAI

Reasoning · Current information

$2.00$6.00$3.00167K500K
8

Gemini 2.5 Pro

Google

Complex reasoning · Coding

$1.25$10.00$3.44100K1.05M
9

Claude Sonnet 5

Anthropic

Coding · Data analysis

$2.00$10.00$4.00100K1M
10

GPT-5.6 Terra

OpenAI

Balanced value · Coding

$2.50$15.00$5.6367K1.05M
11

Claude Sonnet 4.6

Anthropic

Debugging · Implementation

$3.00$15.00$6.0067K1M
12

Claude Opus 4.8

Anthropic

Long tasks · Coding

$5.00$25.00$10.0040K1M
13

Claude Opus 5

Anthropic

Agentic coding · Long tasks

$5.00$25.00$10.0040K1M
14

GPT-5.6 Sol

OpenAI

Complex coding · Agents

$5.00$30.00$11.2533K1.05M
15

Claude Fable 5

Anthropic

Deep reasoning · Long-horizon agents

$10.00$50.00$20.0020K1M

Input is what you send; output is what the model writes back. Typical cost estimates one million total tokens as 75% input and 25% output. API prices can vary with caching, batching, long prompts, tools, and service tiers.

Lowest typical cost

Gemini 3.1 Flash-Lite

$0.56 / 1M total

Most output per dollar

Gemini 3.1 Flash-Lite

667K tokens / $1

Largest working memory

GPT-5.6 Sol

1.05M