Model comparisons

Compare the capability signal.

Public benchmark data helps you understand capability. Your private task telemetry tells you which model actually works best for you.

ModelRelative scoreScore

GPT-5.6 Sol

OpenAI

96
2

Claude Opus 4.8

Anthropic

94
3

Grok 4.5

xAI

89
4

Gemini 3.6 Flash

Google

84

This is a clearly labeled demo composite used to exercise the multi-source benchmark architecture. Production benchmark feeds are the next integration step; no score is presented as an official provider claim.

Best for coding

GPT-5.6 Sol

96 / 100

Best instruction following

Claude Opus 4.8

97 / 100

Fastest

Gemini 3.6 Flash

96 / 100