Describe what you need.
Get the model that fits.
Type it in plain English. We rank the models on real test results.
How this works, and why we think it's good
Plain answers, with real examples.
What you are looking at
Every number on the Leaderboards page is a real published measurement: benchmark scores from about 60 tests run by 19 independent organisations, each read from that organisation's own data, plus cost and speed from Artificial Analysis and from the benchmarks that publish their own. Nothing there is generated by an AI.
How the Ask tab works
Type a sentence describing what you need. A small, fast AI model called Jev (from TypeSafe) reads it and makes three judgements: what kind of model you're after, which of our measurements actually matter for that need, and which models simply can't do the job. Ordinary code — no AI — then does the ranking. Jev never invents a score.
Why this beats a single leaderboard, with examples
We look at coding benchmarks and price, and nothing else — a model's maths or writing scores don't move the ranking. A model that tops the general leaderboard but is expensive or mediocre at coding won't come out on top here.
We switch entirely to speech-to-text benchmarks, and drop any model that can't tell speakers apart — even if it's an excellent transcriber overall.
Anything closed-source, or too large to run locally, is ruled out before ranking even starts — regardless of how well it scores.
Things to keep in mind
- Amber or hollow means estimated. If a model was never tested on the measurement shown, we estimate it from similar models; if the source itself published only an estimate, we mark that too. Cost is never estimated: a model with no measured cost is left off cost charts until it has one.
- We don't trust a maker's word for its own model. We've seen a vendor's self-reported score sit well above what independent testing found for the same model, so a maker's own published figures for its own model are never used — only independent, third-party results.
- Data quality. Every benchmark number is read from the organisation that measured it, never from a site that republishes it, and on every update each source is compared row by row with what we show.
- Cost means something different on each benchmark (per task, per test, a whole run), so each benchmark that publishes a cost has its own honest cost axis rather than being squeezed into one number.
- Image, video and speech models have far fewer measurements than language models: usually one quality rating, a price, and a list of features. New releases show as "not yet rated" until tested.
Sources: Artificial Analysis, arena.ai, OpenRouter, and the leaderboards named on each chart. Research gathered with Exa.