AI Model Benchmarks
Intelligence, speed, price, context, verbosity and trustworthiness β sortable. Click a column header to sort; the deeper a cell's shade, the better that model ranks in that column.
My Score weights:
Intelligence
Price
Speed
Context
Verbosity
Reset to equal
Modeli
Data agei
Intelligencei
Speed (tok/s)i
Latency (s)i
Price ($/1M)i
Contexti
Intelligence per $i
Tok/taski
Omnisciencei
Codingi
T-Bench Hardi
LiveBenchi
Agentic codingi
Cost/successi
Reasoningi
Mathematicsi
Code geni
Data analysisi
Languagei
Instruction foll.i
My Scorei
Intelligence, speed and price data by Artificial Analysis .
Context windows from OpenRouter ; "β" means the model isn't listed there.
Price is Artificial Analysis' blended rate (3:1 input:output); hover a price for the input/output split.
Latency = median seconds until the answer starts appearing, thinking time included; hover a
value to see how much of it was thinking.
Intelligence per $ = Intelligence Index Γ· blended price.
Tok/task = output tokens the model spends per Intelligence Index task (verbosity; lower
means denser answers), read from Artificial Analysis' models page β available only for the
~28 models featured there, as are Omniscience and hallucination rate.
Omniscience is AA's knowledge-and-trustworthiness index: it rewards correct answers and
admitting ignorance, and penalises confident wrong answers, so it can go negative; the
"% hall." figure is how often the model answered wrongly instead of saying it doesn't know.
Coding view ranks models by AA's Coding Index and shows Terminal-Bench Hard
(agentic software-engineering tasks; AA doesn't publish SWE-bench).
LiveBench columns come from LiveBench (release 2026-06-25,
source ), a contamination-free benchmark:
fresh questions each month with objective ground-truth answers, so no model has seen them
before and no AI judges the results. Its overall score is the mean of its seven category
scores; the categories are available as columns in the Columns menu. Cost/success is
LiveBench's cost per successful task β measured token usage at real prices, divided by
the tasks the model got right, so failed attempts are paid for too. LiveBench has run
roughly 33 of the models here; the rest show "β". A model is only matched when the
reasoning effort matches on both sides, so LiveBench-at-xhigh is never shown against a
row that Artificial Analysis measured at max.
The modality and open/closed filters use AA's models page where available and OpenRouter's
listing otherwise (a Hugging Face link there = downloadable weights); models whose
status is unknown are hidden while a filter is active.
My Score is a weighted geometric mean of each model's percentile ranks on the four
metrics (price inverted, so cheaper ranks higher), using the slider weights; a metric a model
lacks (e.g. unknown price or verbosity) counts as the 50th percentile, so missing data
neither helps nor hurts.
Weights reset to their defaults when the page reloads.
The comparison chart at the bottom holds up to five models at once. Its "overall
profile" spokes are percentile ranks across every model here, because raw dollars,
tokens per second and context tokens can't share one axis; its LiveBench spokes are
the real category scores, which already do. Spokes that are naturally lower-is-better
(price, latency, verbosity) are reversed and renamed, so on every spoke further out
is better. Selections aren't saved between visits.
Data age answers a question the sources don't: none of them publishes a
"last measured" date, so the only way to tell whether a refresh brought new figures or
the same ones again is to remember the previous values and compare. Every refresh does
that, and records the day each number moved β which matters now that labs update a model
behind the same name rather than shipping a new one, so a score can change with nothing
announcing it. The column shows the benchmark scores; hover a value to see price and
speed separately, and what changed. Speed and latency are rolling medians that drift on
their own, so they move far more often than a score does and are deliberately kept out of
the headline figure. A "β₯" means the numbers haven't moved since that model first appeared
here, and were most likely settled before that β it's a floor, not an age. This page has
only been keeping the log since 2026-07-31, so those floors will sharpen into real dates as
models are re-evaluated.