Nine major AI models, four benchmarks, and a ranking that reverses on every tab. The “smartest model” is not a stable fact — it is a function of which test you run. This chart maps the three variables that actually matter: cost, speed, and the benchmark that resembles your use case.
LLM Comparison
Comparing large language models is usually done by two numbers: a benchmark score and a price per million tokens. Both invite a precision the underlying question doesn’t support. Benchmark scores measure performance on specific standardised tasks under controlled conditions; not performance on the actual work at hand. Price per million tokens ranks the rate, not the bill: what you pay depends on context length, the ratio of input to output tokens, and how many tokens your task genuinely consumes. These aren’t footnotes. For teams spending real money on model selection, they determine whether the cheaper-looking option is cheaper at all. This cluster examines the LLM market where the popular metrics get the mechanism wrong.
The token-price illusion: what “cost per million tokens” actually measures
Headline price per million tokens ranks the rate, not the bill. Disaggregate the two and the table reorders.
