Which AI model, for what job?
Alyriaz's living ranking: what each major model is worth, what it costs, and which one to choose for what you want to do. One edition a month, and new models join it every week.
What do you want to do?
There is no best model in general: there is the right model for a task. Pick yours.
Loading the radar…
The price map
Each dot is a model, placed by its price per million tokens read (the scale is logarithmic: each step costs ten times more than the previous one). Hover over or tap a dot.
How much would it cost you?
Advanced mode: set every parameter yourself
All models
Prices in dollars per million tokens, read (input) and written (output). Click a column heading to sort.
| Model | Price in / out ($ per million tokens) | Best for |
|---|---|---|
| Claude Opus 5.5 (Anthropic) | $4 / $20 | Fable 5.1 level, 60% cheaper. Agentic coding: 66.4% on Terminal-Bench 4.0 according to Anthropic, ahead of Astra. Cache reads at $0.20. |
| GPT-6 Astra (OpenAI) | $10 / $50 | OpenAI's top model. Research mathematics (97.6% claimed on FrontierMath tier 4) and science in the terminal. First model rated "critical" in cybersecurity. |
| Claude Fable 5.1 (Anthropic) | $10 / $50 | Anthropic's top model. Top of the independent general index at our last check, but the most expensive per task. |
| Claude Mythos 5.1 (Anthropic) | $10 / $50 | Restricted access: cybersecurity and life sciences. Restricted access because of its capabilities in cybersecurity and life sciences. Same price as Fable 5.1. |
| Gemini 4 Argon (Google) | Restricted access: reserved for cybersecurity defenders. Google's most powerful model, which it claims beats Anthropic and OpenAI on several tests. Answers of up to one million tokens; distributed first through the Fairwind programme. | |
| Claude Sonnet 5.5 (Anthropic) | $2 / $10 | Almost Opus 5.5, at half the price. Across 44 professions tested by Anthropic, it almost matches Opus 5.5. The everyday model. |
| GPT-6.1 Sol (OpenAI) | $2 / $10 | Almost Astra for code, at a fifth of the price. Almost matches Astra at writing code and using a computer, according to OpenAI. Rated "critical" in cybersecurity, like Astra. |
| Gemini 3.6 Flash (Google) | $1.5 / $7.5 | Google's everyday model. Better and cheaper than the 3.5 Flash it replaces: 17% fewer output tokens on the Artificial Analysis index, and output drops from $9 to $7.50. |
| Claude Haiku 5.5 (Anthropic) | $0.10 / $0.50 | Anthropic's small model, at the price of GPT-6 Luna. Ten times cheaper than Haiku 4.5 below 100,000 tokens per request ($0.50 / $2.50 above). Adjustable effort level; ahead of Luna on the tests Anthropic selected. |
| GPT-6 Luna (OpenAI) | $0.10 / $0.50 | Volume and simple tasks. At $0.10 per million tokens, it changes the economics of sorting, extracting and filtering large amounts of data. |
| Kimi K3 (Moonshot AI) | $3 / $15 | The largest open model ever released. 2.8 trillion parameters and one million tokens of context. At release, just behind the best closed models on several independent rankings; it always thinks at maximum effort, so cheaper per token, not necessarily per task. |
| Qwen 3.8-Max (Alibaba (Qwen)) | 2.4 trillion parameters, open weights. According to Alibaba, it improves at programming, office work and document research. In practice, only for organisations able to host it. | |
| DeepSeek V4 Flash (DeepSeek) | A Chinese open model for a few cents. Its price went from about $0.14 to $0.27 per million tokens on August 14. Still very low, but an introductory price can double overnight: you can also host it yourself. | |
| Mistral Large 4 (Mistral AI) | The European giant, weights promised by the end of October. About 1 trillion parameters with 49 billion active, trained and hosted in Europe. 38 points on the Artificial Analysis index, against 9 for Large 3 and 58 for Claude Opus 5.5. | |
| Step 5 Preview (StepFun) | $1 / $2.7 | 600 billion parameters at $1 per million. Mixture of experts (27 billion active parameters), one million tokens of context, open access from day one. |
| LongCat-2.0 (Meituan) | $0.75 / $2.95 | The mystery model "Owl Alpha", made by Meituan. A specialist in agentic coding, 1.6 trillion parameters with about 48 billion active. Cache reads cost nothing, an asset for agents running in loops; weights promised under the MIT licence. |
| Gemini 3.5 Flash-Lite (Google) | $0.30 / $2.5 | Google's ultra-cheap model. 350 output tokens per second and adjustable thinking, for translation, document analysis and high volumes. |
| GLM-5.3-Flash (Z.ai) | Open, MIT licence, served at scale. 320 billion parameters with 18 billion active, one million tokens of context. Runs on more than 100,000 Chinese accelerators. | |
| Qwen 3.8 27B (Alibaba (Qwen)) | The open model that fits on a single graphics card. 27 billion parameters under the Apache 2.0 licence, built-in vision, 262,000 tokens of context. Enough to analyse confidential documents without them leaving your machines. | |
| K2 Horizon (Institute of Foundation Models) | A full range of six open models. From 0.9 to 375 billion parameters, all under the Apache 2.0 licence: pick the size that fits your hardware. | |
| Qwen3.8-Omni-Flash (Alibaba (Qwen)) | One model for text, images, sound and video. Natively handles four modalities, where you used to need several models chained together. | |
| GPT-6 Sol (OpenAI), replaced | $2 / $10 | Replaced a week later by GPT-6.1 Sol. Best score-to-cost ratio on automation: 33.2% for 27 cents per task. |
| Claude Opus 5 (Anthropic), replaced | $5 / $25 | Replaced by Opus 5.5. Opus 5.5 claims a 40% lower real cost than this model. |
| Grok 4.6 (xAI), replaced | $2 / $6 | Replaced by Grok 4.7 in September. At release, level with GPT-5.6 Sol on the Artificial Analysis index, with 500,000 tokens of context. Watch the tiers: the price per token changes beyond a context threshold. |
The latest releases
Before you trust a ranking
Scores have become hard to compare. Five reasons to keep a cool head, including about this one.
The effort level
The same model can gain or lose dozens of points depending on how much thinking time it is given, and cost five to ten times more. Learn more
The harness
Rankings compare a model together with the software around it, never the model alone. Learn more
Worn-out tests
An old test ends up mostly measuring that the model has already seen it during training. Learn more
The source
A table published by a vendor picks its own tests. We always separate what it claims from what third parties measure.
Newer does not mean better
A new model can do worse than the one it replaces on your specific task. Learn more