Anthropic’s Claude Opus 5.5 tops all three major AI leaderboards this month, but the best model for you depends on how much you value speed, price and context length.

Why the leaderboards disagree
Search for the “best AI model” and you will find several rankings that never quite match. That is not an error. Each site measures something different, and the differences matter when you choose a model.
Three sources shaped this article, each as a snapshot taken on 11 October 2026:
- Artificial Analysis publishes an Intelligence Index, a single 0–100 score built from a bundle of tests. It also tracks real API behaviour: output speed in tokens per second, time to the first token, total response time and the average cost of finishing a benchmark task.
- LLM Stats produces its own composite score and splits it into reasoning, coding and agent sub-scores. It also runs head-to-head voting arenas, so human preference feeds into its picture.
- BestAIModels repackages Artificial Analysis quality data into a shortlist of 30 models, then adds its own blended price per million tokens, speed and context figures.
Because scores come from different test mixes, the same model can land on 58, 60.5 or 57.6 depending on the page. Treat the numbers as a way to compare models inside one table, not as absolute truths across tables. The order at the top, though, is remarkably stable, and that is the real story.
The intelligence race: who sits at the top
On the Artificial Analysis Intelligence Index, the top spot belongs to Claude Opus 5.5 running at its maximum reasoning setting, with a score of 58. Claude Sonnet 5.5 at max effort follows closely on 56, level with Opus 5.5 at the “xhigh” setting. LLM Stats agrees on the order of the leaders: Opus 5.5 scores 60.5 there, Sonnet 5.5 58.3 and OpenAI’s GPT-6 Astra 58.2, so the top three are separated by only a couple of points.
The table below lists 160 models from the Artificial Analysis index. The first ten rows are the strongest entries; the 150 rows after them add other models, one row each at its best-scoring setting, from highest to lowest score. A dash means the source reported no value, and “~0” means a cost that rounds to zero. Cost per task is the average spend to complete the index’s tasks, so it reflects how many tokens a model burns while thinking, not just its list price.
| Model (setting) | Maker | Index score | Cost per task (USD) | Output speed (tokens/s) |
|---|---|---|---|---|
| Claude Opus 5.5 (max) | Anthropic | 58 | 5.98 | 96 |
| Claude Sonnet 5.5 (max) | Anthropic | 56 | 5.46 | 141 |
| Claude Opus 5.5 (xhigh) | Anthropic | 56 | 3.46 | 81 |
| Claude Opus 5.5 (high) | Anthropic | 54 | 1.82 | 77 |
| Claude Fable 5.1 (max) | Anthropic | 53 | 7.63 | 70 |
| GPT-6 Astra (max) | OpenAI | 53 | 3.26 | 47 |
| Gemini 4 Argon (high) | 53 | 1.99 | not reported | |
| GPT-6.1 Sol (max) | OpenAI | 52 | 0.72 | 56 |
| Muse Spark 1.3 (max) | Meta | 48 | 1.60 | 175 |
| Grok 4.7 (xhigh) | SpaceXAI | 46 | 3.74 | 74 |
| MiMo-V2.6-Pro | Xiaomi | 46 | 0.13 | 43 |
| Qwen3.8 Max (0902) | Alibaba | 45 | 5.41 | 36 |
| GLM-5.3 (max) | Z AI | 45 | 2.01 | 83 |
| Step 5 Preview | StepFun | 44 | 1.03 | 87 |
| Kimi K3 (max) | Kimi | 44 | 2.00 | 41 |
| Claude Haiku 5.5 (max) | Anthropic | 43 | 0.21 | 240 |
| GPT-5.6 Terra (max) | OpenAI | 42 | 1.40 | 108 |
| GLM-5.3-Flash | Z AI | 42 | 0.25 | 53 |
| Ling 3.1 Flash | InclusionAI | 41 | 0.99 | 214 |
| Gemini 3.8 Flash (high) | 41 | 1.24 | 125 | |
| Qwen3.8 2.4T A95B | Alibaba | 40 | 2.16 | 37 |
| Qwen3.8-Flash-Next | Alibaba | 40 | 0.37 | 56 |
| DeepSeek V4.1 Flash (max) | DeepSeek | 39 | 0.27 | 217 |
| Mistral Large 4 Preview | Mistral | 38 | 1.13 | — |
| GPT-6 Luna (max) | OpenAI | 38 | 0.07 | 139 |
| MiMo-V2.6-Flash | Xiaomi | 38 | 0.06 | 58 |
| DeepSeek V4 Pro 0813 (max) | DeepSeek | 36 | 0.67 | 89 |
| DeepSeek V4 Flash Vision (max) | DeepSeek | 35 | 0.31 | 221 |
| JT-4.1 Flash 236B A21B | China Mobile | 34 | — | — |
| Qwen3.8 27B (xhigh) | Alibaba | 34 | 1.01 | 46 |
| Motif 3 | Motif Technologies | 34 | — | — |
| GPT-5.3 Codex (xhigh) | OpenAI | 33 | — | 91 |
| K2 Horizon 375B A23B | Institute of Foundation Models | 31 | ~0 | 119 |
| Gemini 3.1 Pro Preview | 30 | 1.30 | 114 | |
| MiniMax-M3 | MiniMax | 29 | 0.51 | 92 |
| Nex-N2-Pro | Nex AGI | 28 | — | — |
| Solar Pro 4 | Upstage | 28 | — | 104 |
| Quasar 438B (max) | Multiverse Computing | 27 | 2.02 | 126 |
| Apodex 1.1 | Apodex | 26 | 0.46 | — |
| GPT-5.5 Instant (June 2026) | OpenAI | 26 | 0.69 | 139 |
| Kimi K2.7 Code | Kimi | 26 | 0.54 | 79 |
| Inkling Small | Thinking Machines | 26 | 0.09 | 160 |
| Hy3 | Tencent | 25 | 0.07 | 83 |
| Qwen3.7 Plus | Alibaba | 25 | 0.22 | 54 |
| Inkling (xhigh) | Thinking Machines | 25 | — | 151 |
| Ling-3.0-flash-VL | InclusionAI | 25 | — | 145 |
| Solar Mini 4 | Upstage | 24 | 0.37 | 74 |
| Nemotron 3 Ultra | NVIDIA | 23 | 0.60 | 142 |
| Ling-3.0-flash-Fin | InclusionAI | 23 | — | 321 |
| Gemini 3.5 Flash-Lite | 22 | 0.19 | 365 | |
| KAT-Coder-Pro V2 | KwaiKAT | 22 | — | — |
| A.X-K2 | SK Telecom | 21 | — | — |
| o3 | OpenAI | 20 | — | 116 |
| Qwen3.5 Omni Plus | Alibaba | 20 | — | 89 |
| K-EXAONE 2.0 | LG AI Research | 20 | — | — |
| LongCat 2.0 | LongCat | 19 | 0.15 | — |
| JT-35B-Flash | China Mobile | 19 | — | — |
| Qwen3.6 35B A3B | Alibaba | 18 | 0.48 | 127 |
| Qwen3.5 397B A17B | Alibaba | 18 | 0.47 | 87 |
| Muse Glimmer (high) | Meta | 17 | 0.06 | 112 |
| Ring-2.6-1T | InclusionAI | 17 | 0.29 | 116 |
| Doubao Seed Code | ByteDance Seed | 17 | — | — |
| Gemma 4 26B A4B | 17 | — | — | |
| Qwen3.5 122B A10B | Alibaba | 16 | 0.32 | 130 |
| Gemma 4 31B | 15 | ~0 | 35 | |
| Mistral Medium 3.5 | Mistral | 14 | 0.50 | 163 |
| ERNIE 5.0 Thinking Preview | Baidu | 14 | — | — |
| Nova 2.0 Pro Preview (medium) | Amazon | 14 | — | 120 |
| Gemma 4 12B | 14 | — | 114 | |
| Nova 2.0 Omni (medium) | Amazon | 14 | — | — |
| Apriel-v1.6-15B-Thinker | ServiceNow | 13 | — | — |
| Nova 2.0 Lite (high) | Amazon | 13 | — | 188 |
| Command A+ | Cohere | 13 | ~0 | 175 |
| Nemotron 3.5 Lightning | NVIDIA | 13 | 0.09 | 307 |
| Nemotron 3 Super | NVIDIA | 13 | 1.64 | 152 |
| Granite 4.2 30B | IBM | 13 | 0.07 | 75 |
| EXAONE 4.5 33B | LG AI Research | 13 | — | — |
| Qwen3.5 4B | Alibaba | 13 | — | 23 |
| Qwen3.5 Omni Flash | Alibaba | 12 | — | 232 |
| MiniCPM5-2B | OpenBMB | 12 | — | — |
| Magistral Medium 1.2 | Mistral | 12 | — | — |
| gpt-oss-120b (high) | OpenAI | 12 | 0.11 | 184 |
| Nemotron Cascade 2 30B A3B | NVIDIA | 12 | — | — |
| HyperNova 60B 2605 (high) | Multiverse Computing | 12 | — | — |
| Mercury 2.5 | Inception | 12 | 0.12 | 818 |
| Mistral Small 4 | Mistral | 11 | 0.02 | 158 |
| Qwen3 Next 80B A3B | Alibaba | 11 | — | 163 |
| Granite 4.2 8B | IBM | 11 | 0.02 | 52 |
| Ling 3.0 Tiny | InclusionAI | 11 | ~0 | 55 |
| Trinity Large Thinking | Arcee AI | 11 | 0.12 | 339 |
| HyperCLOVA X SEED Think (32B) | Naver | 11 | — | — |
| Mi:dm K 2.5 Pro | Korea Telecom | 11 | — | — |
| INTELLECT-3 | Prime Intellect | 11 | — | — |
| K2 Think V2 | Institute of Foundation Models | 11 | — | — |
| LongCat Flash Lite | LongCat | 11 | — | — |
| Qwen3.5 9B | Alibaba | 11 | 0.21 | 64 |
| Nemotron 3 Nano Omni 30B A3B | NVIDIA | 10 | — | 240 |
| Llama 4 Maverick | Meta | 10 | — | 35 |
| North Mini Code | Cohere | 10 | ~0 | 49 |
| gpt-oss-20b (high) | OpenAI | 9 | 0.01 | 175 |
| Qwen3 Coder Next | Alibaba | 9 | 0.55 | 88 |
| Granite 4.2 3B | IBM | 9 | 0.01 | 215 |
| Nemotron 3 Nano | NVIDIA | 9 | 0.02 | 219 |
| Nova Premier | Amazon | 9 | — | — |
| Gemma 4 E4B | 9 | — | 27 | |
| Sarvam 105B (high) | Sarvam | 9 | — | — |
| MiniCPM5-1B | OpenBMB | 9 | — | — |
| Magistral Small 1.2 | Mistral | 9 | — | — |
| Llama 4 Scout | Meta | 8 | — | 73 |
| Llama 3.3 70B | Meta | 8 | — | 88 |
| Gemma 4 E2B | 8 | — | — | |
| Nanbeige4.1-3B | Nanbeige | 8 | — | — |
| LFM2.5-2.6B | Liquid AI | 8 | — | — |
| EXAONE 4.0 32B | LG AI Research | 8 | — | — |
| Hermes 4 70B | Nous Research | 8 | — | — |
| Falcon-H1R-7B | TII UAE | 8 | — | — |
| Qwen3 Omni 30B A3B | Alibaba | 8 | — | 103 |
| Step3 VL 10B | StepFun | 8 | — | — |
| Llama Nemotron Ultra | NVIDIA | 8 | — | — |
| ERNIE 4.5 300B A47B | Baidu | 8 | — | — |
| Command A | Cohere | 7 | — | 55 |
| Hermes 4 405B | Nous Research | 7 | — | 39 |
| NVIDIA Nemotron Nano 12B v2 VL | NVIDIA | 7 | — | — |
| NVIDIA Nemotron Nano 9B V2 | NVIDIA | 7 | — | 18 |
| Kimi Linear 48B A3B Instruct | Kimi | 7 | — | — |
| Llama 3.1 405B | Meta | 7 | — | — |
| LFM2.5-8B-A1B | Liquid AI | 7 | — | — |
| Ring-flash-2.0 | InclusionAI | 7 | — | — |
| Olmo 3.1 32B Think | Allen Institute for AI | 7 | — | — |
| Qwen3.5 2B | Alibaba | 7 | — | — |
| Sarvam 30B (high) | Sarvam | 7 | — | — |
| Celeris-1 | Celeris | 6 | 0.05 | 1,523 |
| Nova Micro | Amazon | 6 | — | 279 |
| Phi-4 Mini | Microsoft | 6 | — | 45 |
| Ministral 3 14B | Mistral | 6 | 0.02 | 71 |
| Olmo 3.1 32B Instruct | Allen Institute for AI | 6 | — | — |
| R1 1776 | Perplexity | 6 | — | — |
| Llama 3.2 90B (Vision) | Meta | 6 | — | — |
| DeepHermes 3 – Mistral 24B | Nous Research | 6 | — | — |
| Jamba 1.7 Large | AI21 Labs | 6 | — | — |
| Granite 4.0 H Small | IBM | 6 | — | 15 |
| LFM2 24B A2B | Liquid AI | 6 | — | — |
| Phi-4 | Microsoft | 6 | — | 40 |
| Phi-4 Multimodal | Microsoft | 6 | — | — |
| MiniCPM-V 4.6 1.3B | OpenBMB | 6 | — | — |
| Jamba Reasoning 3B | AI21 Labs | 6 | — | — |
| Reka Flash 3 | Reka AI | 6 | — | 94 |
| Olmo 3 7B Think | Allen Institute for AI | 6 | — | — |
| Molmo 7B-D | Allen Institute for AI | 6 | — | — |
| Qwen3.5 0.8B | Alibaba | 6 | — | — |
| Ministral 3 3B | Mistral | 5 | 0.01 | 222 |
| Ministral 3 8B | Mistral | 5 | 0.01 | 101 |
| Llama 3.2 11B (Vision) | Meta | 5 | — | 17 |
| Exaone 4.0 1.2B | LG AI Research | 5 | — | — |
| Olmo 3 7B | Allen Institute for AI | 5 | — | — |
| LFM2.5-1.2B-Thinking | Liquid AI | 5 | — | — |
| Jamba 1.7 Mini | AI21 Labs | 5 | — | — |
| Granite 4.0 H 1B | IBM | 5 | — | — |
| Gemma 3 270M | 5 | — | — | |
| Apertus 70B Instruct | Swiss AI Initiative | 5 | — | — |
Two patterns stand out. First, Anthropic fills five of the top ten rows, and every Claude entry in this snapshot carries a “with fallback” label. Second, the gap between first and tenth is only 12 points, yet the price gap is huge: GPT-6.1 Sol at max effort reaches 52 for about 72 cents per task, roughly one-eighth of what the leader spends for 6 extra points.
Reasoning effort: one model, many price tags
Modern flagship models are no longer a single product. Vendors let you dial how long a model thinks before it answers, and the leaderboard lists each setting as its own row. This changes how you should read the rankings: “Claude Opus 5.5” is really five different cost-and-quality choices.
| Setting | Opus 5.5 score | Opus 5.5 cost per task (USD) | GPT-6.1 Sol score | GPT-6.1 Sol cost per task (USD) |
|---|---|---|---|---|
| max | 58 | 5.98 | 52 | 0.72 |
| xhigh | 56 | 3.46 | 51 | 0.39 |
| high | 54 | 1.82 | 50 | 0.32 |
| medium | 51 | 1.34 | 48 | 0.21 |
| low | 42 | 0.55 | 42 | 0.13 |
The curve flattens quickly. Dropping Opus 5.5 from max to high gives up 4 points but cuts the cost per task by about 70 percent. Dropping GPT-6.1 Sol from max to high costs only 2 points and saves more than half the spend. The steepest cliff is at the bottom: both models lose roughly 6 to 9 points when moved from medium to low.
For most real work, the sweet spot is therefore one or two notches below the top setting. Reserve maximum effort for problems where a wrong answer is expensive, such as complex code changes or multi-step research, and use medium or high for everything else.
How each vendor is doing
Anthropic has the broadest presence at the top. Opus 5.5 and Sonnet 5.5 hold the first two places on the Artificial Analysis index, and Fable 5.1 sits just behind. The newest arrival is Claude Haiku 5.5, released on 7 October: at max effort it scores 43 while costing about 21 cents per task and generating around 240 tokens per second, which makes it a strong small-model option. LLM Stats also lists Claude Mythos Preview at 54.4, but flags it as unreleased, so it cannot be used yet.
OpenAI competes with three tiers. GPT-6 Astra is the premium option and, per LLM Stats, posts the best GPQA Diamond result in its tracker at 96.0 percent. GPT-6.1 Sol, which arrived on 29 September, delivers scores close to Astra’s at a fraction of the price: LLM Stats lists its launch pricing at $2 input and $10 output per million tokens, against $10 and $50 for Astra. Below them, GPT-5.6 Terra and the GPT-6 Luna family cover cheaper, faster work.
Google made news with Gemini 4 Argon, announced on 30 September. It scores 53 on the Artificial Analysis index at high effort, putting it level with Fable 5.1 and Astra, but speed and latency are still blank because access is limited. Google’s strength today is the Flash line: Gemini 3.8 Flash scores 41 at high effort with output near 125 tokens per second.
Meta is a surprise contender. Muse Spark 1.3 reaches 48 at max effort, and its xhigh setting trades 3 points for a jump to 280 tokens per second.
SpaceXAI has Grok 4.7, which scores 46 at its top settings and offers a 500k-token window, smaller than the 1M seen on most rivals.
Open-weight and Asian labs are closing in. LLM Stats labels Kimi K3 (52.1), GLM-5.3 (51.5) and Qwen3.8 Max (51.1) as open source, placing them within about three points of GPT-6.1 Sol (53.9) on its composite. Artificial Analysis scores the same models lower, at 44 to 45, so check which scale you are reading. Xiaomi’s MiMo-V2.6-Pro (46) is notable for costing only 13 cents per task, and DeepSeek V4.1 Flash reaches 39 at 27 cents with about 217 tokens per second.
Mistral Large 4, in public preview since 6 October, scores 38 on the Artificial Analysis index, so it is a competitive European option rather than a frontier leader.
Speed and latency: fast is not the same as smart
Two different things get called “speed”. Output speed is how many tokens per second a model streams once it starts writing. Latency is how long you wait before the first piece of the answer appears. A model can be strong at one and poor at the other.
On output speed, the chart-topper is Celeris-1 at roughly 1,500 tokens per second on Artificial Analysis (BestAIModels records about 1,800). Mercury 2.5 follows at around 820, then Gemini 3.5 Flash-Lite at 365. The catch is quality. Celeris-1 scores only 6 on the intelligence index and Mercury 2.5 scores 12, so they suit simple, high-volume jobs such as tagging, routing or short rewrites rather than analysis.
If you want speed without giving up too much capability, look one tier up:
| Model | Index score | Output speed (tokens/s) | Cost per task (USD) |
|---|---|---|---|
| Claude Haiku 5.5 (max) | 43 | 240 | 0.21 |
| Ling 3.1 Flash | 41 | 214 | 0.99 |
| DeepSeek V4.1 Flash (max) | 39 | 217 | 0.27 |
| Muse Spark 1.3 (xhigh) | 45 | 280 | 1.37 |
| Claude Sonnet 5.5 (max) | 56 | 141 | 5.46 |
Latency tells a different story, and it is where reasoning settings bite hardest. Artificial Analysis measures the time to the first chunk of output, which includes thinking. At maximum effort, Claude Opus 5.5 needs about 662 seconds before its first chunk and Sonnet 5.5 about 487 seconds. Move Opus 5.5 down to high and that falls to about 36 seconds; Sonnet 5.5 at high takes about 12 seconds and at low under one second. Non-reasoning models are fastest of all: the lowest-latency entries in this snapshot are Gemini 2.5 Flash-Lite and Gemini 2.5 Flash in non-reasoning mode, followed by Claude 4.5 Haiku.
The practical rule is simple. Anything a person waits on live, such as a chatbot or an autocomplete, should use a low-latency setting. Maximum-effort modes belong in background jobs where nobody is watching a spinner.
Cost and value: where the budget goes furthest
The absolute cheapest entry on Artificial Analysis is GPT-6 Luna at low effort, at under half a cent per task, followed by IBM’s Granite 4.2 3B and Mistral’s Ministral 3 3B at about one cent. All three are tiny-budget tools with modest scores (22, 9 and 5), so “cheapest” is only useful when the job is easy.
The more interesting question is which models give strong scores for little money. These stand out:
| Model (setting) | Index score | Cost per task (USD) |
|---|---|---|
| GPT-6.1 Sol (xhigh) | 51 | 0.39 |
| GPT-6.1 Sol (high) | 50 | 0.32 |
| GPT-6.1 Sol (medium) | 48 | 0.21 |
| MiMo-V2.6-Pro | 46 | 0.13 |
| GLM-5.3-Flash | 42 | 0.25 |
| Claude Haiku 5.5 (high) | 38 | 0.08 |
| GPT-6 Luna (max) | 38 | 0.07 |
| MiMo-V2.6-Flash | 38 | 0.06 |
GPT-6.1 Sol is the clear value pick among near-frontier models: its medium setting scores 48, higher than every Gemini and Grok entry in this snapshot’s top bracket except Gemini 4 Argon, for about 21 cents a task. Further down, MiMo-V2.6-Pro reaches 46 for just 13 cents, though its output speed is a slow 43 tokens per second.
BestAIModels, which prices by the million tokens rather than per task, names GLM-5.3-Flash its best-value model at about $0.12 per million tokens, while also noting its very high speed of more than 1,200 tokens per second. Remember that token price and cost per task can point in different directions: a cheap model that thinks for a very long time may cost more than a pricier one that answers briefly.
Context windows: one million is the new baseline
The context window is the amount of text a model can consider at once, covering your prompt, any documents and its own reply. A year ago a million tokens was a headline feature. Today it is the default for nearly every leading model: all the Claude models, the GPT-6 and GPT-5.6 families, Gemini, Muse Spark, Kimi K3 and GLM-5.3 are listed at about 1M or slightly above in the Artificial Analysis table.
The exceptions are more interesting than the rule:
| Model | Context window | Note |
|---|---|---|
| Llama 4 Scout | 10M (Artificial Analysis) | BestAIModels lists 1.3M; the sources disagree |
| Grok 4 Fast Reasoning | 2.0M (LLM Stats) | Largest figure in that tracker |
| Kimi K3 | 1.05M | Open-weight, near-frontier score |
| Qwen3.8 Max | 984k | Just under 1M |
| Grok 4.7 | 500k | Half the usual window |
| Qwen3.8 27B | 256k | Smaller open model |
| Command A+ | 192k | Cohere’s latest |
The sources disagree on the leader because they define “largest” differently: one counts vendor-advertised limits, another reports figures from its own testing. If you plan to feed whole books or codebases into a model, check the vendor’s documentation and run a test, because quality often drops long before the stated limit is reached.
For most blog, SEO and support workflows, 200k tokens is already generous. A 1M window mainly pays off for large document sets, long chat histories and codebase-wide analysis.
Which model should you pick?
The picks below are our reading of the numbers above, not a ranking published by any of the three sources. Match the model to the job and the setting to the deadline.
| Your job | Sensible choice | Why |
|---|---|---|
| Hardest reasoning, research or complex code, quality first | Claude Opus 5.5 (high or xhigh) | Top scores on both Artificial Analysis and LLM Stats; high keeps cost per task near $1.82 |
| Strong all-rounder at a lower price | Claude Sonnet 5.5 or GPT-6.1 Sol (high) | Sonnet scores 56 at max with 141 tokens/s; Sol scores 50 at about $0.32 per task |
| Long articles, product descriptions, bulk content drafts | GPT-6.1 Sol (medium) or Claude Haiku 5.5 | Good quality per dollar, quick output |
| Real-time chat or live assistants | A low-effort or non-reasoning setting | Lowest time to first token |
| Massive volume of simple tagging, classification or routing | GPT-6 Luna (low), Mercury 2.5 or Celeris-1 | Cents or fractions of a cent per task, very fast |
| Self-hosting or open weights | Kimi K3, GLM-5.3 or Qwen3.8 | Within a few points of closed leaders on LLM Stats |
| Very large document sets | Any 1M-context model, then test | Window size alone does not guarantee recall |
A good habit is to run your own ten sample tasks through two or three candidates before committing. Leaderboards narrow the field quickly, but only your prompts reveal which model matches your tone, your formatting rules and your accuracy bar.
What leaderboards cannot tell you
Benchmarks are a useful compass, not a verdict. Several limits are worth keeping in mind:
- Scores shift with the test mix. A model that leads on reasoning can trail on writing quality or tool use.
- Many vendor-reported results are self-reported. LLM Stats flags several launch figures this way, so independent reruns may differ.
- Some rows are provisional. Gemini 4 Argon has no speed data yet, and its advertised introductory price was not yet purchasable at the time of the LLM Stats write-up.
- Rankings move weekly. Six notable models launched in the fortnight before this snapshot, so revisit the data before any big purchase.
Frequently asked questions
Which AI model ranks first right now?
Claude Opus 5.5 leads on Artificial Analysis (58 at max effort), on LLM Stats (60.5) and on BestAIModels (57.6).
Which is the fastest model?
Celeris-1 for raw output speed, though it scores only 6 on the intelligence index. Among capable models, Muse Spark 1.3 at xhigh (280 tokens per second) and Claude Haiku 5.5 (about 240) are quick.
Which is the most affordable?
GPT-6 Luna at low effort costs under half a cent per task, but for stronger results GPT-6.1 Sol at medium or MiMo-V2.6-Pro offers far better balance.
Which open-weight model is best?
On LLM Stats, Kimi K3 leads the open group at 52.1, with GLM-5.3 and Qwen3.8 Max close behind.
Does a higher reasoning setting always help?
Not always. It raises the score but also the cost and the wait, and the gains shrink at the top end.
How current is this data?
It reflects snapshots from 11 October 2026 and will date quickly.


