LLM Leaderboard 2026: Top AI Models by Intelligence, , Speed, Price and Context

Anthropic’s Claude Opus 5.5 tops all three major AI leaderboards this month, but the best model for you depends on how much you value speed, price and context length.

Why the leaderboards disagree

Search for the “best AI model” and you will find several rankings that never quite match. That is not an error. Each site measures something different, and the differences matter when you choose a model.

Three sources shaped this article, each as a snapshot taken on 11 October 2026:

  • Artificial Analysis publishes an Intelligence Index, a single 0–100 score built from a bundle of tests. It also tracks real API behaviour: output speed in tokens per second, time to the first token, total response time and the average cost of finishing a benchmark task.
  • LLM Stats produces its own composite score and splits it into reasoning, coding and agent sub-scores. It also runs head-to-head voting arenas, so human preference feeds into its picture.
  • BestAIModels repackages Artificial Analysis quality data into a shortlist of 30 models, then adds its own blended price per million tokens, speed and context figures.

Because scores come from different test mixes, the same model can land on 58, 60.5 or 57.6 depending on the page. Treat the numbers as a way to compare models inside one table, not as absolute truths across tables. The order at the top, though, is remarkably stable, and that is the real story.

The intelligence race: who sits at the top

On the Artificial Analysis Intelligence Index, the top spot belongs to Claude Opus 5.5 running at its maximum reasoning setting, with a score of 58. Claude Sonnet 5.5 at max effort follows closely on 56, level with Opus 5.5 at the “xhigh” setting. LLM Stats agrees on the order of the leaders: Opus 5.5 scores 60.5 there, Sonnet 5.5 58.3 and OpenAI’s GPT-6 Astra 58.2, so the top three are separated by only a couple of points.

The table below lists 160 models from the Artificial Analysis index. The first ten rows are the strongest entries; the 150 rows after them add other models, one row each at its best-scoring setting, from highest to lowest score. A dash means the source reported no value, and “~0” means a cost that rounds to zero. Cost per task is the average spend to complete the index’s tasks, so it reflects how many tokens a model burns while thinking, not just its list price.

Model (setting)MakerIndex scoreCost per task (USD)Output speed (tokens/s)
Claude Opus 5.5 (max)Anthropic585.9896
Claude Sonnet 5.5 (max)Anthropic565.46141
Claude Opus 5.5 (xhigh)Anthropic563.4681
Claude Opus 5.5 (high)Anthropic541.8277
Claude Fable 5.1 (max)Anthropic537.6370
GPT-6 Astra (max)OpenAI533.2647
Gemini 4 Argon (high)Google531.99not reported
GPT-6.1 Sol (max)OpenAI520.7256
Muse Spark 1.3 (max)Meta481.60175
Grok 4.7 (xhigh)SpaceXAI463.7474
MiMo-V2.6-ProXiaomi460.1343
Qwen3.8 Max (0902)Alibaba455.4136
GLM-5.3 (max)Z AI452.0183
Step 5 PreviewStepFun441.0387
Kimi K3 (max)Kimi442.0041
Claude Haiku 5.5 (max)Anthropic430.21240
GPT-5.6 Terra (max)OpenAI421.40108
GLM-5.3-FlashZ AI420.2553
Ling 3.1 FlashInclusionAI410.99214
Gemini 3.8 Flash (high)Google411.24125
Qwen3.8 2.4T A95BAlibaba402.1637
Qwen3.8-Flash-NextAlibaba400.3756
DeepSeek V4.1 Flash (max)DeepSeek390.27217
Mistral Large 4 PreviewMistral381.13—
GPT-6 Luna (max)OpenAI380.07139
MiMo-V2.6-FlashXiaomi380.0658
DeepSeek V4 Pro 0813 (max)DeepSeek360.6789
DeepSeek V4 Flash Vision (max)DeepSeek350.31221
JT-4.1 Flash 236B A21BChina Mobile34——
Qwen3.8 27B (xhigh)Alibaba341.0146
Motif 3Motif Technologies34——
GPT-5.3 Codex (xhigh)OpenAI33—91
K2 Horizon 375B A23BInstitute of Foundation Models31~0119
Gemini 3.1 Pro PreviewGoogle301.30114
MiniMax-M3MiniMax290.5192
Nex-N2-ProNex AGI28——
Solar Pro 4Upstage28—104
Quasar 438B (max)Multiverse Computing272.02126
Apodex 1.1Apodex260.46—
GPT-5.5 Instant (June 2026)OpenAI260.69139
Kimi K2.7 CodeKimi260.5479
Inkling SmallThinking Machines260.09160
Hy3Tencent250.0783
Qwen3.7 PlusAlibaba250.2254
Inkling (xhigh)Thinking Machines25—151
Ling-3.0-flash-VLInclusionAI25—145
Solar Mini 4Upstage240.3774
Nemotron 3 UltraNVIDIA230.60142
Ling-3.0-flash-FinInclusionAI23—321
Gemini 3.5 Flash-LiteGoogle220.19365
KAT-Coder-Pro V2KwaiKAT22——
A.X-K2SK Telecom21——
o3OpenAI20—116
Qwen3.5 Omni PlusAlibaba20—89
K-EXAONE 2.0LG AI Research20——
LongCat 2.0LongCat190.15—
JT-35B-FlashChina Mobile19——
Qwen3.6 35B A3BAlibaba180.48127
Qwen3.5 397B A17BAlibaba180.4787
Muse Glimmer (high)Meta170.06112
Ring-2.6-1TInclusionAI170.29116
Doubao Seed CodeByteDance Seed17——
Gemma 4 26B A4BGoogle17——
Qwen3.5 122B A10BAlibaba160.32130
Gemma 4 31BGoogle15~035
Mistral Medium 3.5Mistral140.50163
ERNIE 5.0 Thinking PreviewBaidu14——
Nova 2.0 Pro Preview (medium)Amazon14—120
Gemma 4 12BGoogle14—114
Nova 2.0 Omni (medium)Amazon14——
Apriel-v1.6-15B-ThinkerServiceNow13——
Nova 2.0 Lite (high)Amazon13—188
Command A+Cohere13~0175
Nemotron 3.5 LightningNVIDIA130.09307
Nemotron 3 SuperNVIDIA131.64152
Granite 4.2 30BIBM130.0775
EXAONE 4.5 33BLG AI Research13——
Qwen3.5 4BAlibaba13—23
Qwen3.5 Omni FlashAlibaba12—232
MiniCPM5-2BOpenBMB12——
Magistral Medium 1.2Mistral12——
gpt-oss-120b (high)OpenAI120.11184
Nemotron Cascade 2 30B A3BNVIDIA12——
HyperNova 60B 2605 (high)Multiverse Computing12——
Mercury 2.5Inception120.12818
Mistral Small 4Mistral110.02158
Qwen3 Next 80B A3BAlibaba11—163
Granite 4.2 8BIBM110.0252
Ling 3.0 TinyInclusionAI11~055
Trinity Large ThinkingArcee AI110.12339
HyperCLOVA X SEED Think (32B)Naver11——
Mi:dm K 2.5 ProKorea Telecom11——
INTELLECT-3Prime Intellect11——
K2 Think V2Institute of Foundation Models11——
LongCat Flash LiteLongCat11——
Qwen3.5 9BAlibaba110.2164
Nemotron 3 Nano Omni 30B A3BNVIDIA10—240
Llama 4 MaverickMeta10—35
North Mini CodeCohere10~049
gpt-oss-20b (high)OpenAI90.01175
Qwen3 Coder NextAlibaba90.5588
Granite 4.2 3BIBM90.01215
Nemotron 3 NanoNVIDIA90.02219
Nova PremierAmazon9——
Gemma 4 E4BGoogle9—27
Sarvam 105B (high)Sarvam9——
MiniCPM5-1BOpenBMB9——
Magistral Small 1.2Mistral9——
Llama 4 ScoutMeta8—73
Llama 3.3 70BMeta8—88
Gemma 4 E2BGoogle8——
Nanbeige4.1-3BNanbeige8——
LFM2.5-2.6BLiquid AI8——
EXAONE 4.0 32BLG AI Research8——
Hermes 4 70BNous Research8——
Falcon-H1R-7BTII UAE8——
Qwen3 Omni 30B A3BAlibaba8—103
Step3 VL 10BStepFun8——
Llama Nemotron UltraNVIDIA8——
ERNIE 4.5 300B A47BBaidu8——
Command ACohere7—55
Hermes 4 405BNous Research7—39
NVIDIA Nemotron Nano 12B v2 VLNVIDIA7——
NVIDIA Nemotron Nano 9B V2NVIDIA7—18
Kimi Linear 48B A3B InstructKimi7——
Llama 3.1 405BMeta7——
LFM2.5-8B-A1BLiquid AI7——
Ring-flash-2.0InclusionAI7——
Olmo 3.1 32B ThinkAllen Institute for AI7——
Qwen3.5 2BAlibaba7——
Sarvam 30B (high)Sarvam7——
Celeris-1Celeris60.051,523
Nova MicroAmazon6—279
Phi-4 MiniMicrosoft6—45
Ministral 3 14BMistral60.0271
Olmo 3.1 32B InstructAllen Institute for AI6——
R1 1776Perplexity6——
Llama 3.2 90B (Vision)Meta6——
DeepHermes 3 – Mistral 24BNous Research6——
Jamba 1.7 LargeAI21 Labs6——
Granite 4.0 H SmallIBM6—15
LFM2 24B A2BLiquid AI6——
Phi-4Microsoft6—40
Phi-4 MultimodalMicrosoft6——
MiniCPM-V 4.6 1.3BOpenBMB6——
Jamba Reasoning 3BAI21 Labs6——
Reka Flash 3Reka AI6—94
Olmo 3 7B ThinkAllen Institute for AI6——
Molmo 7B-DAllen Institute for AI6——
Qwen3.5 0.8BAlibaba6——
Ministral 3 3BMistral50.01222
Ministral 3 8BMistral50.01101
Llama 3.2 11B (Vision)Meta5—17
Exaone 4.0 1.2BLG AI Research5——
Olmo 3 7BAllen Institute for AI5——
LFM2.5-1.2B-ThinkingLiquid AI5——
Jamba 1.7 MiniAI21 Labs5——
Granite 4.0 H 1BIBM5——
Gemma 3 270MGoogle5——
Apertus 70B InstructSwiss AI Initiative5——

Two patterns stand out. First, Anthropic fills five of the top ten rows, and every Claude entry in this snapshot carries a “with fallback” label. Second, the gap between first and tenth is only 12 points, yet the price gap is huge: GPT-6.1 Sol at max effort reaches 52 for about 72 cents per task, roughly one-eighth of what the leader spends for 6 extra points.

Reasoning effort: one model, many price tags

Modern flagship models are no longer a single product. Vendors let you dial how long a model thinks before it answers, and the leaderboard lists each setting as its own row. This changes how you should read the rankings: “Claude Opus 5.5” is really five different cost-and-quality choices.

SettingOpus 5.5 scoreOpus 5.5 cost per task (USD)GPT-6.1 Sol scoreGPT-6.1 Sol cost per task (USD)
max585.98520.72
xhigh563.46510.39
high541.82500.32
medium511.34480.21
low420.55420.13

The curve flattens quickly. Dropping Opus 5.5 from max to high gives up 4 points but cuts the cost per task by about 70 percent. Dropping GPT-6.1 Sol from max to high costs only 2 points and saves more than half the spend. The steepest cliff is at the bottom: both models lose roughly 6 to 9 points when moved from medium to low.

For most real work, the sweet spot is therefore one or two notches below the top setting. Reserve maximum effort for problems where a wrong answer is expensive, such as complex code changes or multi-step research, and use medium or high for everything else.

How each vendor is doing

Anthropic has the broadest presence at the top. Opus 5.5 and Sonnet 5.5 hold the first two places on the Artificial Analysis index, and Fable 5.1 sits just behind. The newest arrival is Claude Haiku 5.5, released on 7 October: at max effort it scores 43 while costing about 21 cents per task and generating around 240 tokens per second, which makes it a strong small-model option. LLM Stats also lists Claude Mythos Preview at 54.4, but flags it as unreleased, so it cannot be used yet.

OpenAI competes with three tiers. GPT-6 Astra is the premium option and, per LLM Stats, posts the best GPQA Diamond result in its tracker at 96.0 percent. GPT-6.1 Sol, which arrived on 29 September, delivers scores close to Astra’s at a fraction of the price: LLM Stats lists its launch pricing at $2 input and $10 output per million tokens, against $10 and $50 for Astra. Below them, GPT-5.6 Terra and the GPT-6 Luna family cover cheaper, faster work.

Google made news with Gemini 4 Argon, announced on 30 September. It scores 53 on the Artificial Analysis index at high effort, putting it level with Fable 5.1 and Astra, but speed and latency are still blank because access is limited. Google’s strength today is the Flash line: Gemini 3.8 Flash scores 41 at high effort with output near 125 tokens per second.

Meta is a surprise contender. Muse Spark 1.3 reaches 48 at max effort, and its xhigh setting trades 3 points for a jump to 280 tokens per second.

SpaceXAI has Grok 4.7, which scores 46 at its top settings and offers a 500k-token window, smaller than the 1M seen on most rivals.

Open-weight and Asian labs are closing in. LLM Stats labels Kimi K3 (52.1), GLM-5.3 (51.5) and Qwen3.8 Max (51.1) as open source, placing them within about three points of GPT-6.1 Sol (53.9) on its composite. Artificial Analysis scores the same models lower, at 44 to 45, so check which scale you are reading. Xiaomi’s MiMo-V2.6-Pro (46) is notable for costing only 13 cents per task, and DeepSeek V4.1 Flash reaches 39 at 27 cents with about 217 tokens per second.

Mistral Large 4, in public preview since 6 October, scores 38 on the Artificial Analysis index, so it is a competitive European option rather than a frontier leader.

Speed and latency: fast is not the same as smart

Two different things get called “speed”. Output speed is how many tokens per second a model streams once it starts writing. Latency is how long you wait before the first piece of the answer appears. A model can be strong at one and poor at the other.

On output speed, the chart-topper is Celeris-1 at roughly 1,500 tokens per second on Artificial Analysis (BestAIModels records about 1,800). Mercury 2.5 follows at around 820, then Gemini 3.5 Flash-Lite at 365. The catch is quality. Celeris-1 scores only 6 on the intelligence index and Mercury 2.5 scores 12, so they suit simple, high-volume jobs such as tagging, routing or short rewrites rather than analysis.

If you want speed without giving up too much capability, look one tier up:

ModelIndex scoreOutput speed (tokens/s)Cost per task (USD)
Claude Haiku 5.5 (max)432400.21
Ling 3.1 Flash412140.99
DeepSeek V4.1 Flash (max)392170.27
Muse Spark 1.3 (xhigh)452801.37
Claude Sonnet 5.5 (max)561415.46

Latency tells a different story, and it is where reasoning settings bite hardest. Artificial Analysis measures the time to the first chunk of output, which includes thinking. At maximum effort, Claude Opus 5.5 needs about 662 seconds before its first chunk and Sonnet 5.5 about 487 seconds. Move Opus 5.5 down to high and that falls to about 36 seconds; Sonnet 5.5 at high takes about 12 seconds and at low under one second. Non-reasoning models are fastest of all: the lowest-latency entries in this snapshot are Gemini 2.5 Flash-Lite and Gemini 2.5 Flash in non-reasoning mode, followed by Claude 4.5 Haiku.

The practical rule is simple. Anything a person waits on live, such as a chatbot or an autocomplete, should use a low-latency setting. Maximum-effort modes belong in background jobs where nobody is watching a spinner.

Cost and value: where the budget goes furthest

The absolute cheapest entry on Artificial Analysis is GPT-6 Luna at low effort, at under half a cent per task, followed by IBM’s Granite 4.2 3B and Mistral’s Ministral 3 3B at about one cent. All three are tiny-budget tools with modest scores (22, 9 and 5), so “cheapest” is only useful when the job is easy.

The more interesting question is which models give strong scores for little money. These stand out:

Model (setting)Index scoreCost per task (USD)
GPT-6.1 Sol (xhigh)510.39
GPT-6.1 Sol (high)500.32
GPT-6.1 Sol (medium)480.21
MiMo-V2.6-Pro460.13
GLM-5.3-Flash420.25
Claude Haiku 5.5 (high)380.08
GPT-6 Luna (max)380.07
MiMo-V2.6-Flash380.06

GPT-6.1 Sol is the clear value pick among near-frontier models: its medium setting scores 48, higher than every Gemini and Grok entry in this snapshot’s top bracket except Gemini 4 Argon, for about 21 cents a task. Further down, MiMo-V2.6-Pro reaches 46 for just 13 cents, though its output speed is a slow 43 tokens per second.

BestAIModels, which prices by the million tokens rather than per task, names GLM-5.3-Flash its best-value model at about $0.12 per million tokens, while also noting its very high speed of more than 1,200 tokens per second. Remember that token price and cost per task can point in different directions: a cheap model that thinks for a very long time may cost more than a pricier one that answers briefly.

Context windows: one million is the new baseline

The context window is the amount of text a model can consider at once, covering your prompt, any documents and its own reply. A year ago a million tokens was a headline feature. Today it is the default for nearly every leading model: all the Claude models, the GPT-6 and GPT-5.6 families, Gemini, Muse Spark, Kimi K3 and GLM-5.3 are listed at about 1M or slightly above in the Artificial Analysis table.

The exceptions are more interesting than the rule:

ModelContext windowNote
Llama 4 Scout10M (Artificial Analysis)BestAIModels lists 1.3M; the sources disagree
Grok 4 Fast Reasoning2.0M (LLM Stats)Largest figure in that tracker
Kimi K31.05MOpen-weight, near-frontier score
Qwen3.8 Max984kJust under 1M
Grok 4.7500kHalf the usual window
Qwen3.8 27B256kSmaller open model
Command A+192kCohere’s latest

The sources disagree on the leader because they define “largest” differently: one counts vendor-advertised limits, another reports figures from its own testing. If you plan to feed whole books or codebases into a model, check the vendor’s documentation and run a test, because quality often drops long before the stated limit is reached.

For most blog, SEO and support workflows, 200k tokens is already generous. A 1M window mainly pays off for large document sets, long chat histories and codebase-wide analysis.

Which model should you pick?

The picks below are our reading of the numbers above, not a ranking published by any of the three sources. Match the model to the job and the setting to the deadline.

Your jobSensible choiceWhy
Hardest reasoning, research or complex code, quality firstClaude Opus 5.5 (high or xhigh)Top scores on both Artificial Analysis and LLM Stats; high keeps cost per task near $1.82
Strong all-rounder at a lower priceClaude Sonnet 5.5 or GPT-6.1 Sol (high)Sonnet scores 56 at max with 141 tokens/s; Sol scores 50 at about $0.32 per task
Long articles, product descriptions, bulk content draftsGPT-6.1 Sol (medium) or Claude Haiku 5.5Good quality per dollar, quick output
Real-time chat or live assistantsA low-effort or non-reasoning settingLowest time to first token
Massive volume of simple tagging, classification or routingGPT-6 Luna (low), Mercury 2.5 or Celeris-1Cents or fractions of a cent per task, very fast
Self-hosting or open weightsKimi K3, GLM-5.3 or Qwen3.8Within a few points of closed leaders on LLM Stats
Very large document setsAny 1M-context model, then testWindow size alone does not guarantee recall

A good habit is to run your own ten sample tasks through two or three candidates before committing. Leaderboards narrow the field quickly, but only your prompts reveal which model matches your tone, your formatting rules and your accuracy bar.

What leaderboards cannot tell you

Benchmarks are a useful compass, not a verdict. Several limits are worth keeping in mind:

  • Scores shift with the test mix. A model that leads on reasoning can trail on writing quality or tool use.
  • Many vendor-reported results are self-reported. LLM Stats flags several launch figures this way, so independent reruns may differ.
  • Some rows are provisional. Gemini 4 Argon has no speed data yet, and its advertised introductory price was not yet purchasable at the time of the LLM Stats write-up.
  • Rankings move weekly. Six notable models launched in the fortnight before this snapshot, so revisit the data before any big purchase.

Frequently asked questions

Which AI model ranks first right now?

Claude Opus 5.5 leads on Artificial Analysis (58 at max effort), on LLM Stats (60.5) and on BestAIModels (57.6).

Which is the fastest model?

Celeris-1 for raw output speed, though it scores only 6 on the intelligence index. Among capable models, Muse Spark 1.3 at xhigh (280 tokens per second) and Claude Haiku 5.5 (about 240) are quick.

Which is the most affordable?

GPT-6 Luna at low effort costs under half a cent per task, but for stronger results GPT-6.1 Sol at medium or MiMo-V2.6-Pro offers far better balance.

Which open-weight model is best?

On LLM Stats, Kimi K3 leads the open group at 52.1, with GLM-5.3 and Qwen3.8 Max close behind.

Does a higher reasoning setting always help?

Not always. It raises the score but also the cost and the wait, and the gains shrink at the top end.

How current is this data?

It reflects snapshots from 11 October 2026 and will date quickly.

Scroll to Top