Key takeaways
- Claude Opus 5.5 leads Artificial Analysis’s independent index, at $4 in and $20 out per million tokens.
- GPT-6.1 Sol scores 94.2% on ARC-AGI-2 for $0.25 a task, close to GPT-6 Astra’s 95.0% at $1.12.
- Long prompts reorder the price list, because GPT-6, Gemini 3.1 Pro, Grok 4.7 and Haiku 5.5 charge more above a threshold.
- Within one model, the effort setting changes cost per task by ten times or more.
- Compare benchmark scores inside one leaderboard, since test harnesses move results by tens of points.
As of 8 October 2026, the model with the highest score on Artificial Analysis’s independent Intelligence Index is Claude Opus 5.5, which costs $4 per million input tokens and $20 per million output tokens. Claude Sonnet 5.5 comes second on the same index at $2 and $10, the price of GPT-6.1 Sol. The most expensive tier, Claude Fable 5.1 and GPT-6 Astra at $10 and $50, leads some tests but no longer leads every one. At the low end, GPT-6 Luna and Claude Haiku 5.5 both cost $0.10 and $0.50.
This post puts 23 models from 11 labs side by side. It covers list prices, what three typical requests cost on each model, independent benchmark results, speed, context limits and the open-weight options. Every figure links to its source, vendor-reported numbers are labelled as such, and everything is dated, because prices and line-ups now change every few weeks.
We choose models for the products we build, so the last sections cover how we use these numbers: which costs actually show up on a bill, which benchmarks still separate the models, and the rules we follow when picking one.
Prices per million tokens
The table lists standard pay-as-you-go API prices in US dollars per million tokens, as of 8 October 2026. “Cached” is the price of input served from the provider’s prompt cache. Batch discounts, enterprise agreements and regional surcharges aren’t included.
| Model | Input | Output | Cached input | Context | Notes |
|---|---|---|---|---|---|
| Anthropic1 | |||||
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | 1M | One price across the full context |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | 1M | Fast mode $8 / $40 |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.10 | 1M | |
| Claude Haiku 5.5 | $0.10 | $0.50 | $0.01 | 1M | Prompts over 100K: $0.50 / $2.50 |
| OpenAI2 | |||||
| GPT-6 Astra | $10.00 | $50.00 | $1.00 | 1.05M | Over 272K input: $20 / $75 for the whole request |
| GPT-6.1 Sol | $2.00 | $10.00 | $0.10 | 1.05M | Over 272K input: $4 / $15 |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | 1.05M | Over 272K input: $0.20 / $0.75 |
| Google3 | |||||
| Gemini 3.1 Pro (preview) | $2.00 | $12.00 | $0.20 | 1M | Prompts over 200K: $4 / $18 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | 1M | Introductory price until 31 December 2026, then $1.50 / $7.50 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.03 | 1M | |
| xAI4 | |||||
| Grok 4.7 | $2.00 | $6.00 | $0.50 | 500K | Prompts of 200K or more: $4 / $12 |
| Grok 4.3 | $1.25 | $2.50 | $0.20 | 1M | Prompts of 200K or more: $2.50 / $5 |
| Meta5 | |||||
| Muse Spark 1.3 (preview) | $1.25 | $4.25 | $0.15 | 1M | |
| Mistral6 | |||||
| Mistral Large 4 (preview) | $1.36 | $4.18 | $0.14 | 1M | Half price for two weeks from its 6 October launch |
| Mistral Medium 3.5 | $1.50 | $7.50 | not listed | 256K | |
| Mistral Small 4 | $0.15 | $0.60 | not listed | 256K | |
| DeepSeek7 | |||||
| DeepSeek V4-Pro-0813 | $1.32 | $3.96 | $0.044 | 1M | Half price off-peak |
| DeepSeek V4.1-Flash | $0.30 | $1.20 | $0.006 | 1M | Half price off-peak |
| Alibaba8 | |||||
| Qwen3.8-Max | $2.00 | $6.00 | $0.25 | 1M | International price |
| Qwen3.8-Flash | $0.15 | $0.47 | $0.016 | 1M | |
| Other labs | |||||
| Kimi K3 (Moonshot)9 | $3.00 | $15.00 | $0.30 | 1M | Cache writes billed separately |
| MiMo-V2.6-Pro (Xiaomi)10 | $0.435 | $0.87 | $0.0036 | 1M | |
| GLM-5.3 (Z.ai)11 | $1.40 | $4.40 | $0.26 | 1M |
Some notes on the table:
- Three of these models are labelled preview by their makers: Gemini 3.1 Pro, Muse Spark 1.3 and Mistral Large 4. Preview models can change or move to a new ID with little notice.
- Gemini 4 Argon, announced on 30 September, is first on LMArena’s text leaderboard. Google has opened it only to vetted cyber-defence teams so far, so it has no public price.12
- DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. Outside them, both DeepSeek models cost half the listed price.
- Gemini 3.8 Flash doubles in price on 1 January 2027 unless Google extends the offer.
What three common requests cost
Per-token prices are hard to compare in your head, so here are three requests we see often, priced per thousand requests.
- Chat answer: 8,000 input tokens (instructions plus retrieved context) and a 500-token reply.
- Agent step: 50,000 input tokens, of which 45,000 are a cached prefix, and 2,000 output tokens. The cache is assumed warm, so there are no cache-write fees.
- Long document: a 300,000-token document and a 2,000-token answer.
None of the three includes reasoning tokens, which are billed as output and covered in the next section.
| Model | Chat answer | Agent step | Long document |
|---|---|---|---|
| Claude Fable 5.1 | $105.00 | $161.25 | $3,100.00 |
| Claude Opus 5.5 | $42.00 | $69.00 | $1,240.00 |
| Claude Sonnet 5.5 | $21.00 | $34.50 | $620.00 |
| Claude Haiku 5.5 | $1.05 | $1.95 | $155.00 |
| GPT-6 Astra | $105.00 | $195.00 | $6,150.00 |
| GPT-6.1 Sol | $21.00 | $34.50 | $1,230.00 |
| GPT-6 Luna | $1.05 | $1.95 | $61.50 |
| Gemini 3.1 Pro | $22.00 | $43.00 | $1,236.00 |
| Gemini 3.8 Flash | $7.88 | $14.63 | $232.50 |
| Grok 4.7 | $19.00 | $44.50 | $1,224.00 |
| DeepSeek V4.1-Flash (peak) | $3.00 | $4.17 | $92.40 |
| Qwen3.8-Max | $19.00 | $33.25 | $612.00 |
| Kimi K3 | $31.50 | $58.50 | $930.00 |
Three patterns stand out.
Models with the same list price drift apart on long inputs. GPT-6.1 Sol and Claude Sonnet 5.5 cost exactly the same for the first two requests. On the 300K-token document, Sol passes its 272K threshold and the whole request is billed at $4 and $15, so it costs $1,230 per thousand against $620 for Sonnet. GPT-6 Astra and Claude Fable 5.1 split the same way, at $6,150 and $3,100.
The small models split even earlier. GPT-6 Luna and Claude Haiku 5.5 match on short requests, but Haiku’s higher rate starts at 100K tokens, so the long document costs $155 on Haiku and $61.50 on Luna.
Cached input decides the agent step. Grok 4.7 is cheaper than Claude Sonnet 5.5 for a chat answer, $19 against $21, and dearer for the agent step, $44.50 against $34.50, because a cached token costs $0.50 on Grok and $0.10 on Sonnet. If your product re-sends a long system prompt or conversation on every turn, compare the cached column first.
Three things per-token prices don’t show
Tokenizers differ. The same text becomes a different number of tokens on each provider. Anthropic says the tokenizer it introduced with Claude Opus 4.7 produces approximately 30% more tokens for the same text than its previous one.1 Its model overview puts 1M tokens at roughly 555,000 English words on the current tokenizer, against about 750,000 before.13 Other labs’ tokenizers differ again. Before you compare prices closely, count tokens for a sample of your own prompts with each provider’s token counter.
Reasoning is billed as output. Most of the models in this post think before they answer, and those thinking tokens are billed at the output rate whether you see them or not. They can outnumber the answer. In one of our own pipelines, 187,152 of 572,209 output tokens were reasoning and only 49,544 were prose that shipped, as we described in why our LLM pipeline reported 0 dropped. The effort setting decides how much a model thinks, and the section on effort below shows how much it moves cost.
Thresholds and fees. Four of the model families above charge more once a prompt passes a size threshold: Claude Haiku 5.5 above 100K tokens, Gemini 3.1 Pro above 200K, Grok above 200K and GPT-6 above 272K. Prompt caching has a write cost as well as a read discount. Anthropic charges 1.25 times the input price to write to its 5-minute cache, OpenAI lists cache writes for GPT-6 Astra at $12.50 per million, and Moonshot bills Kimi K3 cache writes at $3 or $6 depending on how long they’re kept.129 Batch processing cuts the price in half at Anthropic and OpenAI, while Grok 4.7 and Kimi K3 have no batch option. Anthropic also charges 1.1 times the normal price when you pin inference to the United States.1
Benchmarks that still separate the top models
Many familiar benchmarks no longer tell frontier models apart. Vals has archived SWE-bench Verified, GPQA, LiveCodeBench, AIME and MMMU-Pro as saturated.14 Epoch AI labels SWE-bench Verified “flawed”,15 and MathArena has deprecated AIME.16 ARC-AGI-2 is close to its ceiling too, with a top score of 95.0%, and ARC Prize now treats ARC-AGI-3 as the harder test.17
The table uses five measures that still spread the field:
- Artificial Analysis Intelligence Index, version 4.3.2, a composite of several evaluations, at each model’s highest reasoning setting.18 The September revision dropped GPQA and added Terminal-Bench 4.0, so these scores can’t be compared with articles written before it.
- LMArena text score, from human preference votes, with style control on (the default). The board is dated 2 October.19
- ARC-AGI-2, on ARC Prize’s semi-private set, with the cost per task at each model’s best setting.17
- Humanity’s Last Exam (HLE), as Artificial Analysis runs it, on a 2,158-question text-only subset with no tools.18
- Terminal-Bench 4.0, as Vals runs it, with the same agent harness for every model, averaged over three runs.20
| Model | AA index | LMArena (rank) | ARC-AGI-2 (cost per task) | HLE | Terminal-Bench 4.0 |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 58 | 1504 (#4) | 93.3% ($0.41) | 61.4% | 65.2%¹ |
| Claude Sonnet 5.5 | 56 | 1471 (#45) | – | – | 64.1%¹ |
| Claude Fable 5.1 | 53 | 1501 (#6) | 90.0% ($3.12) | 59.1% | 58.1%¹ |
| GPT-6 Astra | 53 | 1477 (#29) | 95.0% ($1.12) | 54.7% | 59.6% |
| Gemini 4 Argon (restricted) | 53 | 1525 (#1)² | – | 57.1% | 57.6% |
| GPT-6.1 Sol | 52 | 1483 (#21)³ | 94.2% ($0.25) | 52.9% | 55.1% |
| Muse Spark 1.3 | 48 | 1494 (#9) | – | 48.7% | 24.8% |
| Grok 4.7 | 46 | 1442 (#91) | 61.4% ($2.01)⁴ | 43.1% | 28.8% |
| MiMo-V2.6-Pro | 46 | 1480 (#26) | – | 49.4% | 31.3% |
| GLM-5.3 | 45 | 1478 (#27) | – | 42.3% | 38.9% |
| Qwen3.8-Max | 45 | 1482 (#23) | – | 43.1% | 34.3% |
| Kimi K3 | 44 | 1488 (#16) | 60.4% ($1.59) | 46.9% | 17.2% |
| Claude Haiku 5.5 | 43 | – | – | – | 35.4% |
| Gemini 3.8 Flash | 41 | 1495 (#8)² | 89.2% ($0.40)⁴ | 47.8% | 19.2% |
| DeepSeek V4.1-Flash | 39 | 1474 (#38) | 72.9% ($0.13)⁴ | 39.2% | 19.7% |
| GPT-6 Luna | 38 | 1443 (#90) | 59.3% ($0.06)⁴ | 38.5% | 13.6% |
| Mistral Large 4 | 38 | – | – | 35.0% | 22.7% |
| DeepSeek V4-Pro-0813 | 36 | 1464 (#56) | 61.3% ($0.60) | 41.0% | 14.1% |
| Gemini 3.1 Pro | 30 | 1487 (#18) | 77.1% ($0.96) | 47.0% | 2.5% |
A dash means no published result as of 8 October 2026.
- Artificial Analysis and Vals run Anthropic’s newest models with Anthropic’s default fallback switched on. When a safety classifier declines a request, the API retries it on an older Claude model, and the answer counts towards the newer model’s score.21 Vals reports that counting those attempts as failures would put Opus 5.5 at 58.1% and Fable 5.1 at 50.0% on Terminal-Bench 4.0.20
- LMArena marks this score as preliminary. Argon’s rests on 4,932 votes.
- Based on only 3,071 votes, with a margin of ±11.
- ARC Prize notes partial dataset coverage for this result.
No model leads every column. Claude Opus 5.5 leads the composite index, HLE and Terminal-Bench 4.0. GPT-6 Astra has the best ARC-AGI-2 score, and Gemini 4 Argon is first on LMArena while few people can use it.
On ARC-AGI-2 the cost per task matters as much as the score. GPT-6.1 Sol reaches 94.2% for $0.25 a task, and Claude Opus 5.5 reaches 93.3% for $0.41. GPT-6 Astra’s extra 0.8 points over Sol cost more than four times as much per task, and Fable 5.1 spends $3.12 a task for 90.0%.
The measures also disagree about some models. Gemini 3.1 Pro is 18th on LMArena, where people vote on chat answers, but scores 30 on Artificial Analysis’s index and 2.5% on Terminal-Bench 4.0, which test multi-step tasks. A chat leaderboard and an agent benchmark measure different things, and a product needs the one closest to its own work.
Vendor-reported numbers need a second source
The labs’ own launch numbers use their own test harnesses, prompts and settings, and independent runs often come out lower. These are some of the larger gaps in the data we collected.
| Model | Benchmark | Reported by the lab | Independent results |
|---|---|---|---|
| DeepSeek V4-Pro-0813 | Terminal-Bench 2.1 | 87.9 | 78.7 (Artificial Analysis), 54.7 (Vals) |
| GLM-5.3 | Terminal-Bench 2.1 | 88.2 | 83.9 (Artificial Analysis), 71.5 (Vals) |
| gpt-oss-120b | SWE-bench Verified | 62.4 | 26.0 (swebench.com), 33.6 (Vals) |
| Mistral Medium 3.5 | GPQA Diamond | not reported | 74.8 (Artificial Analysis), 34.8 (Vals) |
The last row is two independent leaderboards disagreeing with each other, which shows how much the harness matters even without a vendor involved.
HLE has the extra problem that “HLE” now names several tests. Labs variously report scores with tools, without tools, on the full set, on a text-only subset and on a cleaned-up “HLE-Verified” set. Scale’s leaderboard gives Claude Fable 5.1 46.5% on the full set at the xhigh setting,22 while Artificial Analysis gives it 59.1% on its text-only subset at max. Both are honest numbers for different tests. Compare models inside one leaderboard, and check which version of a benchmark a claim refers to before you quote it.
Effort settings change cost more than model choice
Most current models let you set how hard they think, usually as an effort level from low to max. Artificial Analysis tests each setting separately, which shows how much the choice matters within one model.
| Model and effort | AA index | Cost per AA task | First token |
|---|---|---|---|
| Claude Opus 5.5, low | 42 | $0.55 | 8.7 s |
| Claude Opus 5.5, medium (default) | 51 | $1.34 | 22.5 s |
| Claude Opus 5.5, high | 54 | $1.82 | 37.6 s |
| Claude Opus 5.5, xhigh | 56 | $3.46 | 132.0 s |
| Claude Opus 5.5, max | 58 | $5.98 | 731.5 s |
| Claude Sonnet 5.5, low | 36 | $0.35 | 1.1 s |
| Claude Sonnet 5.5, medium | 41 | $0.48 | 2.1 s |
| Claude Sonnet 5.5, high (default) | 47 | $0.88 | 13.0 s |
| Claude Sonnet 5.5, xhigh | 52 | $2.01 | 35.9 s |
| Claude Sonnet 5.5, max | 56 | $5.46 | 460.7 s |
The defaults come from Anthropic’s model overview.13 “First token” is Artificial Analysis’s time to the first chunk of the response, which includes thinking time for models that don’t stream their reasoning.
Going from low to max multiplies the cost per task by about 11 on Opus 5.5 and about 16 on Sonnet 5.5. It also turns a wait of seconds into a wait of minutes. GPT-6 Astra shows the same pattern, with a first token after 5.6 s at medium effort and 414.3 s at max.18
The cheaper model per token isn’t always cheaper per task. Opus 5.5 at high effort scores 54 for $1.82 a task, while Sonnet 5.5 at xhigh scores 52 for $2.01. Sonnet 5.5 at max matches Opus 5.5 at xhigh on the index, 56 each, and costs more per task. When you compare two models, compare them at the effort setting you would actually run.
Speed
These are Artificial Analysis’s figures for each lab’s own API, as the median over the previous 72 hours with a 10,000-token prompt.23 Settings differ between models, so the setting is shown next to each name. “First answer token” is the wait before the answer itself starts, after any reasoning.
| Model (setting) | Output tokens per second | First answer token |
|---|---|---|
| Gemini 3.5 Flash-Lite | 346 | 8.6 s |
| DeepSeek V4.1-Flash (reasoning off) | 219 | 1.1 s |
| Mistral Medium 3.5 | 168 | 14.1 s |
| Claude Haiku 5.5 (medium) | 137 | 13.4 s |
| Gemini 3.8 Flash (high) | 125 | 24.3 s |
| Gemini 3.1 Pro | 122 | 23.2 s |
| GPT-6 Luna (low) | 121 | 3.1 s |
| Claude Sonnet 5.5 (medium) | 102 | 2.1 s |
| GLM-5.3 (max) | 79 | 28.0 s |
| Claude Opus 5.5 (medium) | 76 | 22.5 s |
| Grok 4.7 (low) | 69 | 5.1 s |
| Claude Fable 5.1 (medium) | 54 | 7.6 s |
| GPT-6.1 Sol (medium) | 50 | 8.8 s |
| GPT-6 Astra (medium) | 46 | 5.6 s |
| Kimi K3 (max) | 42 | 50.9 s |
| Qwen3.8-Max | 37 | 56.7 s |
Two things matter more than the ranking. Reasoning effort dominates waiting time, as the previous section showed. And open-weight models are often much faster on third-party hosts than on the lab’s own API. Artificial Analysis measured DeepSeek V4.1-Flash at 936.8 tokens per second on one third-party provider, more than four times the first-party figure.24 These numbers also move from day to day, so treat a gap of a few tokens per second as noise.
For a deeper look at latency in a production assistant, including why the median matters less than the slowest tenth of requests, see how we halved an AI stylist’s latency.
Context windows and output limits
A 1M-token context window is now standard at the top. All four current Claude models offer 1M tokens, GPT-6 offers 1.05M with up to 922K of input, Gemini offers 1,048,576, and DeepSeek V4, Qwen3.8, Kimi K3, MiMo-V2.6-Pro and GLM-5.3 all list 1M. The exceptions are Grok 4.7 at 500K, Mistral Medium 3.5 and Small 4 at 256K, and gpt-oss-120b at 131,072. Mistral lists Large 4 at 1M, while Artificial Analysis lists 524K, so check before relying on the larger figure.6
Maximum output varies more. Claude models and GPT-6 return up to 128K tokens per response, Gemini 65,536, Qwen3.8 131K and DeepSeek V4 384K. Kimi K3 defaults to 131,072 and can be raised to 1,048,576.9
A large window doesn’t mean a flat price. Because of the thresholds in the price section, a 300K-token prompt pays double the headline input rate on GPT-6, Gemini 3.1 Pro and Grok, and five times the headline rate on Claude Haiku 5.5. Claude Fable 5.1, Opus 5.5 and Sonnet 5.5, Gemini 3.8 Flash, Qwen3.8-Max and Muse Spark 1.3 charge the same per token across their full window.
Most models take text and images and return text. Gemini 3.x, Muse Spark 1.3 and MiMo-V2.6-Pro also accept audio and video, while DeepSeek V4-Pro and GLM-5.3 take text only.
Open-weight models
The best open-weight model on Artificial Analysis’s index scores 46, against 58 for the top closed model. The open field has moved quickly this year, and the licences need as much attention as the scores.
| Model | AA index | Licence | Notes |
|---|---|---|---|
| MiMo-V2.6-Pro (Xiaomi) | 46 | MIT | Also sold through Xiaomi’s API at $0.435 / $0.87 |
| GLM-5.3 (Z.ai) | 45 | Custom GLM-5.3 licence | LMArena lists it as MIT, but the published licence is custom |
| Kimi K3 (Moonshot) | 44 | Custom Kimi K3 licence | |
| DeepSeek V4.1-Flash | 39 | MIT | |
| Mistral Large 4 | 38 | Not yet named | Mistral says the weights will be released at the end of October |
| DeepSeek V4-Pro-0813 | 36 | MIT | |
| Muse Glimmer (Meta) | 17 | Apache 2.0 | About 30B parameters, dense |
| Gemma 4 31B (Google) | 15 | Apache 2.0 | |
| gpt-oss-120b (OpenAI) | 12 | Apache 2.0 | OpenAI doesn’t sell it; third-party hosts charge a median of $0.15 / $0.59 |
Alibaba also released open weights for a model in the Qwen3.8-Max family, Qwen3.8-2.4T-A95B, under its own custom licence. It is text-only, unlike the API model. Meta hasn’t released a new Llama model in 2026, and Llama 4, from April 2025, is still the latest.5
Custom licences can limit commercial use, the size of company that may use the model, or what you may train with its outputs. Read the licence before you plan to self-host. MIT and Apache 2.0 are the simple cases.
How we choose a model for a product
These are the rules we follow on the AI products we build. They come from running models in production as well as from the tables above.
- Use leaderboards to shortlist, then test on your own tasks. Pick three or four candidates by price tier and context needs, and run them on a set of real inputs with known good outputs.
- Compare cost per completed task. Include reasoning tokens, retries and failed calls. A model with half the per-token price can cost more per task if it needs more effort to get the same result.
- Set the effort level per step. Steps that reshape text or fill a template rarely need much reasoning, so run them on a small model at low effort and keep the larger model for the steps that need judgement.
- Check your prompt sizes against the pricing thresholds. A retrieval step that sometimes sends 300K tokens can double its cost on some providers without any change in your code.
- Keep a second provider wired in. It covers outages, and it gives you somewhere to move a call whose latency degrades. We moved a latency-critical call between endpoints for that reason, as the stylist post describes.
- Recheck every month. Between July and October 2026, OpenAI, Google, xAI, DeepSeek, Alibaba and Mistral all changed prices or replaced models.
- Log what each call costs. A ledger per model and per kind of work catches surprises before the invoice does. How a backfill bug ran up a $1,100 LLM vision bill shows what happens without one.
How this comparison was put together
- Prices come from each provider’s pricing page and model documentation, read on 8 October 2026. They are standard, pay-as-you-go rates in US dollars.
- Benchmarks come from independent leaderboards read on the same day: Artificial Analysis (Intelligence Index v4.3.2, HLE and speed), LMArena’s text leaderboard (dated 2 October), ARC Prize (ARC-AGI-2 semi-private set) and Vals (Terminal-Bench 4.0, dated 7 October). Lab-reported numbers appear only in the section that compares them with independent results.
- Worked costs are our arithmetic from the listed prices. They leave out reasoning tokens, cache writes, batch discounts and taxes.
- Not covered: fine-tuning, embeddings, image and audio generation, enterprise discounts, rate limits and data-residency options beyond the one noted above.
Prices and rankings change often, and corrections are welcome at admin@9io.ai.
-
Anthropic, “Pricing”, Claude API docs, https://platform.claude.com/docs/en/about-claude/pricing ↩↩↩↩
-
OpenAI, “Pricing”, https://developers.openai.com/api/docs/pricing, and model pages for GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna, https://developers.openai.com/api/docs/models/gpt-6-astra ↩↩
-
Google, “Gemini Developer API pricing”, https://ai.google.dev/gemini-api/docs/pricing ↩
-
xAI, model docs for Grok 4.7 and Grok 4.3, https://docs.x.ai/developers/models/grok-4.7 and https://docs.x.ai/developers/models/grok-4.3 ↩
-
Meta, “Models” and “Pricing and rate limits”, https://ai.developer.meta.com/docs/models and https://ai.developer.meta.com/docs/pricing-rate-limits ↩↩
-
Mistral AI, “Mistral Large 4”, 6 October 2026, https://mistral.ai/news/mistral-large-4/, and pricing, https://mistral.ai/pricing ↩↩
-
DeepSeek, “Models and pricing”, https://api-docs.deepseek.com/quick_start/pricing ↩
-
Alibaba Cloud, “Model Studio pricing”, https://www.alibabacloud.com/help/en/model-studio/model-pricing ↩
-
Moonshot AI, “Kimi K3 pricing”, https://platform.kimi.ai/docs/pricing/chat ↩↩↩
-
Xiaomi, “MiMo pay-as-you-go pricing”, https://mimo.mi.com/docs/en-US/price/pay-as-you-go ↩
-
Z.ai, “Pricing”, https://docs.z.ai/guides/overview/pricing ↩
-
Google, “Gemini 4 Argon”, 30 September 2026, https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ ↩
-
Anthropic, “Models overview”, Claude API docs, https://platform.claude.com/docs/en/about-claude/models/overview ↩↩
-
Vals AI, “Benchmarks”, https://www.vals.ai/benchmarks/swebench ↩
-
Epoch AI, “SWE-bench Verified”, https://epoch.ai/benchmarks/swe-bench-verified ↩
-
MathArena, https://matharena.ai/ ↩
-
ARC Prize, “Leaderboard”, https://arcprize.org/leaderboard ↩↩
-
Artificial Analysis, “LLM Leaderboard”, https://artificialanalysis.ai/leaderboards/models, and “Intelligence Index methodology”, https://artificialanalysis.ai/methodology/intelligence-benchmarking ↩↩↩
-
LMArena, “Text Arena leaderboard”, https://arena.ai/leaderboard/text ↩
-
Vals AI, “Terminal-Bench 4.0”, https://www.vals.ai/benchmarks/terminal-bench-4 ↩↩
-
Anthropic, “Refusals and fallback”, Claude API docs, https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback ↩
-
Scale AI, “Humanity’s Last Exam leaderboard”, https://labs.scale.com/leaderboard/humanitys_last_exam ↩
-
Artificial Analysis, “Performance benchmarking methodology”, https://artificialanalysis.ai/methodology/performance-benchmarking ↩
-
Artificial Analysis, “DeepSeek V4.1 Flash: API provider benchmarking”, https://artificialanalysis.ai/models/deepseek-v4-1-flash/providers ↩
Frequently asked questions
Which LLM API is the cheapest in October 2026?
GPT-6 Luna and Claude Haiku 5.5 cost the least, at $0.10 per million input tokens and $0.50 per million output tokens. Haiku 5.5 scores 43 on Artificial Analysis’s index against 38 for Luna, but its price rises to $0.50 and $2.50 on prompts over 100K tokens. DeepSeek V4.1-Flash costs $0.15 and $0.60 off-peak.
Which LLM is the most capable right now?
It depends on the test. As of 8 October 2026, Claude Opus 5.5 leads Artificial Analysis’s Intelligence Index and Vals’ Terminal-Bench 4.0, GPT-6 Astra leads ARC-AGI-2, and Gemini 4 Argon leads LMArena’s text leaderboard, although Argon isn’t generally available.
How much does GPT-6 cost?
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, GPT-6.1 Sol $2 and $10, and GPT-6 Luna $0.10 and $0.50. A request with more than 272K input tokens is billed at a higher rate for the whole request, $20 and $75 on Astra.
How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 and batch requests at half price. The price is the same across its full 1M-token context. Fast mode, on Anthropic’s own API, costs $8 and $40.
Can I trust the benchmark scores AI labs publish?
Treat them as a starting point. Labs run their own harnesses and settings, and independent runs of the same benchmark can differ by tens of points, as with DeepSeek V4-Pro on Terminal-Bench 2.1 (87.9 reported, 54.7 on Vals). Compare models inside one independent leaderboard, then test on your own tasks.
Why is my LLM bill higher than price times tokens?
Usually because of reasoning tokens, which are billed as output, plus long-prompt surcharges, cache-write fees and differences between tokenizers. Log input, cached, output and reasoning tokens for every call to see which one it is.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.