All articles
Comparisons·7 min read

The 2026 AI Price War: What Cheaper Frontier Models Change

On 22 September 2026 Anthropic and OpenAI both cut prices. Mozilla measures the open-weight gap at 4.4 months. What that means for your token bill.

By The omnirouter team

Abstract dark navy chart with descending neon price curves above a stack of shrinking glass panels

The short version

On 22 September 2026, Anthropic and [OI] released cheaper models on the same day. Anthropic's Opus 5.5 landed at $4 per million input tokens and $20 per million output — 20% below Opus 5. [OI] added GPT-6 Sol at $2 / $10 and GPT-6 Luna at $0.10 / $0.50, pricing Luna at a twentieth of Sol.

Three days earlier, on 15 September, Mozilla published its State of Open Source AI report. Its headline finding: the capability gap between frontier closed models and the best open-weight models has narrowed to 4.4 months.

Read together, those two facts describe the AI model price war better than either one alone: capability is converging faster than price is.

A price cut is three different things

Every "−90%" headline is really a mix of three changes with very different half-lives.

1. A real serving-cost reduction. [OI] credits caching and inference improvements for the Luna and Sol rates. This part is structural — better cache reuse and cheaper kernels do not un-invent themselves.

2. A promotional rate. Some comparison prices are explicitly promotional, which is why the same table looks different a quarter later. Treat any "vs list" figure as a claim about today, not a contract.

3. A land-grab. When two labs cut prices within hours of each other, part of that is a fight for default-model position, not a step change in the cost of compute.

None of this makes the cuts fake. It makes them dated. The practical consequence is blunt: a cost model built on who is cheapest this month is a cost model you rebuild every month.

The number that actually decides your bill

Headline prices are per token. Your bill is per task — and a task is not a fixed number of tokens.

A model that is 3× cheaper per million tokens but needs 2.5× more tokens, retries and tool calls to finish the same job is a rounding error in your favour at best. This is exactly why Anthropic can claim "40% less for typical workloads" beside list prices that fell only 20%: the remainder is fewer tokens per task, not a lower rate.

So the only honest measurement is cost per completed task, and it needs three inputs:

  • tokens consumed on success,
  • tokens burned on attempts that failed but still billed,
  • human time spent re-running what failed.

Most teams track the first and not the other two. The other two are where the money goes.

Why open-weight models became the default

Mozilla's recommendation is unusually direct: most organizations should use open models as the default for most work, reserving closed frontier models for a narrow band — expert professional work, high-intensity retrieval, long context.

There is a supply argument for that, and it matters more than the price argument.

Open weights cannot be withdrawn. When Moonshot shipped Kimi K2.6 — 58.6 on SWE-Bench Pro against GPT-5.4's 57.7 — the weights went to Hugging Face under an open licence, and the older K2 series was deprecated on a published date. A model you can self-host does not fail over to a worse provider at 2am because someone else's capacity tightened.

That pattern held all through September:

  • Tencent Hy4 preview — 770B total parameters, 49B active, 1M-token context, Apache 2.0. Tencent showed it beating its own Hy3 in every category and taking 10 of 12 comparisons against larger-named frontier models.
  • Shanghai AI Lab's Atria Dawn Preview — a 744B agentic model under MIT, posted to GitHub and Hugging Face with no launch announcement at all.

Large, permissively licensed, shipped without ceremony: that is the open-weight tier in 2026.

Where the frontier still earns its premium

This is the part price-war coverage tends to skip. Cheap models did not make expensive models pointless.

Pay for the frontier when the task has most of these properties:

  • A wrong answer is expensive. Legal, medical, financial, security-sensitive reasoning.
  • The context is genuinely long and load-bearing — not long because it is padded.
  • Failure is not cheaply observable. If you cannot tell a bad answer from a good one without a human, the cheaper model is a false economy.
  • You are buying the ecosystem. Compliance packaging, support and accountability are purchases, not free extras that arrive with weights.

Everything else — classification, extraction, summarization, first-draft code, agentic grunt work, high-volume batch — is where the cheap tier has already won.

Routing is the only durable advantage

Per-token prices will keep moving. The capability gap will keep narrowing and occasionally widening. Neither is something a normal team controls.

What you control is which model handles which request. Route by task:

  • Default → a stable open-weight or cheap-tier model you have measured on your own workload.
  • Escalate → the frontier model, only for requests that justify it.
  • Validate → when a deterministic check can grade the cheap model's output, you often do not need the premium call at all.

This is a routing problem, not a purchasing problem. It is also why the interesting layer in 2026 is not the model list but the router.

What this means at omnirouter

omnirouter is a cheap, prepaid, per-token gateway. We are not a frontier lab and we do not present ourselves as one — we are a place to buy tokens at a low rate, with no subscription.

Two things about our supply are worth stating plainly, because they follow from everything above.

Our frontier-tier access is cheap but volatile. When we price GPT-6 Luna, GPT-6 Sol or Opus-tier access far below list, that price rests on aggregator capacity we do not own. Capacity moves, and quality can vary with it. We would rather you knew that up front.

Open-weight is the stable backbone. GLM, DeepSeek, Qwen, Kimi and MiniMax are the models we can stand behind for consistency, and they are also where the price-to-capability curve is best. If you are building something that has to keep working, make one of those your default and treat the frontier tier as an escalation path.

Concretely, from our catalog:

ModelInput / 1MOutput / 1MTier
GPT-6 Luna$0.01$0.05frontier, volatile
DeepSeek V4.1 Flash$0.0085$0.03open-weight
Kimi K2.6$0.51$1.65open-weight
GLM 5.3$0.80$0.90open-weight
GPT-6 Sol$0.18$0.90frontier, volatile

Prices and availability change. The live list on our models page is the only version worth trusting.

What to do this week

  1. Measure cost per completed task — not per token — for your three most common jobs.
  2. Set a cheap default and log every escalation. Never escalating means your default is wrong; always escalating means you have not tested the cheap tier.
  3. Assume one dependency will move. Pick a model you could switch to in an afternoon.
  4. Do not build a budget on a promo price. Model the steady-state rate.

The AI model price war is good news for anyone paying per token. It is only durable good news for the teams that route around it.

Sources

ai pricingopen-weight modelsllm routingmodel comparison2026