What Just Happened

A September 2026 roundup of model releases counted more than 20 new AI models shipped between September 12 and 25 alone, from 28 companies over the month. The headline number is the spread: among the five top models it tracked, the most expensive (Claude Fable 5.1 at $11.90 per million tokens) costs about 119 times the cheapest (Muse Spark 1.3 at $0.10 per million tokens).

Prices are falling at the bottom too. GPT-6 Luna set a new API floor of $0.10 per million input tokens and $0.50 per million output tokens, undercutting DeepSeek V4.1-Flash's off-peak pricing of $0.15 and $0.60. At the top end, Claude Fable 5.1 cut its cache-read price from $1.00 to $0.25 per million tokens.

Why the Spread Matters More Than the Floor

  • One model no longer fits every job. With a 119x gap, sending routine work like tagging, summarizing, or first-draft replies to a premium model is paying luxury prices for commodity tasks.
  • Caching changes the math. A 75% cut in cache-read cost rewards workflows that reuse the same long context, such as support agents that keep re-reading the same product documentation.
  • Cheap is not the same as right. The cheapest model is only a bargain if it is accurate enough for the task. Quality checks on your own data decide that, not a price list.
Quick Insight

Model choice is becoming a routing decision, not a vendor decision. The businesses that benefit most will be the ones that can send each task to the cheapest model that passes their own quality bar, and swap models as prices keep moving.

What We'd Tell a Client Right Now

Do not lock your product to one model. Log what each AI task costs, test two or three models against real examples from your business, and build a thin layer that lets you switch. With prices moving this fast, the ability to change models cheaply is worth more than any single price cut.