← All posts

The Tuesday Intelligence Went on Sale

2026-09-23 · Tech news / LLM economics (frontier price war)

On September 22, 2026, the two leading AI labs shipped new frontier models on the same afternoon — and both cut prices. OpenAI launched GPT-6 Sol ($2/$10 per million tokens, down 50% from GPT-5.6 Sol's promo prices) and GPT-6 Luna ($0.10/$0.50). Anthropic launched Claude Opus 5.5 at $4/$20, a 20% list cut from Opus 5's $5/$25, with cache reads slashed 60% to $0.20 per million.

This is not a normal price war. A normal price war means someone is hurting. This one means inference got structurally cheaper — and both labs decided to pass it through. Here's what a token actually costs, why the price is falling, and why the price list still lies to you.

ELI5: the restaurant kitchen

The bill. When you call an API, you pay per token — roughly 4 characters of text. But tokens aren't free to produce: every one is computed by a GPU doing matrix math. Your prompt is processed once ("prefill"), and then the model generates output one token at a time, each token requiring a full pass through the model ("decode"). Decode is the expensive part — it's serial, and you can't parallelize it. Output tokens cost ~5× input tokens across almost every price list. That's the fundamental asymmetry.

Diagram 1

The printing press. Why is intelligence getting cheaper? Four forces:

The lunch menu. Think of it like a restaurant: the flagship model is the chef's tasting menu; the mini tier is the lunch special, same kitchen, smaller portions. Both labs are expanding the lunch menu.

How it works: the anatomy of an API bill

Your invoice has four lines, and they don't move together:

Diagram 2

Then come the pricing knobs on top:

Anthropic claims Opus 5.5 is ~40% cheaper on typical workloads and 30%+ faster on output than Opus 5. Note the word typical — the fine print matters, as we'll see.

State of the art: the September 22 scoreboard

Model Input $/M Output $/M Cache read $/M Headline benchmark
GPT-6 Sol 2 10 — DeepSWE 68.8%; "half as many mistakes" as predecessor on factuality evals (OpenAI)
GPT-6 Luna 0.10 0.50 — DeepSWE 66.6%; AutomationBench 33.2% at $0.27/task
Claude Opus 5.5 4 20 0.20 Terminal-Bench-Science 0.1: 58.7% (Opus 5: 29%); SWE-bench Pro 89.9%; AutomationBench 40%
Fable 5.1 (Sept 1) 10 50 0.25 Terminal-Bench 4.0: 55.8%

(All benchmark figures are vendor-reported. OpenAI's cuts are attributed to caching + inference efficiency.)

Capability per dollar is exploding. Opus 5.5 scores 58.7% on Terminal-Bench-Science 0.1 — double Opus 5's 29% — while costing 20% less on the price list. Anthropic claims it matches Astra on Terminal-Bench 4.0 for ~40% of the cost and beats GPT-5.6 Sol by 11 points on CursorBench for ~a third of the cost. OpenAI claims Sol makes half as many factual errors as its predecessor while costing half as much. Capability per dollar is improving faster than prices are falling.

Diagram 3

But the per-token vs per-task trap. Artificial Analysis found that at max effort, Fable 5.1 costs ~20% MORE per task than Fable 5 — because it emits ~1.7× as many output tokens, and output tokens are the expensive ones. Its per-token prices are lower; its task bill is higher when it thinks hardest. Anthropic itself concedes the point: "benchmark margins have become a less reliable guide to real-world differences." Optimize your bill for tokens-per-task, not dollars-per-token.

Diagram 4

The benchmark-gaming caveat. Every lab picks the benchmark where it wins: OpenAI leads with factuality and DeepSWE, Anthropic with Terminal-Bench and SWE-bench Pro. METR and Frontier Design ran pre-release evals of Opus 5.5 but haven't published results as of Sept 22 — treat every figure above as vendor-reported, with error bars. It's marketing with a p-value until independent labs weigh in.

The open-weights counter-pressure. Meanwhile Stepfun's Step 5 preview — a 600B sparse MoE with 27B active parameters and 1M-token context — prices at $1/$2.70 per million, scores an Artificial Analysis index of 44 (matching Kimi K3 at ~4× the size), and plans open weights on October 15. If that ships, the price floor drops out of the whole market. The labs cutting prices this week know it.

Safety, quickly: Opus 5.5 shows an 85% reduction in containment-boundary circumvention attempts, ties Fable 5.1 for the lowest attack success rate on Gray Swan's prompt-injection test, and — notably — its thinking mode cannot be disabled. Sonnet 5.5 and Haiku 5.5 land in the coming weeks.

Takeaways

  1. Two frontier launches, one afternoon, every price cut: GPT-6 Sol $2/$10 (−50%), GPT-6 Luna $0.10/$0.50, Opus 5.5 $4/$20 (−20% list, −60% on cache reads). Cache cuts are the real weapon — that's where agentic costs live.
  2. Your bill has four lines, and cache reads are now the interesting one: Opus 5.5 cache reads are 5% of input price; Fable 5.1's are 2.5%. Architect your agent loops to re-read, not re-send.
  3. Capability per dollar is improving faster than prices: Opus 5.5 doubles Opus 5's Terminal-Bench-Science score (29% → 58.7%) at a lower list price; Sol halves both errors and price. The frontier is deflationary.
  4. Dollars-per-token lies; dollars-per-task tells the truth: Fable 5.1 is 25% cheaper on typical workloads but ~20% pricier per task at max effort because it burns 1.7× the output tokens. Watch the output token multiplier, not the price list.
  5. Treat vendor benchmarks as vendor-reported until METR and Frontier Design publish. And watch October 15: if Step 5's open weights ship at $1/$2.70, this week's cuts were just the opening bid.

Sources: Reuters (OpenAI GPT-6 launch pricing, Sept 22, 2026); Anthropic announcement (Sept 22, 2026, via PYMNTS/AIWeekly coverage); Gangsta AI (GPT-6 benchmark coverage); Digital Applied (pricing breakdown); Artificial Analysis (Fable 5.1 token-usage analysis; Step 5 index score).


Suggested Medium tags: artificial-intelligence, large-language-models, llm, openai, anthropic

← All posts