OpenAI's 80% price cut, and what it doesn't change
Three weeks after GPT-5.6 launched, OpenAI cut its two cheaper tiers by up to 80%. The official story is efficiency; the list is really converging on open weights. The price was never the moat.
Three weeks after launching GPT-5.6, OpenAI cut the price of its two cheaper tiers. GPT-5.6 Luna dropped 80%, from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output. Terra fell 20%, to $2 and $12. The flagship, Sol, stayed exactly where it was: $5 and $30. In the same week, the premium 'Priority Processing' tier was renamed Fast mode, at twice the price for up to two and a half times the speed. The cheap end got five times cheaper. The expensive end got more expensive.
Why now
The stated reason is efficiency. OpenAI says Sol, running inside its own coding tools, rewrote the company's production GPU kernels and cut serving costs by around a fifth, with further gains from speculative decoding. Sam Altman's line on X: 'we want to offer the best price/intelligence tradeoff at every level.'
The timing tells a fuller story. Two weeks before the cut, Moonshot released Kimi K3 with open weights. GLM-5.2 has been sitting within a point of the frontier on coding benchmarks since mid-June. Both roughly match the capabilities OpenAI charges for, and both can be downloaded. Meanwhile, enterprises are revolting on spend: Uber reportedly burned through its annual AI budget in four months, and OpenAI shipped hard spending caps on its own API eight days before cutting prices.
Already cheaper
Even after an 80% cut, Luna is only level with what open weights already cost. DeepSeek's V4 Flash runs at $0.14 per million input tokens on the official API. GLM-5.2 is $1.40 with output at $4.40. Open-weight models served by third parties, or on your own hardware, sit below all of them. A frontier lab did not drag prices down to a new floor. It conceded that the floor was poured by open models months ago, and its price list is converging on it.
That is what a commodity looks like. When the value migrates out of the token itself, the margin moves to the ends of the list: a budget tier priced to match software anyone can run, and a flagship tier, unchanged, where the premium still lives. Price cuts did not change OpenAI's business model. They admitted where it is.
What the price never bought
We made this argument when GLM-5.2 landed, and again when K3 launched days before the cut: the premium on a closed API was always for convenience, not quality. What remained constant through every reduction is what actually leaves the building. At $6 per million output tokens your prompts went to someone else's computers. At $1.20 they still do. The invoice fell by four fifths and the data path did not move an inch.
A few days ago we put K3 into production on PrivateMind, and in the first few hours it served 3 billion tokens for Options Technology teams at a cost set by hardware we own, on a meter we control. That is the alternative reading of this week's news. API prices will keep sliding toward the cost of serving, because open weights keep releasing at the frontier. When price is no longer the reason to rent, the only question left is the one we started with: who can see your data.