OpenAI Slashes GPT-5.6 Luna Price By 80% As China's Cheaper Models Close In
OpenAI just cut the price of its cheapest GPT-5.6 model by 80 percent, three weeks after launch - the clearest signal yet that the company that kicked off the generative AI boom is being dragged into a price war it did not start.
On Thursday the company said GPT-5.6 Luna, its speed-focused model, now costs 20 cents per million input tokens and $1.20 per million output tokens, down from $1 and $6. GPT-5.6 Terra, the mid-tier model, gets a 20 percent trim to $2 per million input tokens and $12 per million output. Pricing for the flagship GPT-5.6 Sol remains unchanged.
"Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost," OpenAI said in its announcement.
The move arrives as companies that once encouraged unconstrained "tokenmaxxing" have started looking hard at AI bills that sometimes run into the billions. Enterprises want clearer returns before committing to the most expensive models, and they now have more alternatives than they did in the early ChatGPT era.
OpenAI is framing the cuts as the fruit of efficiency work rather than a margin sacrifice - a day earlier the company said GPT-5.6 had helped make itself cheaper to run, and it is passing those gains through to how usage is counted in Codex and ChatGPT Work subscriptions, with the lower prices rolling out on AWS as well.
Chinese open-weight models have closed the gap quickly. Moonshot AI's Kimi K3, released earlier this month, has beaten some leading proprietary systems on industry benchmarks and can be run on a company's own infrastructure. That development helped spur a round of competitive responses.
The scale of the challenge is hard to overstate. At 2.8 trillion parameters, K3 is the largest open-weight model ever released, and blind developer testing put it in first place in LMArena's front-end coding arena, ahead of Anthropic's frontier Claude Fable 5. Demand has been heavy enough that Moonshot has capped new subscriptions and API access over capacity constraints, and the company's daily revenue has grown roughly sixfold since launch as it seeks a $50 billion valuation ahead of a potential Hong Kong IPO. The economics underneath are brutal: on Artificial Analysis's cost-per-task index, K3 completes a task for 94 cents and DeepSeek V4 Pro for four cents, versus $1.04 for OpenAI's flagship Sol and $1.80 for Anthropic's Claude Opus 4.8 - and Moonshot reports cache-hit rates above 90 percent in coding workloads that cut K3's effective input cost to 30 cents per million.
Anthropic followed with Claude Opus 5, which it pitches as approaching the performance of its top-tier Fable 5 model at half the price - and beating it outright on some knowledge-work benchmarks - while holding the same rate card as its Opus 4.8 predecessor. Microsoft has been loudly promoting its own cheaper models, including MAI-Cyber-1-Flash, a new cybersecurity-focused offering that AI chief Mustafa Suleyman says delivers "world-leading performance at 50% of the cost." Google, meanwhile, launched a trio of new Gemini Flash models this month and claimed its top Flash model undercuts Kimi K3 and other Chinese systems on a per-task basis.
OpenAI's GPT-5.6 family consists of three tiers: Sol (highest capability), Terra (balanced), and Luna (fastest). By aggressively discounting the middle and low ends while leaving the top model alone, the company is trying to keep volume customers from migrating to open-weight or rival proprietary systems without fully abandoning the premium pricing that funds frontier research.
Whether the cuts are enough to slow the shift toward cheaper alternatives remains to be seen. What is clear is that the era of unconstrained AI spending is giving way to a more pragmatic one - one in which even OpenAI has to compete on price.

