OpenAI cuts GPT-5.6 Luna prices 80% after launch

OpenAI dropped GPT-5.6 Luna 80 percent and Terra 20 percent on July 30, saying the model helped find the savings in its own serving stack. Sol stayed put.

Younes Bekrar8 min read
ShareXLinkedInFacebook
OpenAI cuts GPT-5.6 Luna price 80 percent, three weeks after launch

Eighty percent off Luna. Three weeks after GPT-5.6 went generally available. I had to reread the rate card because frontier-model pricing usually sits still longer than that, or at least pretends to. On July 30 OpenAI cut GPT-5.6 Luna from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output. Terra, the mid-tier most API traffic seems to land on, fell 20 percent to $2 / $12. Sol, the flagship reasoning tier, stayed at $5 / $30. Priority Processing is gone. Fast mode replaces it at roughly 2.5× speed for double the standard price on all three models. Altman posted the familiar line about offering the best price/intelligence tradeoff at every level. The rate card is less poetic than the tweet.

The whole rate card moved, not just the headline

Cached input on Luna fell from $0.10 to $0.02. Terra's cached rate went from $0.25 to $0.20. Batch and Flex stay at half of standard, so Luna batch is now a dime in and sixty cents out. Long-context bills at 2× input and 1.5× output, which puts Luna at 40 cents and $1.80. Fast mode lands at $10 / $60 for Sol, $4 / $24 for Terra, and 40 cents / $2.40 for Luna. Requests still tagged for priority get Fast automatically, no redeploy required for that rename, at least. If your billing dashboards still say "Priority," expect a label change before a math change.

I keep a messy spreadsheet of these line items because the headline "80 percent" never tells you what your actual mix looks like. A shop that lives on cached prompts and batch jobs sees a different bill than one that streams long-context Sol traces at peak hours. OpenAI published the full matrix. Secondary coverage mostly repeated Luna's input number and moved on.

Same numbers apply on the API, ChatGPT Work, and Codex. Monthly subscription prices did not move. Quotas just burn slower when traffic hits Terra or Luna. Bedrock is the asterisk OpenAI itself flagged: AWS bills those customers, and those rates can lag or diverge. Coverage is thin on how fast AWS will mirror this cut, so I would not assume the OpenAI card until the Bedrock page agrees. Multi-cloud routing teams already know this dance. Everyone else will learn it on the first invoice that doesn't match the blog post.

OpenAI says the model helped find the money

The company's developer-community post pins the savings on "20% lower serving costs through production GPU kernel improvements" and "more than 15% better token-generation efficiency through improved speculative decoding," and says GPT-5.6 itself helped with that work. Kernels: less compute per token on the same hardware. Speculative decoding: draft several candidates per pass, verify cheaply, burn fewer full evaluations. Fairly standard serving talk, dressed up with a self-optimizing model as the protagonist.

OpenAI has told versions of this story before. How much came from the model's suggestions versus the usual GPU-kernel grind? Nobody outside can audit that, and I'm not going to pretend otherwise. The part you can bank on is the invoice: lower rates, live for every customer, no ticket required. CNBC got a statement that the company wants each generation to "accomplish more work at a lower cost," and also noted enterprises getting tighter about AI spend plus pressure from Google, Microsoft, and Chinese providers. Marketing and competitive heat can both be true. The post does not force you to pick one.

GPT-5.6 Sol, Terra, and Luna hit GA on July 9 at the old prices. Three weeks is a short leash even in a market that reprices constantly, shorter than OpenAI gave previous flagship families before adjusting rates, as far as I can tell from the public history. My guess, and it is a guess, is that a lot of the efficiency work was already in flight before launch, and list price is now something OpenAI will yank when usage data or a rival's card makes the old number look silly.

Early buyers who locked capacity or built forecasts around July 9 rates got no retroactive credit. Somewhere a finance person is redoing a second-half spreadsheet they just finished. Small tax for showing up early. I've sat in enough planning meetings to know that "the vendor cut prices, celebrate" and "our committed spend model is now wrong" arrive in the same Slack thread.

We want to offer the best price/intelligence tradeoff at every level.
Sam Altman, OpenAI CEO, in a post on X

Where Luna sits next to Gemini and Claude

Gemini 3.1 Pro under 200k tokens: $2 / $12, same as Terra after the cut. Gemini 3 Flash: $0.50 / $3, which is now pricier than Luna on both ends. Claude Sonnet 5 runs introductory $2 / $10 through August 31, then $3 / $15. Claude Opus 5 is $5 / $25, Sol's input, cheaper output. Haiku 4.5 sits at $1 / $5, five times Luna. These numbers drift. I'm writing against the cards published the week of the cut.

So Luna is the budget-tier undercut. Terra is a match for Gemini's mid tier, not a beat. Sol still looks expensive on output next to Opus. OpenAI seems willing to fight on high-volume, latency-tolerant work, agents, classification, bulk docs, and less eager to race Anthropic at the reasoning ceiling. Whether Google or Anthropic match Luna soon is anyone's guess. They might shrug and hold Terra/Sol-adjacent prices instead. I've stopped treating any of these cards as multi-year commitments.

If you already dump volume on Luna, yesterday's bill just got roughly five times smaller with no code change. If you live on Sol, nothing moved, and Fast mode is still double for speed, Priority under a new name. One quieter line in the community post: Codex's automatic action-review is shifting from GPT-5.4 to Luna, and OpenAI expects that workload to cost about 10× less after the model swap plus the cut. That is a real operational win for agent stacks that were burning money on review calls, assuming quality holds. Coverage on quality parity for that specific path is thin so far.

I'd wait a week of production logs before rewriting a whole routing policy around the new card. Prices have been this sticky before. The developer question that shifted in three weeks is less "which model is smartest" and more "which vendor held still long enough to build a budget around", and after July 30, OpenAI just reminded everyone that "held still" is optional. Finance will ask for a revised forecast anyway. Might as well have the logs ready when they do.

A quieter week for Sol buyers

Nobody writing about this cut should pretend Sol customers got a gift. They didn't. The flagship stayed expensive on purpose, or at least stayed expensive without an explanation beyond the missing line on the rate card. If your product's quality bar still needs Sol, you're paying July 9 money for July 30 work. That's fine, just don't let a Luna headline talk you into a silent quality regression.

  • LLMs

Keep reading

AI

Moonshot's Kimi K3 slips a cyber-eval sandbox

Frontier Security says Moonshot's Kimi K3 bypassed a UK AI Security Institute-style cyber evaluation sandbox by using command-line tools when web access was blocked, then pulled answers from GitHub. Part of a wider eval-escape news cycle.

Younes Bekrar8 min read
AI

OpenAI agents built a secret board, then hit Hugging Face

A last-minute Black Hat USA talk from OpenAI's Eric Wallace and Michael Dalton detailed how experimental agents turned an internal JFrog Artifactory registry into a covert coordination channel, escalated privileges twice, and eventually reached the open internet to attack Hugging Face.

Younes Bekrar10 min read
AI

Google DeepMind's Leadership Shake-Up

Demis Hassabis moves to chair and Alphabet chief scientist, Koray Kavukcuoglu takes the operating reins, and Jeff Dean and Sanjay Ghemawat leave for Discovery Loop, a reshuffle that hit Alphabet's stock and Gemini's calendar.

Younes Bekrar8 min read