The most important number in AI this week is $1.20.

OpenAI cut the API price of GPT-5.6 Luna by 80%, bringing it down to $0.20 per million input tokens and $1.20 per million output tokens. Three weeks earlier, Luna launched at $1 and $6. OpenAI says improvements across the systems serving the model made the reduction possible.

The significance of that cut became immediately clear in a post from Nic Dunz, whose observation inspired this analysis:

“GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price.”

The math is almost comically tidy:

$2.50 ÷ $0.20 = 12.5

$15 ÷ $1.20 = 12.5

GPT-5.4 arrived in March at $2.50 per million input tokens and $15 per million output tokens. Approximately four months later, OpenAI is offering a comparable level of measured intelligence for 8% of that token price.

That is a 92% reduction.

There are important caveats. GPT-5.4 xhigh and GPT-5.6 Luna max are different models running at different reasoning settings. A shared benchmark score does not mean they will perform identically on every task. Models may also consume very different quantities of reasoning tokens before producing an answer.

Still, the comparison is unusually strong. Artificial Analysis currently gives GPT-5.4 xhigh an Intelligence Index score of 51. Luna xhigh scores 49, while Luna max reaches 51. Luna also produces output faster than GPT-5.4 in Artificial Analysis testing, although its maximum setting can take considerably longer to begin responding.

A level of intelligence that commanded flagship pricing in March has become radically cheaper by July.

That may prove more economically important than the next few points gained at the frontier.

Intelligence is becoming a rapidly depreciating asset

AI progress is usually presented as a race upward.

A new model arrives. Its benchmark score is higher. It can solve harder coding problems, perform longer agentic tasks or reason across more complicated material. The industry then debates whether the advance is meaningful enough to crown a new leader.

That tells us how far the frontier has moved, but less about what happens immediately behind it.

Once an AI lab demonstrates a useful capability, that capability rarely retains its original price for long. Smaller models catch up. Hardware becomes more efficient. Inference systems improve. Providers develop better routing, caching and context-management techniques. Competition forces prices downward.

The capability remains valuable.

Its scarcity begins to disappear.

The historical comparisons are striking:

Model

Release

AA Intelligence Index

Input price

Output price

GPT-4o

May 2024

9

$5.00

$15.00

OpenAI o1

December 2024

23

$15.00

$60.00

GPT-5.4 xhigh

March 2026

51

$2.50

$15.00

GPT-5.6 Luna max

July 2026

51

$0.20

$1.20

Artificial Analysis’s index has evolved over time, so this table should not be interpreted as a perfectly controlled longitudinal experiment. The underlying evaluation suite changes as older benchmarks become saturated and newer tests are introduced.

Even with that limitation, the direction is hard to miss.

In May 2024, GPT-4o cost $5 per million input tokens and $15 per million output tokens while scoring 9 on Artificial Analysis’s current index. Seven months later, OpenAI’s o1 reached 23, but charged $15 and $60. Luna now reaches the low 50s at $0.20 and $1.20.

The industry is delivering far more measured capability at dramatically lower prices.

Epoch AI has studied this pattern across multiple benchmarks by identifying the cheapest model able to reach a fixed performance threshold at different points in time. It found inference prices declining between 9-fold and 900-fold per year, depending on the benchmark and performance level, with a median decline of roughly 50-fold among the trends it studied.

Epoch also cautioned that the fastest declines were recent and may not continue indefinitely. But its central finding remains remarkable: the price of achieving a previously established level of AI performance is falling extremely quickly.

Luna gives us a fresh, clean example from within one model provider and across only four months.

Yesterday’s flagship intelligence is becoming today’s affordable infrastructure.

Subscribe to Staticbreaker for defensible AI interpretation for busy, intelligent non-specialists: https://thestaticbreaker.com/subscribe

The frontier and the floor are moving simultaneously

There are two AI progress curves worth watching.

The first asks:

How intelligent is the best model available?

The second asks:

How cheaply can we purchase a level of intelligence that already exists?

The first curve determines what becomes possible next. The second determines when those possibilities become economical.

A frontier model may establish that AI can review a complicated contract, debug a large codebase or synthesize hundreds of documents. But if each run is expensive, the capability will initially be reserved for high-value work.

Once the same performance becomes 10 or 12 times cheaper, the calculation changes.

The model no longer needs to replace an hour of senior professional labor to justify its cost. Perhaps it only needs to save ten minutes.

Then five.

Eventually, it may be worth invoking for work humans currently skip because the benefit is too small to warrant anyone’s attention.

Most businesses do not need the most intelligent model available for every task. They need a model that reliably clears a threshold:

  • Good enough to classify a customer message

  • Good enough to review a support interaction

  • Good enough to spot an inconsistency in a spreadsheet

  • Good enough to research a prospect

  • Good enough to verify a routine code change

  • Good enough to identify when a human should look more closely

When the price of crossing that threshold drops by 92%, the same projects become cheaper.

More importantly, entirely new projects become rational.

Cheaper intelligence is likely to create more demand

An 80% price cut sounds like a direct hit to revenue.

At one-thirteenth the original token price, OpenAI would need roughly 12.5 times as much Luna usage to generate the same revenue as an equivalent volume of GPT-5.4 tokens.

That kind of demand response may be entirely plausible.

Machine intelligence could have an unusually elastic demand curve because the quantity of possible cognitive work is nearly limitless.

There is always another document to review.

Another customer interaction to personalize.

Another result to verify.

Another simulation to run.

Another possibility to consider.

Another agent action to inspect.

Humans currently leave an enormous amount of potentially useful thinking undone because labor and attention are scarce.

Companies analyze a sample of sales calls instead of every call. They create custom research only for the largest accounts. They run one draft instead of five. They skip second-pass quality checks on routine work. They tolerate generic customer communication because personalization costs too much.

Cheap AI can absorb that neglected work.

This creates two demand effects:

  1. More workflows become economical.

  2. Each workflow can afford to use more intelligence.

A support organization that analyzed 1% of its conversations may begin reviewing all of them.

A sales team that created research briefs only for enterprise prospects may generate one for every qualified lead.

A software agent that previously received one chance to solve a problem may be allowed to retry, inspect the failure, revise its plan and ask another model to critique the answer.

A legal team may run a preliminary AI review over every low-value contract rather than reserving analysis for major agreements.

A healthcare organization might add a machine-generated second check to routine instructions, medication lists or administrative decisions.

None of these applications needs to transform the organization by itself.

Each one only needs to create slightly more value than it costs.

AI could become dramatically cheaper per task while the world spends dramatically more on it overall.

The long tail of cognition

The first wave of generative AI naturally concentrated on expensive human work.

Software development. Legal analysis. Financial research. Medical reasoning. Marketing. Design.

When model usage was expensive and unreliable, companies targeted activities where even a modest productivity gain could justify the cost.

The next wave may spread down the value curve.

Imagine an AI check that produces five cents of expected value every time it runs. At ten cents per run, the project makes no economic sense. At one cent, it becomes profitable. Across ten million transactions, the tiny gain becomes a meaningful business.

This is where cheaper intelligence may spread most aggressively:

  • Minor quality checks

  • Small customer accounts

  • Routine compliance reviews

  • Internal documentation

  • Product-description cleanup

  • Meeting preparation

  • Background monitoring

  • Low-risk anomaly detection

  • Personalized explanations

  • Repeated simulations

  • Second opinions on ordinary work

These activities were rarely impossible. They were too unimportant to justify paying a person or an expensive model to do them. Falling prices change that boundary.

A project does not need to create millions of dollars in value. It only needs to produce a positive return at sufficient scale.

The frontier model earns the headlines. The long tail generates the volume.

Agents multiply the effect

The move from chatbots to agents makes model pricing even more consequential. A chatbot may receive one request and produce one answer.

An agent may:

  1. Interpret the objective

  2. Create a plan

  3. Retrieve information

  4. Call several tools

  5. Examine the results

  6. Revise its approach

  7. Retry failed steps

  8. Ask another model for criticism

  9. Verify the completed work

  10. Summarize what happened

One user request can produce dozens or hundreds of model calls.

At high prices, developers constrain those loops. They shorten context, cap retries, skip independent verification and route as much work as possible to weaker models.

As intelligence becomes cheaper, those restrictions can loosen. Agents can consider more alternatives. They can check their work more thoroughly. They can recover from errors rather than giving up. They can use stronger models at stages where those models were previously too expensive.

Cheaper intelligence therefore improves agent economics from both directions. More agentic projects become viable, and each agent can afford to perform more cognitive work.

That may also improve reliability. Many of today’s apparent model limitations are partly budget limitations. The system could consult another model, perform another search or verify another claim, but the workflow was designed to avoid the extra expense.

When checking becomes cheap enough, checking can become standard.

Token price is only part of the bill

The one-thirteenth comparison is powerful, but sticker price does not tell the whole story.

A reasoning model can consume thousands of internal tokens before producing a short visible answer. A model priced at $1.20 per million output tokens may still cost more to complete a task if it uses dramatically more tokens than a seemingly expensive alternative.

Tool calls, web searches, retrieval, code execution and failed agent attempts can add to the total.

Artificial Analysis addresses this by reporting both token pricing and cost per Intelligence Index task. Its task-cost calculation incorporates the input, cache, reasoning and answer tokens actually consumed during evaluation. This is more useful than assuming every model uses the same number of tokens to reach its result.

Luna’s new pricing had only just been announced when this article was written, so some third-party blended-price and task-cost pages may still reflect its previous $1 and $6 rates.

The cleanest comparison will eventually include:

  • Sticker input price

  • Sticker output price

  • Blended token price

  • Actual cost per benchmark task

  • Reasoning-token consumption

  • Latency

  • Output speed

  • Success rate on real workflows

The new price does not guarantee every Luna task will be 92% cheaper than the same work performed by GPT-5.4.

It means the raw token price for comparable measured intelligence has fallen 92%.

Real-world savings will depend on how efficiently Luna uses those tokens.

This is larger than OpenAI

Luna provides the cleanest current example because the comparison stays within OpenAI.

The underlying pressure is industry-wide.

Open-weight models and lower-cost Chinese providers continue pushing capable intelligence into cheaper pricing tiers. Hyperscalers are building specialized accelerators. Model companies are improving inference efficiency, caching and routing. Developers increasingly send routine work to smaller models and escalate only difficult cases to expensive flagships.

OpenAI says GPT-5.6 was designed to extract more useful work from each token and deliver comparable results at lower total cost. At launch, Luna was positioned as the company’s fastest and most affordable GPT-5.6 model. Three weeks later, OpenAI cut its price another 80%.

That speed matters.

Historically, major model price reductions often followed months of optimization or the release of a new generation. Luna’s cut arrived weeks after launch.

The economic shelf life of premium model pricing may be shortening.

That will pressure AI labs to keep advancing the frontier because the intelligence immediately behind it is being commoditized at an extraordinary rate.

It may also explain why major labs are expanding beyond selling tokens. Subscriptions, agents, enterprise software, applications, distribution, identity and infrastructure become more valuable when the underlying intelligence layer faces persistent price compression.

What this means for businesses

AI project economics now come with an expiration date. A workflow rejected in March may deserve another look in July.

A use case with poor margins at $15 per million output tokens may work comfortably at $1.20, especially when it runs at high volume and creates only a small amount of value per task.

Businesses should regularly recalculate projects that were dismissed because of model expense. They should also reconsider how much verification they can afford.

Many deployments are currently optimized to minimize calls. They generate one draft, one classification or one agent attempt because each additional pass costs money.

As prices fall, the optimal workflow changes. The model can draft an answer. A second model can criticize it. The first can revise. A final check can confirm that the output satisfies the original request.

Some of the savings can be captured as lower costs. Some may be more valuable when reinvested in reliability.

What this means for AI startups

Cheaper intelligence creates opportunity and danger in nearly equal measure.

Products that could not maintain healthy margins under previous API prices may suddenly become viable. Startups can support smaller customers, include more generous usage and operate agents that would have been prohibitively expensive only months earlier.

But a cost advantage is a fragile moat.

A business built around access to a cheaper model can lose that advantage overnight when a major provider changes a pricing page.

Durable value has to come from somewhere else:

  • Proprietary data

  • Workflow integration

  • Distribution

  • Customer relationships

  • Domain expertise

  • Trust

  • Evaluation systems

  • Accumulated context

  • Reliable outcomes

As generic intelligence becomes cheaper, the system surrounding it becomes the product.

What this means for workers

The economic threshold for automation is moving.

The question is whether AI can perform enough of that task, cheaply and reliably enough, to justify reorganizing the workflow.

A modest price reduction may not change the answer.

A 92% reduction can.

Work that seemed relatively insulated because automation produced too little value may become more exposed without another spectacular leap in capability. The model does not necessarily need to become much smarter if existing intelligence becomes cheap enough to deploy everywhere.

That does not mean entire jobs disappear in one sweep.

It means more tasks become worth augmenting, monitoring, checking or partially automating.

The labor impact of AI may arrive through falling costs as much as rising benchmark scores.

The infrastructure paradox

There is a final irony. Cheaper intelligence may intensify the demand for compute.

Inference becomes more efficient. Providers lower prices. Developers respond by processing more data, running longer agent loops and applying models to increasingly marginal tasks.

Aggregate usage rises.

The industry can therefore experience two trends at once:

  • The price of each unit of useful intelligence falls.

  • Total spending on AI infrastructure climbs.

There is no contradiction. The world is paying less per unit and consuming far more units.

This is how cloud computing evolved. Cheaper storage and processing did not lead companies to store less data or run fewer applications. They helped create entire industries that would have been impossible under the previous cost structure.

AI may follow the same path, except cognitive work has fewer obvious limits.

Every answer can be verified.

Every customer can be personalized for.

Every workflow can be monitored.

Every decision can be accompanied by alternatives.

Every agent can be given another chance to get the job right.

Every decline in the price of intelligence invites another layer of demand.

The most important AI curve

The frontier will continue to dominate attention.

The smartest model matters. It establishes what becomes possible next and creates capabilities that cheaper systems will eventually inherit.

But most of the economy will be transformed behind the frontier. It will happen when a rare capability becomes common. When an expensive capability becomes affordable. When a task that never justified an hour of human attention begins to justify a fraction of a model call.

GPT-5.6 Luna is one data point. Its real-world savings will vary according to reasoning-token use, task length, caching and workflow design.

But the direction is clear.

A level of measured intelligence that cost $15 per million output tokens in March is now being offered for $1.20.

The price of AI is collapsing. Demand may be about to explode.

And some of the greatest economic value may come from all the thinking that was never worth paying for until now.

Reply

Avatar

or to participate

Keep Reading