← Back to Blog
Case StudyJuly 2026Nicholas S.

GPT-5.6 Luna: A Cheap Agent Outclassed by Chinese Models

Net worth$13,540
ROI+577%
Margin37.9%
Current cost~$2.20

OpenAI positions GPT-5.6 Luna as the fastest and most affordable member of the new family. The obvious promise is not that a small model will beat Sol. It is that better intelligence per dollar will make capable agents cheap enough to run everywhere.

On FoodTruck Bench, Luna built a profitable business. It finished with $13,540 in net worth, $13,536 in profit, and a current-price workload cost of approximately $2.20.

It also finished below Qwen 3.6 35B, MiMo V2.5 Pro, and DeepSeek V4 Pro. The problem was not raw speed or tool access. Luna repeatedly identified inventory as the bottleneck, bought every capacity upgrade anyway, and left 8,593 customers unserved while the kitchen often sat below capacity.

The surprising part is not that a small model lost to a flagship. It is that Luna could find a much stronger strategy and did not do so reliably.

Key findings

The setup

FoodTruck Bench gives a model $2,000, a truck, and 30 simulated days in Austin. The model controls menu, prices, purchasing, staffing, location, upgrades, and financing. Ingredients perish. Weather and events shift demand. Every decision changes the state inherited by the next day.

The primary metric is final net worth. This rewards an agent that can connect diagnosis to action across a long horizon. A model does not get credit for writing the correct postmortem if tomorrow's order repeats yesterday's mistake.

OpenAI describes GPT-5.6 Luna as the family's fastest and most affordable option, priced at $1 per million input tokens, $0.10 for cached input, and $6 for output. The analyzed workload comes to approximately $2.20 at that rate.

The present evidence base is limited. The goal of this article is to examine the observed policy in detail, not to pretend we have a narrow confidence interval.

Where Luna sits on the benchmark

Luna's relevant comparison is not only Sol. It is the cluster of models that can plausibly be used as inexpensive agents today.

ModelFinal net worthAPI costNet worth per $1Analyzed tokens
GPT-5.6 Sol$53,229$13.44$3,9604.41M
Grok 4.5$34,586$3.98 current-price estimate$8,6864.64M
DeepSeek V4 Pro$27,142$3.51$7,7395.77M
Gemma 4 31B$24,878$0.21$119,2601.30M
Kimi K3$22,404$7.15$3,1323.31M
MiMo V2.5 Pro$22,388$2.41$9,2903.82M
Qwen 3.6 35B-A3B$15,317$0.68$22,6823.42M
GPT-5.6 Luna$13,540$2.20 current-price estimate$6,1663.23M
Luna enters below the affordable fieldFinal net worth against nearby small, Chinese, and frontier models.

At current pricing, Luna produces about $6,166 of simulated net worth per API dollar. That is better than Sol, Kimi, GPT-5.5, and Opus. It is still 34% below MiMo, 73% below Qwen, and 95% below Gemma on the same efficiency measure.

The token workload is equally important. Luna processed about 3.23 million input, output, and reasoning tokens. Qwen used 3.42 million and MiMo 3.82 million. Luna was not dramatically smaller in workload. It simply converted a similar amount of inference into less business value.

A business that stalled in the middle

Luna's net-worth curve has three phases.

It starts well, reaching $3,543 by Day 5. It then spends most of the month on a shallow plateau. Net worth is $4,278 on Day 10, $5,912 on Day 15, $6,781 on Day 20, and $8,003 on Day 25. Only in the final five days does the curve steepen, ending at $13,540.

A long plateau, then a late recoveryLuna net worth through Day 30.

The late acceleration matters. Luna was not incapable of running the business. It eventually found a profitable rhythm. The earlier decisions gave that rhythm too little time to compound.

Seven days lost money. Day 21 was the worst location error: the truck went to the industrial zone on a Sunday, generated $87 in revenue, and lost $584. Day 17 produced only eight servings because inventory made almost the entire menu unavailable. Day 13 had 915 customers of demand and only 42 servings.

Those were not random demand failures. They were days when the model's operational state made available demand impossible to capture.

Luna knew

The most important quotes are not confused. They are correct.

On Day 3 Luna wrote that “ingredient stockouts” had “limited sales.” On Day 17 it recorded that “the bottleneck was inventory/stockouts.” On Day 25 it called the truck “massively underutilized.”

On Day 30 it finally summarized the accumulated waste:

“FOOD WASTE was $613.66... major cash leak.”

Luna was not blind. It saw the bottleneck almost word for word.

The failure lived between analysis and the next action.

The run ended with:

Under-ordering and high waste may sound contradictory. They are not. Luna often ordered too little of the ingredients that completed popular recipes while carrying too much of other perishable stock. Inventory accuracy is about mix as much as volume.

That is precisely the kind of problem an agent should solve by joining sales history, recipe requirements, and stockout data. Luna described each piece but did not stabilize the combined policy.

Eight upgrades, three full days

Luna spent $5,150 on all eight available upgrades by Day 20.

Capacity and infrastructure upgrades are valuable when the truck is physically unable to meet demand. Luna's daily notes repeatedly said the opposite: ingredients were the binding constraint and the truck was underused.

Only three days reached at least 95% capacity utilization.

Infrastructure bought versus infrastructure usedBoth models bought every upgrade. Only Grok consistently filled the truck.
LunaGrok 4.5

This is a capital-allocation error, not merely an inventory error. Cash that could have financed a larger and better-balanced order went into a kitchen that often lacked the components required to cook.

The comparison with Grok 4.5 is useful. Grok also bought every upgrade. It then operated at or above 95% capacity on 21 days. For Grok, capacity investment served an observed constraint. For Luna, it often served a generic plan.

The distinction is simple:

A good upgrade is not an upgrade with positive theoretical value. It is the upgrade that removes the bottleneck you actually have.

Luna followed a familiar “buy the tech tree” sequence without proving that the next node was limiting the business.

The operational gapStockout days and loss days in the selected result.

The $10 ceiling

Luna's menu prices barely moved.

For most of the month, burgers and chicken tacos sat at $10, fries at $5, lemonade at $4, and soda at $3. The custom Austin Cheese Burger also sold for $10. A new recipe became another item, not a new price tier.

Day 22 was a six-times-demand Food Truck Rally. Demand reached 1,609 customers. Luna served 306, turned away 1,303, and generated $2,087 in revenue and $1,061 in profit. Revenue per serving was $6.82.

On the same seed, GPT-5.6 Sol served 366 customers at higher prices. It generated $5,445 in revenue and $4,001 in profit.

Day 22 Food Truck RallyGPT-5.6 LunaGPT-5.6 Sol
Demand1,609429
Servings306366
Unserved demand1,30363
Revenue$2,087$5,445
Profit$1,061$4,001
Revenue per serving$6.82$14.88

The demand rows should not be read as a direct popularity comparison because each model's prior choices alter demand. The unit economics are comparable. Sol earned 2.18 times more revenue per serving.

Luna carried a low-price default into an environment where demand repeatedly overwhelmed available supply. Like Kimi K3, it preserved affordable prices for customers it could not serve.

Luna's strongest evidence is another Luna

The most interesting comparator is not Sol or a Chinese model. It is Luna itself.

In another observed full trajectory, Luna finished with $28,868 in net worth, more than twice the selected result. It began with seven dishes instead of four, used prices up to $11.99 immediately, later moved to $12, and earned $8.13 per serving instead of $6.93. Margin reached 51.1%. Waste was $106 instead of $614.

Observed Luna policySelected resultStronger result
Final net worth$13,540$28,868
Revenue per serving$6.93$8.13
Margin37.9%51.1%
Food waste$614$106
Opening menu4 simple items7 items
Early top price$10$11.99

This changes the diagnosis.

Luna's ceiling is not obviously $13,540. The model can discover a richer menu, stronger prices, and much cleaner inventory economics. What it does not yet do reliably is converge on that policy.

For an inexpensive agent, variance is part of the price. A model that costs less per attempt but needs more attempts to produce the strong strategy may not be as cheap as its token rate suggests. More repetition is required before that can be quantified.

The cheap-model comparison

The initial temptation was to write that Chinese models are an order of magnitude cheaper and better. The data supports the “better” part for the relevant examples. It does not support the price claim as a blanket statement at Luna's current rate.

ModelNet worthAPI costNet worth vs Luna
DeepSeek V4 Pro$27,142$3.51+100%
MiMo V2.5 Pro$22,388$2.41+65%
Qwen 3.6 35B-A3B$15,317$0.68+13%
GPT-5.6 Luna$13,540$2.20 current-price estimatebaseline

Qwen is about 3.25 times cheaper and finishes 13% higher. MiMo costs roughly 10% more and finishes 65% higher. DeepSeek costs about 60% more and doubles Luna's result.

Gemma 4 31B, which is not a Chinese model, makes the broader efficiency point more dramatically: about $24,878 in net worth at a recorded cost near $0.21. That is more than ten times cheaper than the current-price Luna estimate and 84% better in result.

The honest conclusion is sharper than a slogan:

Luna's new price is competitive. Its observed business policy is not.

Progress inside the small OpenAI line

Calling Luna simply “a weak small model” would erase real progress. It is substantially more capable than earlier mini-class OpenAI models on this benchmark. It completes the month, produces a meaningful profit, uses the full tool environment, and sometimes finds a strategy that approaches the top non-frontier group.

The disappointment comes from the 2026 comparison set.

Cheap agentic models no longer compete only with the previous mini generation. They compete with Qwen, MiMo, DeepSeek, and efficient open-weight models that have already demonstrated stronger economic policies. Against that field, “better than the old mini” is not enough.

Luna's current rate card solves much of the historical price problem. It does not solve policy selection.

Verdict

GPT-5.6 Luna can run the business. It cannot yet be trusted to find its better business consistently.

The selected result spent $5,150 on capacity while inventory remained the bottleneck, reached full utilization on only three days, stocked out on 23 days, and kept a $10 pricing ceiling through repeated demand spikes. Luna documented these problems in unusually clear language. Then it repeated them.

Another Luna trajectory did more than twice as well. That is encouraging and damning at the same time. The model contains a strong policy, but the current evidence does not show reliable access to it.

At approximately $2.20 for this workload under current pricing, Luna is not outrageously expensive. Qwen remains much cheaper and slightly better. MiMo and DeepSeek cost a similar order of magnitude and produce much stronger results. Gemma shows that order-of-magnitude cost advantages are possible outside the Chinese field too.

We expected more because the affordable frontier is already crowded.

More evidence would make this conclusion much stronger. FoodTruck Bench is independent, and API evaluations are paid out of pocket. More repetitions mean a real distribution instead of a behavioral snapshot. If you want more repetitions, more models, and fewer launch-week conclusions from limited evidence, support the next evaluation batch on Buy Me a Coffee.


Source and pricing note

The selected result is the median observed Luna trajectory by final net worth. Public copy deliberately avoids presenting a non-canonical repetition count. The stored run metadata contains a stale $10.98 cost produced by an incorrect Sol-tier mapping. Repricing the recorded token usage at OpenAI's published Luna rate gives approximately $2.20. Recheck the official GPT-5.6 pricing immediately before publication.