GPT-5.6 Sol vs GPT-5.5: Half the Cost, No Agentic Progress
OpenAI calls GPT-5.6 Sol its most capable model and says it delivers stronger performance per dollar. On FoodTruck Bench, Sol is excellent. It finishes with $53,229 in net worth, stays profitable every day, buys every upgrade without debt, and would enter the leaderboard near the top.
It also loses to GPT-5.5.
Not to GPT-5.5's best result. To its median. Sol finishes 13.3% below it and 7.9% below the lowest GPT-5.5 result in the existing corpus.
This is not a story about a bad model. It is a story about a flagship launch that does not demonstrate progress on a benchmark built for long-horizon agency.
Key findings
- Sol is frontier-tier. It beats Claude Opus 4.6 by 7.5%, GPT-5.2 by 89.6%, and DeepSeek V4 Pro by 96.1% in net worth.
- The immediate predecessor is still ahead. GPT-5.5's median finishes at $61,408, an $8,179 advantage.
- Sol loses the first half. It serves 2,457 customers in the first 15 days against GPT-5.5's 3,628, then recovers strongly in the second half.
- Stockouts are the main regression. Sol records stockouts on 19 days. GPT-5.5's median does so on four.
- The model knew the problem. On Day 3 it wrote that observed sales were capped by stockouts and future orders should follow demand. It repeated the failure for the rest of the month.
- Sol's best behavior is genuinely good. Its event pricing, late-game menu, location choices, and debt-free upgrade path are all frontier quality. The result is a near-miss, not a collapse.
What “no progress” means here
“No progress” is deliberately in quotation marks.
FoodTruck Bench does not measure intelligence in general. It measures whether a model can operate a small business across 30 linked days under uncertainty. A single benchmark cannot invalidate OpenAI's broader GPT-5.6 evaluation suite, where Sol is reported as stronger in coding, computer use, and professional tasks.
The narrower statement is supported:
On this long-horizon economic task, the evaluated GPT-5.6 Sol result does not improve on the established GPT-5.5 distribution.
That wording matters. Sol is clearly better than GPT-5.2 and most of the field. The missing progress is relative to its immediate predecessor, not relative to older OpenAI models.
The current evidence is limited, so the exact gap needs more repetition. But the observed strategy is detailed enough to explain why Sol left money behind.
The leaderboard result
Sol needs two comparison sets at once. GPT-5.5 and Opus show whether it moved the frontier. Grok, DeepSeek, Gemma, and GPT-5.2 show what the same API budget can buy elsewhere.
| Model | Final net worth | API cost | Net worth per $1 | Analyzed tokens |
|---|---|---|---|---|
| GPT-5.5 | $61,408 | $24.63 | $2,493 | 5.03M |
| GPT-5.6 Sol | $53,229 | $13.44 | $3,960 | 4.41M |
| Claude Opus 4.6 | $49,519 | $36.04 | $1,374 | 2.52M |
| Grok 4.5 | $34,586 | $3.98 current-price estimate | $8,686 | 4.64M |
| GPT-5.2 | $28,081 | $7.64 | $3,677 | 4.11M |
| DeepSeek V4 Pro | $27,142 | $3.51 | $7,739 | 5.77M |
| Gemma 4 31B | $24,878 | $0.21 | $119,260 | 1.30M |
Sol generates 59% more net worth per API dollar than GPT-5.5 and almost three times Opus's rate. That is real efficiency progress inside the Western frontier. It is not the best value in the wider field. Grok 4.5 and DeepSeek produce roughly twice Sol's net worth per dollar, while Gemma remains an extreme outlier.
The token comparison clarifies what changed. Both models processed almost the same input, about 3.9 million tokens each, most of it cached. The difference is generated text. Sol produced about 268,000 output tokens, most of that reasoning, against 612,000 for GPT-5.5. It reached 86.7% of GPT-5.5's net worth with less than half as much output.
Sol's margin is slightly higher than GPT-5.5's. Its profit per serving is close too, $7.66 against $7.89. The missing dollars come mostly from throughput. Sol serves 6,901 customers. GPT-5.5 serves 7,694.
The trajectory shows where the gap opened.
Chart key: GPT-5.6 Sol, GPT-5.5 median, Claude Opus 4.6 median.
By Day 10, Sol is $3,790 behind GPT-5.5. By Day 20, it is $10,492 behind. The final ten days narrow the gap, but do not erase it.
Sol built the machine faster than it fed it
The capital plan was strong.
Sol bought all eight upgrades by Day 10 without taking a loan. It reached a full team, expanded cooking capacity, improved storage and marketing, and still finished every day profitable. Those are hard things to do at once. Weaker models often buy upgrades too early and suffocate on fixed costs, or wait so long that they miss the high-demand back half.
Sol's timing was credible. Its inventory plan was not.
The model recorded stockouts on 19 days. On 17 of those days it both stocked out and failed to use 95% of capacity. That combination is worse than a simple capacity limit. It means the kitchen could have cooked more, customers were available, and the truck still could not assemble the menu it had chosen.
GPT-5.5's median recorded stockouts on four days.
| Operating period | GPT-5.6 Sol servings | GPT-5.5 servings |
|---|---|---|
| First 15 days | 2,457 | 3,628 |
| Final 15 days | 4,444 | 4,066 |
Chart key: first series GPT-5.6 Sol, second series GPT-5.5.
Sol's late game was actually stronger on volume. The benchmark had already compounded the early deficit.
It wrote the right rule on Day 3
Sol understood the inventory problem almost immediately:
“Order tomorrow based on demand rather than just prior sales, because observed sales were capped by stockouts.”
GPT-5.6 Sol, Day 3
That is exactly right. Sales history is censored when the truck runs out of food. If 100 customers buy a burrito bowl and the bowl stocks out at noon, ordering for 100 tomorrow does not match demand. It repeats the cap.
Sol then stocked out of burrito bowls on Days 10, 17, 19, 20, and 23 through 29.
This is a familiar pattern in agent benchmarks: a model can generate the correct causal explanation without making that explanation operational. Sol's reasoning was not absent. The rule failed to survive contact with its daily order quantities.
One implementation detail made this worse. Sol appears to have confused cumulative food waste with daily waste. The dashboard's $110.23 figure described waste accumulated across the run. The model treated it as evidence of a recurring daily over-ordering problem and kept orders conservative. In trying to avoid a small loss, it protected itself from the larger profit available through adequate stock.
The custom-recipe detour
Sol created four custom dishes. Together they sold four servings.
The clearest failure came on Day 16, when it replaced the proven menu with a premium custom lineup. The truck served 107 customers, generated $1,129 in revenue, and made $463 in profit at 29% capacity utilization.
On Day 17 it restored the base dishes. The truck served 362 customers and made $2,790 in profit.
Weather and demand conditions are not identical, so this is not a controlled A/B test. The direction is still hard to miss. Sol spent tool calls and menu space on experimental products that had not accumulated popularity while its known products were already inventory-constrained.
This is one place where Kimi K3, a much weaker business operator overall, made the cleaner choice. Kimi did not invent custom recipes. It concentrated on products the market already understood.
Sol's exploration was intellectually plausible and economically mistimed.
Memory went in one direction
Sol made 175 structured memory writes and zero structured retrieval calls.
That does not mean the model had no memory. Notes are also injected into daily context. It does suggest a one-way habit: write a rule, move on, and trust that the surrounding context will surface it later.
On Day 23 Sol wrote:
“Verify inventory and never menu an unavailable item.”
The same class of availability failure continued afterward, including days when lemonade could not be served.
The issue is not that 175 notes were insufficient. It is that a note without an explicit check is not a constraint. Strong long-horizon behavior needs active verification, especially for rules learned from costly mistakes.
Where Sol looked like a flagship
It would be easy to overstate the regression. Sol did several things exceptionally well.
It raised normal main-item prices from roughly $8 to $9.50 early in the month to $12.50 to $14 late. On major events it moved to $14 to $15.50. It selected profitable locations, consolidated its late menu around four core dishes, and never used debt.
Day 22, the Food Truck Rally, was the best example. Sol priced into captive demand, served 366 customers, generated $5,445 in revenue, and made $4,001 in profit. Revenue per serving reached $14.88. Kimi K3 served four more customers on the same seed and made $1,413.
The second half produced $39,409 in profit. Sol was learning.
That is why the headline is not “GPT-5.6 is worse.” Sol is comfortably the second-best model observed on the benchmark by this result. The finding is that its strong recovery was necessary because its first-half execution was below the standard GPT-5.5 had already set.
Price-performance is real
The runtime API cost for this Sol result was $13.44, using OpenAI's published Sol rates of $5 per million input tokens, $0.50 for cached input, and $30 for output.
GPT-5.5's median cost $24.63. Sol delivered 86.7% of its net worth for 54.6% of the API cost.
The gap is almost entirely output. Both models processed a near-identical volume of input, most of it cached, so the input side barely moves the bill. Generated text is where the money is: output tokens cost $30 per million, six times the fresh-input rate and sixty times the cached rate. GPT-5.5 wrote about 612,000 of them. Sol wrote 268,000, less than half. Here is the same result from the run logs, one line per rate component:
| Cost breakdown | Uncached input | Cached input | Output (incl. reasoning) | Total |
|---|---|---|---|---|
| GPT-5.5 | 0.96M × $5 = $4.81 | 2.91M × $0.50 = $1.45 | 0.61M × $30 = $18.37 | $24.63 |
| GPT-5.6 Sol | 0.77M × $5 = $3.84 | 3.16M × $0.50 = $1.58 | 0.27M × $30 = $8.03 | $13.44 |
| GPT-5.6 Luna | 0.61M × $1 = $0.61 | 2.20M × $0.10 = $0.22 | 0.23M × $6 = $1.37 | $2.20 |
GPT-5.5 and Sol share the same rate card, so their comparison is clean: same prices, roughly half the output, roughly half the bill. GPT-5.6 Luna sits on a cheaper tier ($1 input, $0.10 cached, $6 output), so its $2.20 is not a like-for-like figure against the other two. It makes the same point from the other end. In every row the expensive line is output, and the model that generates fewer tokens pays far less, whatever its rate card.
That is meaningful progress per dollar, even though it is not progress in the benchmark's primary score. OpenAI's pricing claim holds up better here than the absolute-capability claim.
There are therefore two defensible headlines:
- Sol did not beat the previous model.
- Sol came close at roughly half the cost.
Both are true. This article focuses on the first because flagship generations are usually judged on capability, not only efficiency.
Verdict
GPT-5.6 Sol is a frontier agent. It built a $53,229 business, finished every day profitable, used no debt, priced major events correctly, and recovered from a weak first half with the strongest late-volume block in the comparison.
It is still behind GPT-5.5.
The gap came from an ordinary operational failure that the model had already explained to itself: inventory should follow demand, not stockout-capped sales. Sol wrote that rule on Day 3 and spent the rest of the month violating it. Nineteen stockout days, an unnecessary custom-recipe detour, and one-way memory writes cost it the throughput needed to match its predecessor.
So where is the progress?
On FoodTruck Bench, it is in price-performance, not in the primary result. Sol produced most of GPT-5.5's value at roughly half the API cost. That is useful. It is not the clean generational capability win implied by a new flagship name.
This conclusion is preliminary. More repetitions could move Sol's stable estimate in either direction. FoodTruck Bench is independent, and API evaluations are paid out of pocket. If you want tighter confidence intervals and faster coverage of new models, support the next evaluation batch on Buy Me a Coffee.
Source and pricing note
Day-level examples use the selected seed-42 GPT-5.6 Sol trace and the published median GPT-5.5 trace. OpenAI's current GPT-5.6 launch page describes Sol as the flagship and lists the current rate card. The local run metadata records $13.44, which agrees with the published Sol rate for the recorded token usage.