← Back to Blog
Case StudyJuly 2026Nicholas S.

Grok 4.5: Real Agentic Progress, Still Behind the Frontier

Net worth$34,586
vs Grok 4.3+24.1%
Profit$34,430
Current cost~$3.98

Grok 4.5 is better than Grok 4.3 on almost every business metric that matters.

It finishes FoodTruck Bench with $34,586 in net worth, up 24.1%. It serves 28.4% more customers, increases profit by 20.1%, raises margin by 5.5 percentage points, and cuts food waste by 77.6%. It stays profitable every day.

That is real generational progress.

It is also still 44% below GPT-5.5 and 30% below Claude Opus 4.6. Grok 4.5 found an effective high-volume strategy, then ran into two ceilings: the simulation's physical capacity and its own reluctance to charge more for each scarce serving.

The result is better than the strongest Chinese model currently tested on the benchmark. The economics are mixed. At current published pricing, Grok is competitive with DeepSeek and MiMo, while Qwen still produces far more business value per API dollar.

Key findings

The setup

FoodTruck Bench asks a model to run a food truck for 30 simulated days. The agent starts with $2,000 and controls inventory, menu, pricing, staffing, location, upgrades, suppliers, and financing. Demand changes with weather, events, reputation, price, and prior decisions.

The benchmark rewards final net worth. High volume helps only when the model converts that demand into strong unit economics and avoids waste.

xAI presents Grok 4.5 as its strongest model for coding, agentic tasks, and knowledge work. This is exactly the kind of launch claim FoodTruck Bench can test narrowly: can a model maintain and improve an economic policy across a month of linked tool use?

The current evidence is limited. The comparison below is a deep read of the observed behavior, not a final estimate of the model's stable median.

Where Grok 4.5 sits on the benchmark

Grok 4.5 is no longer a mid-table model. The full context shows both the progress and the remaining gap.

ModelFinal net worthAPI costNet worth per $1Analyzed tokens
GPT-5.5$61,408$24.63$2,4935.03M
GPT-5.6 Sol$53,229$13.44$3,9604.41M
Claude Opus 4.6$49,519$36.04$1,3742.52M
Grok 4.5$34,586$3.98 current-price estimate$8,6864.64M
GPT-5.2$28,081$7.64$3,6774.11M
Grok 4.3$27,880$3.57$7,8064.85M
DeepSeek V4 Pro$27,142$3.51$7,7395.77M
Gemma 4 31B$24,878$0.21$119,2601.30M
MiMo V2.5 Pro$22,388$2.41$9,2903.82M
Qwen 3.6 35B-A3B$15,317$0.68$22,6823.42M
Grok 4.5 clears the value tierFinal net worth: above GPT-5.2 and the tested Chinese frontier, below the top three Western models.

At the corrected current-price estimate, Grok 4.5 produces about $8,686 of simulated net worth per API dollar. That is 11% better than Grok 4.3, 12% better than DeepSeek, and only 7% below MiMo. Qwen remains 2.6 times more efficient. Gemma remains an outlier that no premium model approaches.

The workload itself barely grew relative to Grok 4.3. Grok 4.5 processed about 4.64 million analyzed tokens against 4.85 million for 4.3. The newer model used roughly half as many reasoning tokens and still produced 24% more net worth. That is a cleaner generational improvement than the raw price comparison suggested.

Grok 4.5 versus Grok 4.3

MetricGrok 4.3Grok 4.5Change
Net worth$27,880$34,586+24.1%
Revenue$61,467$66,065+7.5%
Profit$28,661$34,430+20.1%
Margin46.6%52.1%+5.5 pp
Servings7,1949,234+28.4%
Food waste$2,908$652-77.6%
Revenue per serving$8.54$7.15-16.3%
Profit per serving$3.98$3.73-6.3%
Grok 4.5 progress versus Grok 4.3Grok 4.3 indexed to 100.
Grok 4.3Grok 4.5

Grok 4.5 improved by doing more. It pushed the truck harder, served 2,040 additional customers, and controlled spoilage far better.

The regression in revenue and profit per serving explains why a 28% volume increase became only a 24% net-worth increase. Grok improved the machine without improving the value extracted from each unit of scarce capacity.

A remarkably clean operating curve

The net-worth trajectory is unusually steady:

Grok 4.5 pulled away from 4.3 after Day 15Net worth at five-day checkpoints.
Grok 4.5Grok 4.3DeepSeekSol

Grok finished every day profitable. It took no loans, bought all upgrades by Day 10, built a full team, and used a stable eight-item menu. It negotiated with the Farmers supplier 30 times and accepted 25 deals.

The strategy was not elegant. It was effective:

  1. Keep a broad menu.
  2. Buy capacity early.
  3. Accumulate popularity across established products.
  4. Move to major events and strong locations.
  5. Fill the kitchen almost every day.

Grok avoided one of the most common agent traps in the benchmark: custom-recipe theater. It did not spend the month inventing products nobody knew. It sold proven dishes and increased throughput.

That is a boring strategy. Boring is underrated in long-horizon agents.

It diagnosed the bottleneck and actually changed the business

On Day 9 Grok wrote:

“Demand is there. Execution (inventory + capacity) is bottleneck.”

On Day 16:

“CAPACITY IS THE #1 BOTTLENECK”

Unlike Luna and Kimi K3, Grok's actions broadly matched its diagnosis. It bought the upgrades and then used them. The truck reached at least 95% capacity utilization on 21 days. Luna bought the same upgrade tree and reached that threshold on three.

By Day 30 Grok's note was accurate:

“Capacity is hard limit.”

This is one of the strongest agentic signals in the result. The model did not merely write a sensible postmortem. It converted the postmortem into capital expenditure, staffing, supplier work, and daily volume.

The remaining problem was what to do after the hard limit became real.

The volume ceiling

Across the month, Grok generated demand from 25,840 customers and served 9,234. It left 16,606 unserved.

Late in the run the pattern became extreme:

DayDemandServedUnserved
222,5323992,133
272,5034002,103
281,8044001,404
The hard capacity ceilingLate-run demand split into served and unserved customers.
ServedUnserved

On Day 22, the Food Truck Rally, Grok generated $2,894 in revenue and $1,521 in profit from 399 servings. The model wrote that “events are gold mines if we can scale.”

But the truck could not scale further. It had reached the simulation cap.

At that point the remaining lever was price. Grok's average main-item price across the month was $9.05 and $9.33 late in the run. GPT-5.6 Sol charged $14 to $15.50 at the same Day 22 event and made $4,001 in profit from 366 servings.

Day 22Grok 4.5GPT-5.6 Sol
Servings399366
Revenue$2,894$5,445
Profit$1,521$4,001
Revenue per serving$7.25$14.88
The missing pricing moveRevenue and profit per serving for nearby strategies.
Revenue / servingProfit / serving

Grok served 9% more customers. Sol made 163% more profit.

The model solved the capacity problem as far as the environment allowed. It did not fully make the strategic transition from scaling volume to pricing scarcity.

The eight-item compromise

Grok's broad menu helped demand and product popularity. It also complicated inventory.

Stockouts occurred on 16 days. The model browsed supplier catalogs 56 times and negotiated daily, but still accumulated $652 in food waste. On Day 30 it summarized the issue plainly:

“Waste is killer.”

The waste was much lower than Grok 4.3's $2,908, so this is a success relative to the previous generation. It remained a material loss and a signal of menu complexity.

A broad menu requires more ingredient types, more accurate recipe-level forecasting, and more ways for one missing component to block a product. Grok handled that complexity better than most agents, but the work was expensive in tool use: 1,046 total calls, including 498 information calls, 180 actions, and 311 memory calls.

This was not an agent that failed because it did too little analysis. It worked constantly. The ceiling was the efficiency of the resulting policy.

Above the Chinese field, with mixed economics

By absolute result, Grok 4.5 is stronger than the currently tested Chinese frontier on FoodTruck Bench.

ModelNet worthCurrent or recorded API cost
Grok 4.5$34,586$3.98 current-price estimate
DeepSeek V4 Pro$27,142$3.51
MiMo V2.5 Pro$22,388$2.41
Qwen 3.6 35B-A3B$15,317$0.68

Grok is 27.4% above DeepSeek, 54.5% above MiMo, and 125.8% above Qwen in final net worth.

Under the corrected current-price estimate, Grok's cost is also far more competitive than the raw local tracker implied. The workload cost is in the same broad range as DeepSeek, not four times higher.

Using final net worth divided by API cost, Grok produces about $8,686 per API dollar. DeepSeek produces about $7,739 and MiMo about $9,289. Those three are in the same broad efficiency band.

Qwen remains the striking exception. It produces less than half Grok's final net worth, but its API cost is about one-sixth. That works out to roughly $22,683 in net worth per API dollar, 2.6 times Grok's rate. Small open models can be worse in absolute outcome and still be the rational choice when repetition, throughput, or parallelism matters.

The correct conclusion is not “Grok only matches Chinese models.” It beats them on this result.

The correct conclusion is:

Grok 4.5 has crossed above the tested Chinese frontier in absolute business performance. Its current-price efficiency is competitive with DeepSeek and MiMo, but Qwen remains in another cost class.

Still below the Western frontier

ModelNet worthDifference vs Grok 4.5
GPT-5.5$61,408+77.6%
GPT-5.6 Sol$53,229+53.9%
Claude Opus 4.6$49,519+43.2%
Grok 4.5$34,586baseline
GPT-5.2$28,081-18.8%
DeepSeek V4 Pro$27,142-21.5%

Grok is no longer mid-pack. It clears GPT-5.2 and the tested Chinese models by a meaningful margin. It is also not close to the top three.

The remaining gap is not mysterious. Frontier leaders make much more money per serving. They price constrained demand aggressively, narrow the menu when necessary, and turn scarce kitchen slots into margin. Grok's high-volume policy pushed against the physical cap without completing that last economic step.

Verdict

Grok 4.5 made real progress.

It improved net worth by 24%, profit by 20%, margin by 5.5 percentage points, volume by 28%, and waste by 78% relative to Grok 4.3. It bought capacity early, used it, stayed profitable every day, and ran one of the cleanest operational curves in the benchmark.

The model also showed the difference between diagnosis and adaptation. When Grok identified capacity as the bottleneck, it expanded capacity and filled it. Kimi and Luna often wrote the same diagnosis without completing the action loop.

Then Grok reached the hard cap and stopped one move short. With thousands of customers still waiting, it kept average main prices near $9. Sol charged roughly $15 at the rally and made more than twice the revenue from fewer servings.

So the progress is real, and insufficient.

Grok 4.5 is above GPT-5.2 and the tested Chinese frontier in absolute result. It remains well below GPT-5.5, Sol, and Opus. At current pricing it is much more economically attractive than our original raw tracker suggested. Its efficiency is competitive with DeepSeek and MiMo, while Qwen still offers far cheaper repetition.

We need more evidence before calling the exact rank stable. FoodTruck Bench is independent, and API evaluations are paid out of pocket. If you want more repetitions, more launch-day models, and a tighter estimate of where Grok actually lands, support the next evaluation batch on Buy Me a Coffee.


Source and pricing note

The run metadata stored $14.08 because the local client matched grok-4.5 to an older generic grok-4 price. Using xAI's published short-context Grok 4.5 rates of $2 per million uncached input tokens, $0.30 per million cached input tokens, and $6 per million output tokens, the recorded workload reprices to approximately $3.98. Verify the current xAI pricing documentation immediately before publication.