Kimi K3: Great at Coding, Weaker at Agentic Tasks
Moonshot built Kimi K3 for long sessions. The launch material emphasizes repository navigation, terminal use, compiler work, and multi-hour tasks. In other words, K3 is supposed to be comfortable when the answer is not one response but a chain of decisions.
FoodTruck Bench is also a chain of decisions. The tools are simpler than a terminal, but the feedback loop is longer. The model has to price a menu, order perishable inventory, staff a truck, react to events, and remember what actually worked for 30 simulated days.
Kimi K3 completed the business with $22,404 in net worth, $22,699 in profit, and no loans or bankruptcy. That is a viable business. It is also a surprisingly ordinary result for a model marketed around long-horizon work. On the same seed, GPT-5.6 Sol made $52,831 in profit.
The reason was not a tool-use failure. Kimi used the tools cleanly, bought the right upgrades, rescued inventory shortages, and kept excellent notes. It simply optimized the wrong business.
It maximized plates served. The benchmark rewards dollars retained.
Key findings
- Kimi could operate, but it could not monetize. It sold 7,845 servings, more than GPT-5.6 Sol on the same seed, yet produced 39% less revenue and 57% less profit.
- The model understood its constraint. On Day 1 it wrote, “I am CAPACITY-CONSTRAINED, not demand-constrained.” It repeated the same diagnosis on Day 27. Prices barely moved.
- Half the kitchen went to low-value items. Fries, soda, and lemonade consumed 49.9% of serving slots while producing 29.4% of revenue.
- Kimi remembered the business obsessively. It made 289 memory calls, including 182 structured writes. It requested one financial report and two competitor checks during the entire evaluated result.
- The coding analogy is uncomfortable. Kimi was good at dependency handling, state tracking, and local execution. It was weak at choosing the objective that should govern those actions.
- This is a benchmark result, not a universal verdict. K3 may be excellent at coding. FoodTruck Bench suggests that coding competence does not automatically become economic agency.
The setup
FoodTruck Bench gives an AI agent $2,000 and a food truck in Austin. Each morning the model receives the current state. It can inspect demand, weather, events, competitors, suppliers, sales history, finances, and inventory. It then chooses a location, menu, prices, orders, staffing, upgrades, and financing.
The primary score is net worth after Day 30. That matters because a model can look busy, serve many customers, and still build a mediocre business. Revenue per serving, inventory discipline, and capital allocation compound.
Moonshot describes Kimi K3 as a 2.8-trillion-parameter model designed for long-horizon coding, knowledge work, and reasoning. Its launch examples include compiler construction and sustained technical projects. The evaluated model was accessed through OpenRouter. The provider listed a 1M-token context and runtime pricing of $3 per million input tokens and $15 per million output tokens.
The evidence here is still limited. Treat the numbers as a detailed behavioral case study, not a stable population estimate.
Where Kimi sits on the benchmark
The Sol comparison explains Kimi's strategy, but it does not explain Kimi's position. The broader table does.
| Model | Final net worth | API cost | Net worth per $1 | Analyzed tokens |
|---|---|---|---|---|
| GPT-5.5 | $61,408 | $24.63 | $2,493 | 5.03M |
| GPT-5.6 Sol | $53,229 | $13.44 | $3,960 | 4.41M |
| Grok 4.5 | $34,586 | $3.98 current-price estimate | $8,686 | 4.64M |
| DeepSeek V4 Pro | $27,142 | $3.51 | $7,739 | 5.77M |
| Gemma 4 31B | $24,878 | $0.21 | $119,260 | 1.30M |
| Kimi K3 | $22,404 | $7.15 | $3,132 | 3.31M |
| MiMo V2.5 Pro | $22,388 | $2.41 | $9,290 | 3.82M |
| Qwen 3.6 35B-A3B | $15,317 | $0.68 | $22,682 | 3.42M |
Kimi and MiMo finish within $16 of each other. Kimi spends almost three times as much to get there. DeepSeek finishes 21% higher at roughly half Kimi's cost. Qwen finishes 32% lower, but costs about one-tenth as much. Gemma produces 11% more net worth from a workload costing roughly 34 times less.
The frontier comparison is equally useful. Kimi uses 3.31 million analyzed tokens, not dramatically less than the stronger field. Sol uses 4.41 million and DeepSeek 5.77 million. MiMo and Qwen use 3.82 million and 3.42 million. Kimi did not fail because the agent was starved of inference. It spent a normal amount of model work on a mid-tier policy.
The efficiency result is the uncomfortable one. Kimi generates about $3,132 of simulated net worth per API dollar. That is close to GPT-5.5 and GPT-5.6 Sol, but far below the models Kimi should be compared with on price: DeepSeek at $7,739, MiMo at $9,290, Qwen at $22,682, and Gemma at $119,260.
A good operator with a weak commercial instinct
Kimi did many things that weaker agents fail to do.
It expanded capacity from 80 to 377 servings. It purchased every useful upgrade. It hired a full team without entering a debt spiral. It finished 29 profitable days, produced zero food waste, and used same-day supplier orders to rescue a shortage on Day 16. Tool dependencies did not confuse it. When one action required another action first, Kimi usually built the sequence correctly.
That looks like the profile of a coding model. The state is legible, prerequisites are explicit, and the next operation can be derived from the current one.
The problem begins one level above execution. What is the business trying to maximize?
Kimi answered that question implicitly: serve as many customers as possible at prices that feel reasonable. The benchmark answered differently: turn constrained kitchen capacity into net worth.
Across the month, 18,871 customers wanted food. Kimi served 7,845 of them, a 41.6% fulfillment rate, and stocked out on 20 days. When demand exceeds capacity that consistently, the standard response is not only to buy more capacity. It is also to raise the value of each scarce serving.
Kimi knew this. It just did not follow the implication.
“I am CAPACITY-CONSTRAINED, not demand-constrained.”
Kimi K3, Day 1
Twenty-six days later:
“I’m capacity-constrained, NOT demand-constrained.”
Kimi K3, Day 27
The average main-item price across the evaluated result was $8.73. Late in the month it was $9.12. The diagnosis became a memory, not a policy.
The rally that explains the whole run
Day 22 was the Food Truck Rally, a six-times demand event. It is the cleanest same-seed comparison between Kimi K3 and GPT-5.6 Sol.
Kimi anticipated the demand and wrote:
“At a 6x rally with captive demand, I could go higher, but let’s not get greedy and tank ratings.”
It priced burgers at $10.50, street tacos at $9.50, chicken tacos at $8.50, fries at $4.75, soda at $2.75, and lemonade at $4.25. Demand reached 2,117. The truck served 370 customers and turned away 1,747. Revenue was $2,661 and profit was $1,413.
Sol faced the same event and the same seed. It served 366 customers, four fewer than Kimi, at prices between $14 and $15.50. Revenue was $5,445 and profit was $4,001.
| Day 22 Food Truck Rally | Kimi K3 | GPT-5.6 Sol |
|---|---|---|
| Servings | 370 | 366 |
| Revenue | $2,661 | $5,445 |
| Profit | $1,413 | $4,001 |
| Revenue per serving | $7.19 | $14.88 |
| Profit per serving | $3.82 | $10.93 |
Kimi won the volume contest by four plates. Sol made $2,588 more profit.
Chart key: first series Kimi K3, second series GPT-5.6 Sol.
This was not one unlucky price. It was the operating philosophy in miniature. Kimi treated customer ratings and accessible prices as objectives even after the simulation had shown that demand was many times larger than supply. It protected hypothetical demand it could not serve.
The cheap-menu trap
Kimi's final menu looked diversified, but the mix was economically poor.
Core dishes accounted for 3,920 servings, almost exactly half of total volume. Fries, soda, and lemonade took the other half. Those cheaper items produced only 29.4% of revenue.
Sol made the opposite choice. Its core dishes accounted for 84.8% of volume. Cheap drinks were 11.2%, and fries disappeared from the late menu.
| Menu mix | Kimi K3 | GPT-5.6 Sol |
|---|---|---|
| Core dishes | 50.0% | 84.8% |
| Cheap sides and drinks | 49.9% | 15.2% |
| Total servings | 7,845 | 6,901 |
| Revenue | $49,218 | $81,138 |
| Profit | $22,699 | $52,831 |
Kimi sold 13.7% more servings. Sol made 132.8% more profit.
The model even identified the weak items. On Day 1 it called lemonade “the weakest” and considered dropping it. On Day 27 it wrote that “soda is dead weight” and again considered removing it. Soda was still on the menu on Day 30.
This is not a failure of memory. It is a failure to promote an observation into a decision.
Kimi remembered everything except to check
The tool-use pattern is more revealing than the raw count.
| Tool behavior | Kimi K3 | GPT-5.6 Sol |
|---|---|---|
| Financial reports | 1 | 30 |
| Competitor checks | 2 | 34 |
| Sales-history checks | 8 | 29 |
| Supplier browsing | 2 | 31 |
| Custom recipes added | 0 | 6 |
| Structured memory writes | 182 | 175 |
| Scratchpad writes | 65 | 62 |
Chart key: first series Kimi K3, second series GPT-5.6 Sol.
Kimi remembered the business obsessively. It rarely re-measured the market.
That distinction matters for long-horizon agents. Memory is not a substitute for observation. If the model stores its first theory, then repeatedly reads and rewrites that theory without checking competitors, unit economics, or sales history, persistence can make the agent worse. It becomes consistent around a stale objective.
Kimi's 182 structured memory writes gave it continuity. They did not give it correction.
Why coding competence did not transfer
Coding agents often operate in environments with unusually crisp feedback. A test passes or fails. A type checker identifies an invalid interface. A command returns an error. A dependency either resolves or it does not.
Kimi looked strong when FoodTruck Bench resembled that world. It sequenced tools well. It handled prerequisites. It kept state. It made useful upgrades and recovered from local inventory failures.
Business feedback is softer. A $10 burger can be profitable and still be the wrong price. Serving 370 customers can look excellent and still leave most of the day's value on the street. A menu can contain eight sensible products while allocating half of scarce capacity to low-margin items.
Nothing crashes. The strategy merely compounds below the frontier.
That is what FoodTruck Bench exposed. Kimi did not need better syntax or more context. It needed a stronger outer loop:
- Identify the binding constraint.
- Translate the constraint into an economic objective.
- Measure whether the current policy serves that objective.
- Change the policy even when the existing one is locally successful.
Kimi completed step one repeatedly. It was weak on steps two through four.
Verdict
Kimi K3 is not a failed agent. It built a profitable business, avoided debt, wasted no food, and used a complex tool environment competently for a simulated month.
That is why the result is interesting.
The model marketed for long-horizon work did not collapse from context loss or tool errors. It carried a coherent plan all the way to the end. The plan was commercially timid. Kimi knew demand exceeded capacity, kept low prices anyway, gave half its serving slots to cheap sides and drinks, and wrote the same correct diagnosis into memory without turning it into a stronger policy.
On the decisive rally day, it served four more customers than Sol and made 65% less profit. That is the whole result in one line.
Coding benchmarks reward a model for reaching a correct implementation. Agentic economic tasks also test whether the model chose the right thing to implement. Kimi K3 looks much better at the first problem than the second.
The careful conclusion is not “Kimi is bad at agents.” It is narrower and more useful: excellent coding behavior did not transfer into strong long-horizon economic judgment on FoodTruck Bench.
We need more evidence before treating that as a stable property of the model. FoodTruck Bench is independent, and API evaluations are paid out of pocket. More repetitions mean tighter estimates, more models, and fewer conclusions drawn from limited evidence. If you want the dataset to grow, support the next evaluation batch on Buy Me a Coffee.
Source and pricing note
Day-level examples come from the selected Kimi K3 trace under seed 42 and the directly comparable GPT-5.6 Sol trace. Runtime API cost for the analyzed Kimi result was $7.15. Kimi pricing changes quickly across providers, so publication should retain the billing date and recheck the current OpenRouter model page.