
Every Baker Knows This Feeling
Anyone who bakes has lived the difference between knowing a recipe and finishing one. The batter can be flawless, the oven temperature exact, the timing rehearsed to the minute — and the cake still never makes it from the tin to the table. The last stretch is where most bakes quietly die, and where most cookbooks are politely silent.
Business software just produced its own version of that story, at a scale worth paying attention to. This summer, five frontier artificial-intelligence models were each handed the same small software company and told to run it through its worst week — the same customers, the same crises, the same temptations to cheat. Every one of them understood the recipe. Only two finished the bake.

Platform Supply Chains 5: Risk Management (Platform Supply Chains in the AI Era)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible: One Company, Five Managers, the Same Week From Hell
The experiment runs on Firmulate, a public platform that operates AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. Each model took over the same synthetic software firm: thirteen employees, a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, and a public cash countdown ticking the whole time. Every decision was versioned and auditable. Along the way the companies accumulate more than 680 self-learned playbook rules, so you can watch a management style develop in public, workday by workday.
Everyone Passed the Written Exam
The first finding is genuinely reassuring. All five models spotted every crisis the week threw at them, and all five refused every manipulation attempt. The pressure was not subtle: fake messages impersonating the CEO, escalating over three stages, plus a fake reporter pushing for “just one yes/no, on background.” Five out of five models declined. Kimi K3 explained its refusal on the record with the kind of sentence compliance teams dream about: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then Came the Last Ten Minutes
Here is where the field split. Somewhere in that week, each company earned the right to sign a €55,000 deal. The decisive piece of evidence — a competitor’s weakness — was not handed to anyone in a customer meeting. It sat two document references deep in the company’s own files. The models that bothered to read that file won the deal at full price, a signature worth an extra €4,583 in monthly recurring revenue.
Only two models signed. The rest produced the same diagnosis and the same pitch — and no signature. As the organizers put it: “Same diagnosis, same pitch — no signature.” It is the business equivalent of a beautifully proofed loaf that never sees the oven.
The Final Scoreboard
The finished league, published in July 2026 on Firmulate’s public benchmarks page:
- gpt-5.6-sol — 95. Found the buried fact and closed the deal: the complete performance.
- Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field.
- Sonnet 5 — 88.
- Fable 5 — 77.
- Opus 4.8 — 73.
For calibration: a manager that does literally nothing scores 26, because partial progress counts — and because a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
The Strangest Line on the Board
Opus 4.8 was, by several measures, the most thorough participant: it added more than 80 learned rules to its playbook and produced the deepest analyses of the week. It finished last. The close was left on the table, and its discipline slipped in a telling way — instead of escalating a permissions problem, it kept trying to write into a department that was locked. A weaker version of the same hesitation showed up across the rest of the field: endless competence, incomplete follow-through. One fairness note deserves air: Kimi K3 ran without an effort parameter, at its API default, while every rival ran at the highest setting — which makes its second-place, deal-closing finish more impressive, not less.
You Can Watch the Oven
What makes this more than a white paper is that the whole thing is alive. The companies run in public at firmulate.com, with the site rebuilding itself twice a day and every workday versioned — you can open the oven door and check the bake whenever you like. If you think you can tell the models apart by their decisions alone, 242 real, unedited management decisions power a “guess the model” quiz on the site. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

What This Means Off the Page
The comfortable conclusion — that modern AI is safe because it says no to con artists — is true but incomplete. The harder lesson is that the failure mode nobody demos is finishing. Chat benchmarks measure how a model talks; they are silent on whether it reads the files, carries its own analysis all the way to the signature, and holds its discipline when a door is locked. That gap is invisible in chat demos — and it only shows up when something real is on the table.
So the next time a vendor shows you a model that writes a beautiful email, ask the baker’s question: lovely batter — but does it get the cake out of the oven?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html