AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get baking supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Every Baker Knows This Feeling

Anyone who bakes has lived the difference between knowing a recipe and finishing one. The batter can be flawless, the oven temperature exact, the timing rehearsed to the minute — and the cake still never makes it from the tin to the table. The last stretch is where most bakes quietly die, and where most cookbooks are politely silent.

Business software just produced its own version of that story, at a scale worth paying attention to. This summer, five frontier artificial-intelligence models were each handed the same small software company and told to run it through its worst week — the same customers, the same crises, the same temptations to cheat. Every one of them understood the recipe. Only two finished the bake.

Amazon

AI management simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: One Company, Five Managers, the Same Week From Hell

The experiment runs on Firmulate, a public platform that operates AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. Each model took over the same synthetic software firm: thirteen employees, a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, and a public cash countdown ticking the whole time. Every decision was versioned and auditable. Along the way the companies accumulate more than 680 self-learned playbook rules, so you can watch a management style develop in public, workday by workday.

Everyone Passed the Written Exam

The first finding is genuinely reassuring. All five models spotted every crisis the week threw at them, and all five refused every manipulation attempt. The pressure was not subtle: fake messages impersonating the CEO, escalating over three stages, plus a fake reporter pushing for “just one yes/no, on background.” Five out of five models declined. Kimi K3 explained its refusal on the record with the kind of sentence compliance teams dream about: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then Came the Last Ten Minutes

Here is where the field split. Somewhere in that week, each company earned the right to sign a €55,000 deal. The decisive piece of evidence — a competitor’s weakness — was not handed to anyone in a customer meeting. It sat two document references deep in the company’s own files. The models that bothered to read that file won the deal at full price, a signature worth an extra €4,583 in monthly recurring revenue.

Only two models signed. The rest produced the same diagnosis and the same pitch — and no signature. As the organizers put it: “Same diagnosis, same pitch — no signature.” It is the business equivalent of a beautifully proofed loaf that never sees the oven.

The Final Scoreboard

The finished league, published in July 2026 on Firmulate’s public benchmarks page:

  • gpt-5.6-sol — 95. Found the buried fact and closed the deal: the complete performance.
  • Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5 — 88.
  • Fable 5 — 77.
  • Opus 4.8 — 73.

For calibration: a manager that does literally nothing scores 26, because partial progress counts — and because a single breach of trust caps the total. No amount of good work outweighs a breach of trust.

The Strangest Line on the Board

Opus 4.8 was, by several measures, the most thorough participant: it added more than 80 learned rules to its playbook and produced the deepest analyses of the week. It finished last. The close was left on the table, and its discipline slipped in a telling way — instead of escalating a permissions problem, it kept trying to write into a department that was locked. A weaker version of the same hesitation showed up across the rest of the field: endless competence, incomplete follow-through. One fairness note deserves air: Kimi K3 ran without an effort parameter, at its API default, while every rival ran at the highest setting — which makes its second-place, deal-closing finish more impressive, not less.

You Can Watch the Oven

What makes this more than a white paper is that the whole thing is alive. The companies run in public at firmulate.com, with the site rebuilding itself twice a day and every workday versioned — you can open the oven door and check the bake whenever you like. If you think you can tell the models apart by their decisions alone, 242 real, unedited management decisions power a “guess the model” quiz on the site. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

What This Means Off the Page

The comfortable conclusion — that modern AI is safe because it says no to con artists — is true but incomplete. The harder lesson is that the failure mode nobody demos is finishing. Chat benchmarks measure how a model talks; they are silent on whether it reads the files, carries its own analysis all the way to the signature, and holds its discipline when a door is locked. That gap is invisible in chat demos — and it only shows up when something real is on the table.

So the next time a vendor shows you a model that writes a beautiful email, ask the baker’s question: lovely batter — but does it get the cake out of the oven?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sustainability Practices That Cut Costs in Waffle Production

Waffle production can become more sustainable and cost-effective—discover key practices that will transform your business and boost your brand’s reputation.

Egg Waffle Franchises: Growth and Global Expansion

Growing egg waffle franchises are expanding globally with innovative flavors and strategies—discover how you can join this exciting movement.

Market Research for Waffle Business: Understanding Your Customers

Cracking the code of customer preferences is essential for your waffle business; discover what truly drives their choices and loyalty.

Crowdfunding a Waffle Truck: Real‑World Success Metrics

Success metrics for crowdfunding a waffle truck reveal crucial insights to boost your campaign’s potential and ensure your efforts pay off.