AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A dessert shop can have a brilliant recipe and still lose the customer: the order goes missing, a supplier problem gets overlooked, or a promising catering deal never gets signed. Firmulate, a live experiment that puts AI models in charge of a simulated company, tests that kind of follow-through under pressure.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get baking supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, served to every model

Firmulate gave each frontier model the same small software company, the same customers, crises and temptations. Every decision was versioned and auditable. The point was to assess management in action, rather than judge a model by how polished its answers sound.

The final Crucible League table, dated July 2026, puts Moonshot’s Kimi K3 in second place with 93 points, just behind gpt-5.6-sol at 95. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The gap between first and second is small; the broader result is striking: the newcomer outscored three of the four Western frontier models in the trial.

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The researchers’ summary captures the gap: “Same diagnosis, same pitch — no signature.”

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The detail hidden in the files

The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found that detail and closed. It also saved the churning customer and resisted all three baits, with only one deviation—the cleanest discipline in the field.

Those baits included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs a finish

Opus 4.8 provides a telling counterpoint. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared across all four.

The experiment’s do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.” That is a demanding standard for an AI workforce, and one that connects directly to the questions a business owner might ask of a kitchen team or a shop manager: can they spot the problem, keep their word and carry the job through?

The simulated company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. The live experiment is watchable at Firmulate.

The results also come with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. And a quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the team before trusting the recipe

K3’s result makes the frontier contest look open, while the unsigned deals show why a strong answer is not the same as a completed job. For a bakery, café or any business considering AI agents, the practical question is how a model handles your own customers, records and pressure points. Firmulate says enterprises can run the wargame against a read-only export of their business; nothing writes back to real systems. The league table and plain-language findings are at Firmulate’s benchmarks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Market Research for Waffle Business: Understanding Your Customers

Cracking the code of customer preferences is essential for your waffle business; discover what truly drives their choices and loyalty.

Crowdfunding a Waffle Truck: Real‑World Success Metrics

Success metrics for crowdfunding a waffle truck reveal crucial insights to boost your campaign’s potential and ensure your efforts pay off.

The Psychology of Menu Design: Placing Waffles for Maximum Profit

I’m about to reveal how strategic waffle placement on menus can significantly boost your profits and influence customer choices.

Starting a Waffle Food Truck: A Step-by-Step Guide

Discover essential steps to launch your waffle food truck and unlock the secrets to delicious success that awaits you!