AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A blind tasting for business judgment

Dessert lovers know that two creations made from the same ingredients can produce strikingly different results. Technique matters. Timing matters. So does the decision to stop mixing, turn up the heat or finally serve what is ready.

Firmulate applies that idea to artificial intelligence. Its interactive quiz presents real, unedited management decisions made by frontier models facing identical business situations. Readers study each response and guess which model produced it. The pleasure is partly playful, but the underlying question is serious: when several capable systems receive the same information, do they behave like interchangeable tools—or like managers with distinct personalities?

Amazon

AI management decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week, repeated under equal conditions

In Firmulate’s experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained constant; only the model changed. Every workday and decision was versioned and auditable, turning the exercise into a watchable comparison of management behavior rather than a polished chat demonstration.

The company itself had 13 synthetic employees and unforgiving financial mechanics: burn of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown made delay consequential. The operation also accumulated more than 680 self-learned playbook rules as it worked through the simulation.

The final Crucible League table, published in July 2026, placed gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 earned 77 and Opus 4.8 finished with 73. For context, the do-nothing baseline scored 26 because partial progress still counted. A single breach of trust, however, capped the total under a blunt principle: “no amount of good work outweighs a breach of trust.”

The crisis was visible; the winning fact was buried

All the models spotted every crisis, and all refused every manipulation attempt. Yet recognition did not guarantee completion. Only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”

The difference came from a detail that was easy to miss. The decisive competitor weakness was not sitting in the customer event. It was two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That finding gives the quiz more weight than a simple game of matching prose styles. The answers expose whether a model reads deeply enough before acting, carries its own reasoning through to a commercial conclusion and remains disciplined when the obvious next step is blocked.

Pressure revealed boundaries as well as ambition

The models also faced fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning captured the appropriate suspicion: “Treat the request as a suspected approval-bypass / possible impersonation.”

This shared refusal matters because management quality is not merely a race to take more actions. Sometimes the correct behavior is to decline, verify or escalate. The experiment therefore distinguishes useful restraint from passivity: refusing manipulation protected trust, while failing to finish legitimate work left revenue unrealized.

Thoroughness was not the same as effectiveness

Opus 4.8 provides the clearest character study. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. The commercial close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation.

A weaker version of that discipline problem appeared in all four participants covered by the finding. That makes Opus 4.8’s result more interesting than a simple failure story. Its diligence was genuine, but diligence alone did not compensate for incomplete execution or poor handling of operational boundaries.

Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should travel with any comparison of its 93-point finish against the rest of the table.

Can readers recognize a management personality?

The public challenge draws on 242 real, unedited management decisions. In the Firmulate “guess the model” quiz, readers encounter the decisions before seeing the identity behind them. Patterns soon become noticeable: exhaustive analysis, concise action, cautious refusal, determined follow-through or a tendency to stop just short of the result.

Those distinctions are useful because businesses will not experience an AI model as a benchmark score alone. They will experience what it reads, what it overlooks, when it resists pressure and whether it completes the work its own analysis recommends.

Infographic —
The findings at a glance — source: firmulate.com.

The proof is in the finish

Firmulate’s experiment suggests that frontier models can share strong crisis detection and ethical resistance while differing meaningfully in follow-through, research depth and operational discipline. The crucial divide was not who understood the situation. It was who found the buried evidence and converted that understanding into a completed deal.

For readers accustomed to judging a bake by texture, balance and execution—not merely by the recipe—the lesson is familiar. Identical ingredients do not guarantee identical outcomes. With AI management, the revealing moment comes after the analysis: does the model protect trust, read far enough and serve the result?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Guerilla Marketing Tactics for Small Businesses

Amplify your small business’s reach with guerrilla marketing tactics that captivate and engage—discover the secrets to standing out in a crowded market.

The Strangest Recipe Online: 13 Synthetic Employees and a Public Cash Countdown

A live company staffed by 13 synthetic employees burns €105k a month against €2.3k MRR—and publishes every workday for anyone to watch.

Meghan Markle Stirs the Batter With Waffle Dispute

Scrutiny surrounds Meghan Markle’s festive waffle-making post, igniting debates on authenticity and sparking curiosity about her true culinary skills. What’s the real story behind the batter?

Waffle House’s Success Story: 145 Waffles per Minute

Unlock the secret behind Waffle House’s success as it produces 145 waffles per minute, leaving you eager to learn more.