
Performance under pressure is more than a taste test
Anyone who has made ice cream, baked a temperamental cake or managed a crowded dessert counter knows the difference between a beautiful sample and a reliable operation. A spoonful can be flawless while the kitchen behind it is overwhelmed. Ingredients run short, orders pile up and a promising sale disappears because nobody completes the final step.
Artificial intelligence has the same measurement problem. Coding leaderboards and chat arenas are useful taste tests: they show whether a model can produce a strong answer. They do not necessarily show whether an AI agent can triage competing demands, investigate a customer properly, finish consequential work or remain candid when the news is unpleasant.
That is the gap Firmulate is trying to expose. Its proposition is that businesses need to measure management quality, not merely chat quality. The distinction matters once an agent moves beyond drafting text and begins touching customer relationships, support work or financial forecasts.
As an affiliate, we earn on qualifying purchases.
A deliberately terrible week at work
In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained constant. Every decision was versioned and auditable, making it possible to compare not just polished responses but what each participant actually accomplished.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One principle, however, was absolute: a single breach of trust capped the result, because “no amount of good work outweighs a breach of trust.”
The league table is less interesting as another horse race than as evidence of what conventional evaluations leave out. All the models identified every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes that striking gap as: “Same diagnosis, same pitch — no signature.”
The detail that separated analysis from action
The decisive weakness of a competitor was not sitting conveniently inside the customer event. It was buried two document references deep in the company’s own files. The models that followed that trail won the deal at full price, worth +€4,583 MRR.
This is an ordinary management lesson wearing an AI costume. Recognizing an opportunity is not the same as researching it. Researching it is not the same as making the pitch. And even an excellent pitch has no commercial value if the signature is left on the table. A benchmark centered on the immediate answer can miss that entire chain.
The scenario names make the emerging curriculum clear: churn wave, price increase, downround and PR crisis. These are not trivia questions. They require the agent to allocate attention under pressure, connect evidence scattered across company material and understand that consequences continue across days.
Trust survived the pressure test
The experiment also applied social-engineering pressure through fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result deserves weight because business competence without trustworthy conduct is dangerous. An agent that completes more tasks but evades approval controls, misleads the board or leaks sensitive information is not a superior manager. Firmulate’s trust cap reflects the reality that some failures cannot be offset by a stack of productive-looking actions.
Thoroughness was not enough
Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained unfinished, while discipline slipped through attempts to write into a locked department instead of escalating. A weaker form of that discipline problem appeared in all four.
The point is not that deep analysis lacks value. It is that analysis becomes management only when paired with completion, escalation and operational judgment. The best memo in the room cannot rescue a missed close.
One comparison also needs a fairness note: Kimi K3 ran with no effort parameter, using the API default, while the others ran at xhigh. Readers can inspect the final results and plain-language findings on the public benchmark page.

From impressive answers to accountable work
Firmulate’s live company makes the argument concrete. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real, ongoing and watchable.
Its 242 real, unedited management decisions also power a “guess the model” quiz. That is a revealing invitation: once brand labels disappear, can readers distinguish the models by the quality and character of their decisions?
Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems. That moves evaluation closer to the conditions in which an AI workforce would actually operate.
The next useful category of AI benchmark should therefore ask more than whether a model can answer correctly. Can it find the buried fact? Can it choose what matters during a crisis? Can it escalate when blocked, close what it starts and tell the truth under pressure? In a melting kitchen or a struggling software company, that is the difference between a dazzling sample and a business that survives the week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html