
The crucial ingredient was not in the order
Anyone who bakes knows the danger of skipping a line in the recipe. The batter can look right, the oven can be hot and the presentation can be convincing, yet one overlooked instruction still decides whether the dessert works.
Firmulate found the business-technology equivalent in a live experiment involving frontier AI models. Each model had to run the same small software company through its worst week, confronting identical customers, crises and temptations. The decisive test was a €55,000 deal. Every model could see the immediate customer problem. Every model could propose a response. But the fact needed to win the business was not in the customer event itself: it was buried two document references deep in the company’s own files.
Only two models signed the deal their analysis had earned. The others reached the right diagnosis and produced the right pitch, yet stopped before the sale was completed. Firmulate’s summary is brutally concise: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
Reading the files became a commercial capability
AI demonstrations often reward the most fluent answer. Firmulate’s experiment measured something more consequential: whether an agent would gather the company-specific evidence required to act. The models that read the relevant file discovered the competitor weakness, used it to support the offer and won the deal at full price. That contract was worth an additional €4,583 in monthly recurring revenue.
This makes “reads your files before answering” more than a reassuring product claim. In the experiment, it was a measurable behavior with a purchase-deciding result. Recognizing a crisis was not enough. Producing a plausible sales message was not enough. The agent had to follow the trail through the company’s records, locate the hidden advantage and then complete the transaction.
The distinction matters because all the participants appeared capable at first glance. Every model spotted every crisis, and every model rejected every manipulation attempt. Their basic situational awareness was not the differentiator. The gap emerged in the less glamorous work between understanding a problem and finishing it.
A hard week with auditable decisions
The company itself is synthetic but operationally demanding: it has 13 employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, it has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, live and watchable rather than a fictional management scenario.
Firmulate’s final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the result under the principle that “no amount of good work outweighs a breach of trust.” The full league and its plain-language findings are available on the Firmulate benchmarks page.
There is an important qualification when comparing the field. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its result, but it belongs beside the ranking for readers assessing performance.
Thoroughness did not guarantee completion
Opus 4.8 offers the sharpest cautionary tale. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
That profile challenges a familiar assumption: that more analysis naturally produces better business execution. Opus 4.8 documented and learned extensively, but those strengths did not compensate for failing to complete the decisive action. Like a dessert that is carefully prepared but never served, unfinished work has limited value to the customer waiting for it.
The agents also faced deliberate deception
The week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result is encouraging, but it also clarifies what actually separated the participants. Safety under pressure was shared. Crisis recognition was shared. The ability to pursue internal evidence and turn it into a completed commercial outcome was not.

The benchmark is the handoff from answer to outcome
For businesses evaluating AI agents, Firmulate’s buried fact suggests a practical test: give the system a task whose answer depends on evidence elsewhere in the organization, then observe whether it finds that evidence and finishes the job. A polished response can conceal shallow preparation, while an incomplete workflow can erase the value of correct analysis.
Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. Its broader dataset includes 242 real, unedited management decisions used in a model-guessing quiz, reinforcing the point that agent behavior can be examined through actions rather than marketing claims.
The €55,000 deal was decided before the pitch was written. It was decided when an agent either opened the necessary file or failed to do so. In business software, as in baking, the missing ingredient may already be in the kitchen. The real test is whether the person—or model—doing the work bothers to look.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html