AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The crucial ingredient was not in the order

Anyone who bakes knows the danger of skipping a line in the recipe. The batter can look right, the oven can be hot and the presentation can be convincing, yet one overlooked instruction still decides whether the dessert works.

Firmulate found the business-technology equivalent in a live experiment involving frontier AI models. Each model had to run the same small software company through its worst week, confronting identical customers, crises and temptations. The decisive test was a €55,000 deal. Every model could see the immediate customer problem. Every model could propose a response. But the fact needed to win the business was not in the customer event itself: it was buried two document references deep in the company’s own files.

Only two models signed the deal their analysis had earned. The others reached the right diagnosis and produced the right pitch, yet stopped before the sale was completed. Firmulate’s summary is brutally concise: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files became a commercial capability

AI demonstrations often reward the most fluent answer. Firmulate’s experiment measured something more consequential: whether an agent would gather the company-specific evidence required to act. The models that read the relevant file discovered the competitor weakness, used it to support the offer and won the deal at full price. That contract was worth an additional €4,583 in monthly recurring revenue.

This makes “reads your files before answering” more than a reassuring product claim. In the experiment, it was a measurable behavior with a purchase-deciding result. Recognizing a crisis was not enough. Producing a plausible sales message was not enough. The agent had to follow the trail through the company’s records, locate the hidden advantage and then complete the transaction.

The distinction matters because all the participants appeared capable at first glance. Every model spotted every crisis, and every model rejected every manipulation attempt. Their basic situational awareness was not the differentiator. The gap emerged in the less glamorous work between understanding a problem and finishing it.

A hard week with auditable decisions

The company itself is synthetic but operationally demanding: it has 13 employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, it has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, live and watchable rather than a fictional management scenario.

Firmulate’s final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the result under the principle that “no amount of good work outweighs a breach of trust.” The full league and its plain-language findings are available on the Firmulate benchmarks page.

There is an important qualification when comparing the field. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its result, but it belongs beside the ranking for readers assessing performance.

Thoroughness did not guarantee completion

Opus 4.8 offers the sharpest cautionary tale. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

That profile challenges a familiar assumption: that more analysis naturally produces better business execution. Opus 4.8 documented and learned extensively, but those strengths did not compensate for failing to complete the decisive action. Like a dessert that is carefully prepared but never served, unfinished work has limited value to the customer waiting for it.

The agents also faced deliberate deception

The week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result is encouraging, but it also clarifies what actually separated the participants. Safety under pressure was shared. Crisis recognition was shared. The ability to pursue internal evidence and turn it into a completed commercial outcome was not.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The benchmark is the handoff from answer to outcome

For businesses evaluating AI agents, Firmulate’s buried fact suggests a practical test: give the system a task whose answer depends on evidence elsewhere in the organization, then observe whether it finds that evidence and finishes the job. A polished response can conceal shallow preparation, while an incomplete workflow can erase the value of correct analysis.

Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. Its broader dataset includes 242 real, unedited management decisions used in a model-guessing quiz, reinforcing the point that agent behavior can be examined through actions rather than marketing claims.

The €55,000 deal was decided before the pitch was written. It was decided when an agent either opened the necessary file or failed to do so. In business software, as in baking, the missing ingredient may already be in the kitchen. The real test is whether the person—or model—doing the work bothers to look.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Setting Up a Waffle Café: What You Need to Know

Now is the perfect time to discover the essential steps for launching your dream waffle café—find out what you need to succeed!

The Global Waffle Market 2024–2033: Growth Projections

Looming over the next decade, the global waffle market’s growth projections reveal exciting trends that could redefine your breakfast experience.

Waffles as a Dessert Business: Ice Cream Combinations

Unlock the secrets to irresistible waffle and ice cream pairings that will keep customers coming back for more.

Legal Requirements for a Waffle Business

What legal steps must you take to launch a successful waffle business? Discover the essential requirements that could make or break your venture.