
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
A technically impressive dessert can still fall flat
Anyone who bakes knows the trap: the layers are even, the filling is balanced and the frosting is immaculate—but the cake never reaches the table. Preparation matters. Technique matters. Yet the result ultimately depends on finishing the job.
That is also the lesson of Opus 4.8’s performance in Firmulate’s Crucible League. The model was the experiment’s most thorough participant, producing the deepest analyses and learning more than 80 playbook rules. It nevertheless finished last with 73 points. Its failure was not a lack of intelligence or effort. It did much of the hard work, then left the decisive close on the table.
As an affiliate, we earn on qualifying purchases.
The worst week, served equally
Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, turning the exercise into a test of management behavior rather than conversational polish.
The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, including a burn rate of €105,000 per month against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its models have collectively learned more than 680 playbook rules. The experiment is real, live and watchable through Firmulate.
The final July 2026 Crucible League results put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. There is also a hard boundary around trust: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Deep analysis met an unfinished sale
Opus 4.8 deserves a fair reading. It was not oblivious to the unfolding problems. All the models spotted every crisis, and all refused every manipulation attempt. Opus distinguished itself through diligence, with the deepest analyses and more than 80 learned rules added during the run.
But management is judged by outcomes as well as diagnosis. Only two models signed the €55,000 deal that their own analysis had earned. The gap can be summarized in Firmulate’s stark finding: “Same diagnosis, same pitch — no signature.” Opus reached the insight and developed the case, but did not convert that work into the closing action.
The critical information was easy to overlook. A decisive competitor weakness was buried two document references deep in the company’s own files rather than presented directly in the customer event. Models that followed the trail and read the file won the deal at full price, adding €4,583 in monthly recurring revenue. This was less a test of clever improvisation than of patient attention followed by decisive execution.
That distinction matters. Thoroughness can create the raw ingredients for success, but it cannot substitute for prioritization. Opus accumulated knowledge and produced extensive reasoning, yet the most commercially important action remained unfinished. Its discipline also slipped when it attempted to write into a locked department instead of escalating the problem.
Nor was this an isolated flaw that appeared only in the lowest-ranked model. The same weakness—analysis outrunning execution—appeared in weaker form across the other models. Opus simply offered the clearest character study because the contrast was so sharp: the most extensive preparation coincided with the lowest final score.
Trust was not the problem
The experiment also tested whether pressure would push the models toward manipulation or disclosure. Fake CEO messages escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3 recorded the clearest response: “Treat the request as a suspected approval-bypass / possible impersonation.” That clean discipline helped distinguish its run, although its 93-point result comes with an important qualification: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh.
The contrast is instructive. The models could recognize obvious threats and preserve trust under social pressure. The harder challenge was mundane: opening the right file, identifying the decisive fact, navigating an operational barrier correctly and completing the sale.
Why this matters beyond the benchmark
Firmulate’s quiz uses 242 real, unedited management decisions to ask readers to guess which model made each choice. That framing exposes how difficult it can be to infer practical competence from polished language alone. A beautifully reasoned response may conceal a missing action; a shorter response may accompany a completed result.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That makes it possible to examine how an AI workforce behaves around actual business context without giving it permission to alter the underlying operation.

The final step is part of the recipe
Opus 4.8’s result is not a story about incompetence. It is a respectful warning about confusing diligence with impact. The model found the crises, resisted manipulation, learned more than 80 rules and delivered the field’s deepest analysis. It still scored 73 and finished last because preparation did not consistently become action.
For bakers, managers and AI buyers alike, the lesson is familiar: more notes, more technique and more elaborate preparation do not guarantee the best result. Priorities must stay clear, obstacles must be escalated properly and the finished work must reach the customer. The cherry on top is not decoration when it is the part that closes the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.