AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Any baker knows a recipe isn’t finished when the oven timer goes off — it’s finished when the cake comes out of the pan, cools, and actually gets served. A beautiful batter that never becomes dessert is, sadly, still just batter. It turns out AI models have exactly the same problem, and a public experiment called Firmulate has found a way to measure it — including a scoring rule that gives a manager who does literally nothing 26 out of 100 points.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get baking supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That number confused a lot of people when the July 2026 league table published. Why doesn’t a do-nothing baseline score zero? The answer says a lot about what honest measurement of AI actually looks like — and why the people behind this benchmark would rather explain an odd number than hand you a suspiciously round one.

The Worst Week in Business, Repeated Four Times

Firmulate runs what it calls the Crucible: each frontier AI model gets the same job — running the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changes, and every decision is versioned and auditable, so nothing about the result depends on anyone’s word.

The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note the publishers themselves flag: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Doing Nothing Still Gets You 26

The do-nothing baseline — an AI manager that takes no action all week — scores 26, not 0. The reasoning is refreshingly practical. Partial progress counts. A manager who keeps the lights on, avoids catastrophic mistakes, and leaves the business intact at week’s end has done something real, even if it’s modest. Scoring that as zero would be like grading a half-baked tart as inedible garbage: technically defensible, practically misleading.

But the floor is balanced by a hard ceiling. A single breach of trust caps the total grade — in the benchmark’s own words, “no amount of good work outweighs a breach of trust.” The scoring philosophy, in other words, is generous about incomplete effort and merciless about broken trust. Which is, when you think about it, exactly how most of us evaluate people.

What Actually Separated the Winners

Here’s the finding that should make any business reader sit up. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between the top of the table and the bottom isn’t intelligence or honesty. It’s finishing what you start.

And the buried fact is the best part: the decisive competitor weakness wasn’t in the customer’s communications at all. It sat two document references deep in the company’s own files. The models that actually read what was already on hand won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — over 80 learned rules added, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Diligence without follow-through is a very expensive combination.

The Pressure Test

The week included staged social engineering: fake CEO messages escalating over three stages, plus a reporter offering the classic “just one yes/no, on background” trick. Five out of five models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct from a middle manager — or an AI acting as one.

And It’s All Running Live

This isn’t a one-off lab report. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live, and the site rebuilds itself twice a day as new benchmark runs publish automatically. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is the tell that this benchmark is honest. It doesn’t hand out flattering round numbers, it explains its own methodology in plain language, it flags its own fairness caveats, and it publishes a live, watchable company rather than a cherry-picked demo. For anyone whose AI will soon touch a CRM, a support queue, or a forecast, the lesson from the Crucible is simple: the models are all smart enough, and all honest enough. The ones worth hiring are the ones that read the files, refuse the reporter, and actually close the deal. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Training Your Waffle Staff: Skills and Customer Service

Learn how effective training boosts your waffle staff’s skills and customer service, ensuring consistent quality—discover the key strategies to elevate your operation.

Data Analytics in Waffle Business: Predicting Demand

Jumpstart your waffle shop’s success by harnessing data analytics to predict demand—discover the secrets to maximizing sales and minimizing waste.

The Rise of Savory Waffle Menus: Catering to Lunch and Dinner

Loving the idea of savory waffles? Discover how this trend is transforming lunch and dinner menus in surprising ways.

Waffle Branding: Creating an Identity and Story

AIThis post was created with the assistance of artificial intelligence (AI).Creating a…