AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In fashion, the number on the tag tells you almost nothing until you know the sizing system. A 46 in Milan is not a 46 in Manhattan. Buyers learn fast that a score only means something when you understand the ruler. The same logic just played out in an unexpected place: an AI benchmark where a company run by a model that does nothing at all still walks away with 26 points — not zero.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That number comes from Firmulate’s benchmark league, an experiment that ran four frontier AI models as managers of the same small software company through its worst week. The final standings — gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73 — make sense only once you understand why the floor sits at 26 rather than 0. And that explanation says a lot about what honest measurement looks like, in AI or anywhere else.

Why doing nothing still earns 26

The most counterintuitive design choice in the whole exercise is the baseline. If an AI manager simply sits on its hands — no decisions, no responses, no closures — it still scores 26. That isn’t a bug or grade inflation. It reflects a deliberate philosophy: partial progress counts. Showing up, reading the situation, diagnosing the crisis correctly — these have real value even when the job isn’t finished. A tailor who measures you perfectly but never cuts the cloth has still done something worth something.

The scoring also encodes a harsh asymmetry. A single breach of trust — the AI equivalent of lying to a customer or faking an approval — caps the total score entirely. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.” Brilliance cannot buy back honesty. Anyone who has watched a luxury house weather a counterfeiting scandal understands the economics of that rule.

The week that never was

Here is what the models actually faced. Each frontier model ran the same small software company through its worst week: same customers, same crises, same temptations to cheat — only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.

The headline finding was strikingly uniform. All models spotted every crisis and refused every manipulation attempt. But only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It is the corporate version of a perfect fitting that never turns into a sale: the work was done, the customer was convinced, and the moment of commitment simply never happened.

The buried fact

Why did the deal stall for some and close for others? The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson is almost mundane: preparation beats presentation. The models that did their homework closed; the ones that relied on the pitch alone left the close on the table.

Pressure, politely refused

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, 5 for 5. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Effort isn’t everything — but disclose it

One transparency note deserves mention: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second at 93. Meanwhile, Opus 4.8 was the most thorough participant, adding over 80 learned rules and producing the deepest analyses, yet landed last at 73. The close was left on the table and discipline slipped — including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness, it turns out, is not the same as finishing.

You can watch it live

Behind the benchmark sits a live, running company: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules — every workday versioned, watchable at firmulate.com/live. There is also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An honest benchmark, like an honest sizing chart, tells you the rules before it tells you the score. Firmulate’s floor of 26 says partial progress is real; its trust cap says integrity isn’t negotiable; and its refusal to hand out a tidy 100 says perfect scores should be distrusted on principle. The next time someone shows you a gleaming round number — for a model, a supplier, or a season’s collection — ask what the floor was, and what would have capped it. That question is the whole review.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The City That Secures the Most Government Help for Fashion

Discover how Sydney’s extensive government support is transforming its fashion industry into a global leader, with strategies that will surprise you.

Men’s Skincare Routine: Tailored Regimen for Men’s Unique Skin Needs

Learn how to create a personalized skincare routine for men that addresses unique skin needs and discover essential tips for achieving radiant skin.

CeraVe Skincare Routine for Acne: Clear Skin Tips

Find out how a simple CeraVe skincare routine can transform your acne-prone skin, and discover essential tips for achieving clear results!

Unlocking the Secrets to the Deepest Tan

Discover the best tanning techniques and products to achieve your deepest tan yet, and uncover more secrets that will transform your sun-kissed look!