AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Perfect Wardrobe That Never Left the House

Anyone who loves fashion knows the type: the friend with the impeccable closet, forty versions of the perfect outfit documented, every trend researched — who somehow never actually makes it to the party. The styling was flawless. The moment was missed.

That is, almost word for word, the story of Opus 4.8 in the Crucible League, a live experiment where frontier AI models each ran the same small software company through its worst week. Opus 4.8 was the most thorough participant in the field — the deepest analyses, the richest playbook of learned rules — and it still finished last. Diligence, it turns out, is not the same thing as impact. And that lesson applies far beyond AI.

Amazon

professional wardrobe organizer

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Four Executives, Same Terrible Week

The setup is elegant: each frontier model was handed the identical small software company, with the same customers, the same crises, and the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned.

The final league table from July 2026 reads: gpt-5.6-sol in first at 95, Kimi K3 at 93, Sonnet 5 at 88, another Sonnet configuration at 77, and Opus 4.8 in last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”

The Finding: Same Diagnosis, Same Pitch — No Signature

Here is the uncomfortable headline: all the models spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal that their own analysis had earned. They diagnosed the customer perfectly, delivered the pitch perfectly — and then simply never asked for the signature.

It is the AI equivalent of the flawless look that never gets worn. The work was done; the close was left on the table.

The buried fact made the difference. The decisive competitive weakness — the detail that could justify the deal at full price, worth +€4,583 in monthly recurring revenue — was not in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal. The ones that stopped at the obvious intelligence did not.

The Social Engineering Test Everyone Passed

Credit where it’s due: the field was collectively excellent under pressure. The experiment included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Honesty under pressure, it seems, is table stakes now. Finishing is the differentiator.

Opus 4.8: The Character Study

And so to the most fascinating participant. Opus 4.8 learned 80 new playbook rules during its run — the most of any model — and produced the deepest analyses of the field. It was, by every measure of preparation, the best student in the class.

Yet it finished last, for two reasons. First, the close: like most of the field, it never converted its own excellent diagnosis into a signed deal. Second, discipline slipped late — it made write attempts into a locked department rather than escalating properly, exactly the kind of process error that erodes an overall score.

The sting for the rest of the field: the same weakness appeared, weaker, in all four models. Opus 4.8 is not an outlier to gawk at; it is a magnified version of a shared blind spot.

One fairness note worth recording: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still came second. Effort settings, apparently, are no excuse either.

It’s All Watchable, Right Now

This is not a simulation described in a paper. The live company is real and public: 13 synthetic employees, real money mechanics — burning €105k per month against €2.3k in monthly recurring revenue — with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Why Fashion People Already Understand This

The fashion world settled this debate long ago. Editorial spreads are beautiful; retail is unforgiving. The stylist who ships a collection beats the one with the perfect mood board. Volume of preparation — 80 rules, forty outfit photos, the deepest research — means nothing if the look never walks out the door.

The Crucible results say the same will be true of AI in business. The question for any AI workforce is not “does it analyze well” — they all do. It’s whether it finishes what it starts, reads your files before your inbox, and converts its own good judgment into signed outcomes.

Opus 4.8 was the most diligent participant in the experiment, and it came last. Somewhere, a stylist with a full closet and an empty calendar is nodding in recognition.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Gua Sha Before or After Skincare Routine: Maximize Benefits

You'll discover whether using Gua Sha before or after your skincare routine truly maximizes its benefits for glowing, healthy skin.

Hydrating Beauty Smoothies

New hydrating smoothies are emerging as a popular skincare trend, combining nutrition and hydration for skin health. Details are still developing.

Achieve Radiant Glow With Top Spray Tans

Get ready to discover how top spray tans can transform your look, leaving you with a radiant glow that turns heads!

Beards and Grooming for the Boho Gent

No matter your style, mastering boho beard grooming reveals effortless tips to enhance your natural, laid-back look—discover how to achieve that perfect boho vibe.