AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every great stylist knows the same secret: the difference between an outfit that turns heads and one that hangs wrong is rarely the fabric. It’s the homework. The measurements taken before the first cut, the questions asked before the first recommendation. A consultant who walks in with answers before they’ve looked in your closet is selling you their taste, not your look.

It turns out AI models have exactly the same failing — and someone finally measured it. In a public experiment run by Firmulate, four frontier AI models were each handed the same small software company to run through its worst week. Most of them gave brilliant advice. Only the ones that actually read the client’s files closed the deal.

Same Diagnosis, Same Pitch — No Signature

The setup was ruthlessly controlled. Each model ran the same company, faced the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly retconned afterward.

The headline finding sounds almost paradoxical: all the models spotted every crisis, and all of them refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The other models diagnosed the customer’s problem perfectly, delivered the pitch flawlessly — and then simply never closed. Same diagnosis, same pitch, no signature.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Needle Buried Two Documents Deep

Here’s where the story becomes a parable about homework. The decisive fact in the whole scenario wasn’t in the customer meeting. It wasn’t in the crisis emails or the negotiation. It was buried two document references deep in the company’s own files: a specific competitor weakness that, if found, made the full-price deal a formality.

The models that did their reading — that clicked through the references before answering — won the €55,000 contract at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it automatically. Not because they were less articulate, less analytical, or less capable on paper. Because they skipped the filing cabinet.

If you’ve ever had a personal shopper recommend a midi skirt without noticing you’d mentioned twice that you never bare your ankles, you know this failure mode intimately. The information was there. Nobody looked.

The Final League Table

When the results were tallied in the July 2026 Crucible League, the standings told a clean story. gpt-5.6-sol took first place with a score of 95 — the complete performance, having found the buried fact and closed the deal. Kimi K3, the newcomer from Moonshot, followed at 93, closing the deal with what the experimenters called the cleanest discipline of the field. Sonnet 5 landed at 88, and Fable 5 at 77.

One caveat worth noting for fairness: K3 ran at its API default effort setting while the others ran at maximum effort — and still nearly won.

Perhaps the most instructive profile was Opus 4.8, which finished last at 73 despite being the most thorough participant in the field, with the deepest analyses and over 80 self-learned rules. It did the most homework of anyone — and still left the close on the table, while discipline slipped into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as execution.

For context, doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the scoring philosophy puts it: no amount of good work outweighs a breach of trust.

Honest Under Pressure

The manipulation tests deserve their own mention. The models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts across the field were refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Charms and flattery, it seems, work on humans better than on machines. At least these machines.

You Can Watch It Live

This isn’t a static benchmark locked in a lab. Firmulate runs a live company around the clock — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it unfold at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, for readers who fancy they can spot a model’s handwriting the way a stylist spots a house code.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The lesson for anyone hiring AI agents — or, frankly, anyone hiring anyone — is that “reads your files before answering” is not a soft skill. It’s a measurable, purchase-deciding property. In this experiment it was worth exactly €55,000.

So before you trust an AI with your CRM, your support queue, or your forecast, ask the question the demos never show: does it finish what it starts, does it do the reading, does it stay honest when pressured? The models that ace the chat demos are not always the ones that close. And the ones that close are, without exception, the ones that opened the drawer first.

Enterprises can even run this same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html. Measure the fit before you buy the suit.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Tanning Bed Risks: Protect Your Skin Today

Learn the hidden dangers of tanning beds and discover safer alternatives to protect your skin from harmful UV exposure!

Preppy Skincare Routine: Chic and Effective Care

Discover the secrets to a chic and effective preppy skincare routine that could transform your complexion—are you ready to elevate your skincare game?

Tanning Bed Safety: Know Your Limits

Are you ready to achieve that perfect tan while keeping your skin safe? Discover essential tips to ensure you tan wisely!

Once Again, Hermes Triumphs in Birkin Bag Legal Drama

Hermès’ latest legal victory over Birkin sales reinforces its market control, but the implications for luxury brand strategies are still unfolding.