AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every great stylist knows the same secret: the difference between an outfit that turns heads and one that hangs wrong is rarely the fabric. It’s the homework. The measurements taken before the first cut, the questions asked before the first recommendation. A consultant who walks in with answers before they’ve looked in your closet is selling you their taste, not your look.

It turns out AI models have exactly the same failing — and someone finally measured it. In a public experiment run by Firmulate, four frontier AI models were each handed the same small software company to run through its worst week. Most of them gave brilliant advice. Only the ones that actually read the client’s files closed the deal.

Same Diagnosis, Same Pitch — No Signature

The setup was ruthlessly controlled. Each model ran the same company, faced the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly retconned afterward.

The headline finding sounds almost paradoxical: all the models spotted every crisis, and all of them refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The other models diagnosed the customer’s problem perfectly, delivered the pitch flawlessly — and then simply never closed. Same diagnosis, same pitch, no signature.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Needle Buried Two Documents Deep

Here’s where the story becomes a parable about homework. The decisive fact in the whole scenario wasn’t in the customer meeting. It wasn’t in the crisis emails or the negotiation. It was buried two document references deep in the company’s own files: a specific competitor weakness that, if found, made the full-price deal a formality.

The models that did their reading — that clicked through the references before answering — won the €55,000 contract at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it automatically. Not because they were less articulate, less analytical, or less capable on paper. Because they skipped the filing cabinet.

If you’ve ever had a personal shopper recommend a midi skirt without noticing you’d mentioned twice that you never bare your ankles, you know this failure mode intimately. The information was there. Nobody looked.

The Final League Table

When the results were tallied in the July 2026 Crucible League, the standings told a clean story. gpt-5.6-sol took first place with a score of 95 — the complete performance, having found the buried fact and closed the deal. Kimi K3, the newcomer from Moonshot, followed at 93, closing the deal with what the experimenters called the cleanest discipline of the field. Sonnet 5 landed at 88, and Fable 5 at 77.

One caveat worth noting for fairness: K3 ran at its API default effort setting while the others ran at maximum effort — and still nearly won.

Perhaps the most instructive profile was Opus 4.8, which finished last at 73 despite being the most thorough participant in the field, with the deepest analyses and over 80 self-learned rules. It did the most homework of anyone — and still left the close on the table, while discipline slipped into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as execution.

For context, doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the scoring philosophy puts it: no amount of good work outweighs a breach of trust.

Honest Under Pressure

The manipulation tests deserve their own mention. The models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts across the field were refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Charms and flattery, it seems, work on humans better than on machines. At least these machines.

You Can Watch It Live

This isn’t a static benchmark locked in a lab. Firmulate runs a live company around the clock — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it unfold at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, for readers who fancy they can spot a model’s handwriting the way a stylist spots a house code.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The lesson for anyone hiring AI agents — or, frankly, anyone hiring anyone — is that “reads your files before answering” is not a soft skill. It’s a measurable, purchase-deciding property. In this experiment it was worth exactly €55,000.

So before you trust an AI with your CRM, your support queue, or your forecast, ask the question the demos never show: does it finish what it starts, does it do the reading, does it stay honest when pressured? The models that ace the chat demos are not always the ones that close. And the ones that close are, without exception, the ones that opened the drawer first.

Enterprises can even run this same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html. Measure the fit before you buy the suit.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Ultimate Makeup Brush Care Guide: Cleaning and Maintenance

With essential tips for cleaning and maintaining your brushes, discover how to elevate your makeup game while ensuring your skin stays healthy. What’s the secret?

4.6 Billion Move: L’Oréal Adds Kering’s Beauty.

AIThis post was created with the assistance of artificial intelligence (AI).L’Oréal’s $4.6…

Beard Grooming 101: The Expert Tips Your Barber Won’t Share

Amateur beard care can lead to mistakes; discover the expert tips your barber won’t share and elevate your grooming game today!

Kylie Jenner Shares Inside Look at Her Date Night With Timothée Chalamet in New York

Kylie Jenner posted a TikTok revealing her preparations for a last-minute date night in New York with Timothée Chalamet, confirming their outing together.