
Anyone who has watched a runway show knows the feeling: a look that photographs beautifully and falls apart the moment someone tries to walk across a rainy street in it. Presentation is not performance. That gap — between how something shows and how it holds up — is exactly the problem now facing companies rushing to hire AI agents for real work.
For two years, the industry has judged AI models the way front-row critics judge collections: chat arenas and coding leaderboards, where models answer one polished question at a time. But a live experiment at Firmulate just ran something closer to a wear-test — the same four frontier models given the same small software company to run through its worst week — and the results expose a measurement gap the leaderboards can’t see.
The stress test, not the runway
Firmulate, which bills itself as “the AI company emulator,” gave each model the identical job: run a small software firm through a brutal week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable. The final league table from the July 2026 Crucible run: gpt-5.6-sol in first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73.
Here’s what should stop any executive mid-scroll: all five models spotted every crisis, and all five refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s “just one yes/no, on background” trick. Kimi K3 put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
And yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. A do-nothing baseline still scored 26, because partial progress counts, but a single breach of trust caps the total. As Firmulate puts it: no amount of good work outweighs a breach of trust.
The fact buried in the filing cabinet
The decisive moment of the week wasn’t the churn wave, the price increase, the downround rumor, or the PR crisis. It was a competitor’s weakness sitting two document references deep in the company’s own files — not in the customer event at all. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t left the close on the table.
That is a management failure, not a chat failure. No leaderboard measures whether an agent reads your files before it talks to your customer, or finishes what it starts when nobody is grading a single clever answer.
The cautionary tale
The most striking profile belongs to Opus 4.8: the most thorough participant in the field, with the deepest analyses and over 80 learned rules, and still last place. The close went unsigned, and discipline slipped — it attempted writes into a locked department instead of escalating. Firmulate notes the same weakness appeared, weaker, in all four models. One fairness footnote: K3 ran without an effort parameter while the others ran at xhigh, and still nearly won.
Not a slide deck
The experiment runs on a genuinely live company: 13 synthetic employees, real money mechanics, €105k monthly burn against just €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — rebuilding itself twice a day and watchable at firmulate.com/live. For the quiz-inclined, 242 real, unedited management decisions power a “guess the model” game. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The fashion industry learned long ago that a garment must survive more than the photograph. AI buyers are learning the same lesson: chat quality is the fitting-room mirror; management quality is the garment after a hundred wears. Firmulate’s field suggests the models are honest under pressure and sharp at diagnosis — but only half of them finish the job. Before you hand an agent your CRM, your support queue, or your forecast, ask the question the arenas never do: does it close, and does it stay clean? Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.