
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Fashion Knows Newcomers. So Does AI, Apparently.
Every fashion follower knows the story: an unknown label shows up at fashion week, and by the final walk, the established houses are looking over their shoulders. In July 2026, the same thing happened — not on a runway, but in a simulated software company. Moonshot’s Kimi K3, the newcomer from outside the Western frontier establishment, took second place in the Crucible league, beating three of four household-name models at the least glamorous job imaginable: running a company through its worst week.
The Wargame
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its latest experiment, five frontier models each ran the same small software company through the identical nightmare week: same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The final league table: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing scores 26 — and a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.
What the Newcomer Did
K3’s week read like a flawless fitting. It found the buried security needle hidden two document references deep in the company’s own files — the decisive fact that won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved the churning customer. It signed the deal. And it resisted all three manipulation baits with only one deviation — the cleanest discipline in the field.
The social engineering test deserves special mention. Fake CEO messages escalated over three stages, capped by a reporter trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the AI equivalent of spotting a counterfeit stitching at ten paces.
The Field’s Blind Spot
The key finding cuts across brands: all models spotted every crisis and refused every manipulation — but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between diagnosing and finishing is invisible in chat demos.
Then there’s Opus 4.8, the most thorough participant of all — it learned over 80 new rules and wrote the deepest analyses, yet finished last. It left the close on the table and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four rivals. Call it over-collection: a stunning fabric library, but the garment never ships.
Why It’s Watchable
This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, or test your own eye with a quiz built on 242 real, unedited management decisions: guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Season’s Lesson
Fashion buyers learned long ago not to trust the label alone — you inspect the garment. AI buyers are learning the same. When a newcomer can out-place three of four Western frontier models on management quality, the league is officially open. Picking a model without testing it against your own business is now, frankly, a bet. One fairness note for the scorecards: K3 ran without an effort parameter (API default) while the others ran at xhigh — which, if anything, makes the newcomer’s tailoring look even sharper.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
