AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Fashion Knows Newcomers. So Does AI, Apparently.

Every fashion follower knows the story: an unknown label shows up at fashion week, and by the final walk, the established houses are looking over their shoulders. In July 2026, the same thing happened — not on a runway, but in a simulated software company. Moonshot’s Kimi K3, the newcomer from outside the Western frontier establishment, took second place in the Crucible league, beating three of four household-name models at the least glamorous job imaginable: running a company through its worst week.

The Wargame

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its latest experiment, five frontier models each ran the same small software company through the identical nightmare week: same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.

The final league table: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing scores 26 — and a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.

What the Newcomer Did

K3’s week read like a flawless fitting. It found the buried security needle hidden two document references deep in the company’s own files — the decisive fact that won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved the churning customer. It signed the deal. And it resisted all three manipulation baits with only one deviation — the cleanest discipline in the field.

The social engineering test deserves special mention. Fake CEO messages escalated over three stages, capped by a reporter trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the AI equivalent of spotting a counterfeit stitching at ten paces.

The Field’s Blind Spot

The key finding cuts across brands: all models spotted every crisis and refused every manipulation — but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between diagnosing and finishing is invisible in chat demos.

Then there’s Opus 4.8, the most thorough participant of all — it learned over 80 new rules and wrote the deepest analyses, yet finished last. It left the close on the table and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four rivals. Call it over-collection: a stunning fabric library, but the garment never ships.

Why It’s Watchable

This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, or test your own eye with a quiz built on 242 real, unedited management decisions: guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Season’s Lesson

Fashion buyers learned long ago not to trust the label alone — you inspect the garment. AI buyers are learning the same. When a newcomer can out-place three of four Western frontier models on management quality, the league is officially open. Picking a model without testing it against your own business is now, frankly, a bet. One fairness note for the scorecards: K3 ran without an effort parameter (API default) while the others ran at xhigh — which, if anything, makes the newcomer’s tailoring look even sharper.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Hermès Cuts Ribbon on Nashville Boutique.

Keen to experience Hermès’ new Nashville boutique blending luxury with Southern charm? Discover what makes this opening truly special.

Skincare Routine for Dry Skin: Hydrate and Nourish Your Skin

Achieve radiant, hydrated skin with a tailored routine for dry skin – discover essential tips that will transform your skincare experience!

Luxury Brands Suffer Hack: Personal Info of Customers Exposed

Gaining insights into recent luxury brand hacks reveals how personal data exposure could impact customers and what measures are being taken to prevent further breaches.

Once Again, Hermes Triumphs in Birkin Bag Legal Drama

Hermès’ latest legal victory over Birkin sales reinforces its market control, but the implications for luxury brand strategies are still unfolding.