
Good taste is revealed by the choices made under pressure
In fashion, polish can attract attention, but judgment determines whether a collection, campaign or brand actually works. The same distinction is becoming visible in artificial intelligence. A model may sound composed in a demonstration, yet still fail to complete the consequential task sitting behind the conversation.
Firmulate turns that gap into an unusually revealing interactive article. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how an AI handled a business situation, then try to identify the model from its managerial voice and behavior.
The premise is playful, but the material is not hypothetical. Each frontier model ran the same small software company through its worst week, confronting the same customers, crises and temptations. Every decision was versioned and auditable. What emerges is something closer to a management personality test than a writing contest: thoroughness, brevity, commercial instinct and operational discipline become recognizable traits.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Identical pressure, noticeably different performances
The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. The evaluation also imposed a hard ethical boundary: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
That boundary did not separate the field. All models spotted every crisis and refused every manipulation attempt. The sharper distinction was whether they converted sound analysis into completed business action. Only two signed the €55,000 deal their own work had earned. The experiment’s stark summary is: “Same diagnosis, same pitch — no signature.”
For anyone accustomed to evaluating style, the lesson feels familiar. Recognizing what should happen is not the same as executing it. A beautiful proposal that never reaches the client is still unfinished work. Firmulate’s decisions expose that final gap, where a model must stop discussing the correct move and actually make it.
The detail that changed the commercial outcome
The winning clue was not sitting prominently in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.
This was not a contest of who could produce the most confident sales language. It rewarded the quieter discipline of checking the company’s existing knowledge before acting. In a fashion business, the equivalent might be consulting production notes, account history or an earlier buyer conversation before negotiating. The relevant distinction is not glamour versus caution; it is informed action versus plausible improvisation.
Firm on trust, even when the pressure escalated
The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result matters because the models were not merely asked abstract questions about security. The manipulation attempts arrived amid the company’s operational turmoil. Refusing them showed that ethical and approval boundaries could survive distraction rather than disappearing when commercial pressure increased.
When exhaustive analysis becomes its own weakness
Opus 4.8 offers the most interesting character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.
Its performance complicates the assumption that more detail automatically produces better management. Thoroughness can be valuable, just as craft and finish are valuable in fashion. But the experiment suggests that analysis must remain connected to authority, timing and follow-through. Otherwise, sophistication becomes ornament around an incomplete decision.
Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase its 93 score, but it belongs beside the ranking when readers compare the field.
A company designed to make behavior visible
The live Firmulate company contains 13 synthetic employees and uses real money mechanics. It burns €105k per month against €2.3k MRR, while maintaining a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. The result is a real, watchable experiment in which management behavior unfolds over time rather than being selected from a polished collection of chat responses.

The quiz asks a bigger question than “Which model wrote this?”
Guessing the author of a decision is entertaining because the models develop recognizable tendencies. Yet the reveal points toward a practical issue for any brand considering an AI workforce: which behavior fits the responsibility being delegated?
The Crucible League shows that crisis recognition and resistance to manipulation may be shared strengths, while file-reading, disciplined escalation and closing the loop remain differentiators. The best manager is not necessarily the model with the most elaborate answer. It is the one that finds the buried fact, respects the boundary and completes the valuable work.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. For readers, the 242-decision quiz offers the more immediate experience: look past the verbal silhouette, study the judgment, and decide which AI management personality is really on display.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html