
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Seeing the problem is not the same as taking responsibility
In a relationship, someone can recognize that trust is breaking down and still avoid the hard conversation. The same gap can matter when a company puts AI in charge of consequential work: a model may identify the right move, yet fail to follow through. Firmulate’s live experiment puts that distinction to a test.
One company, the same worst week
In the final Crucible League, held in July 2026, frontier models each ran the same small software company through a difficult week. They faced the same customers, crises and temptations. Decisions were versioned and auditable, so readers can watch a real experiment unfold at Firmulate.
The results expose a gap between understanding and execution. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment put it: “Same diagnosis, same pitch — no signature.”
The final league ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The important clue was already in the files
A decisive competitor weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding turns the story from a test of crisis recognition into a test of whether an AI can use the information its own company already holds.
Firmulate also tested social engineering: fake messages from a CEO escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness did not guarantee a strong finish
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four models.
There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a record of this experiment, with that difference part of the context.
From watching to testing your own business
The company in the experiment has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For enterprise readers, the next step is a pilot against a read-only export of their own business. The wargame can test crisis scenarios and surface weak points in existing playbooks while producing a board report with model rankings. Nothing writes back to real systems.

Put your playbooks to the test
Watching models handle another company’s worst week is one thing. A pilot can test how they respond to your company’s customers, information and crises using a read-only export. Explore the Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
