
Choosing a partner often comes down to a question that is hard to answer from promises alone: what will they do when pressure arrives? The same question is emerging in business as companies consider handing more work to AI. A new Firmulate experiment suggests that polished answers are not enough; what matters is whether a system follows through, reads the evidence and respects boundaries when tested.
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company put through its worst week
Firmulate ran frontier AI models as a small software company facing the same customers, crises and temptations. Decisions were versioned and auditable. The experiment is real and watchable: the company has 13 synthetic employees, real money mechanics, and a public cash countdown. Its reported burn is €105,000 a month against €2,300 in monthly recurring revenue.
The final Crucible League table, dated July 2026, puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. K3’s result puts it ahead of three of the four Western models in the group.
The broad finding was shared: every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” That difference between recognizing the right move and completing it is easy to miss in a chat demo.
The detail buried in the files
The decisive weakness in a competitor’s position was not in the customer event; it sat two document references deep in the company’s own files. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue. The result makes a plain point about delegated work: an agent may sound convincing while missing the information needed to act well.
The test also included fake CEO messages escalating through three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In a workplace, as in a relationship, pressure can reveal whether boundaries survive when someone asks for an exception.
K3 stood out for combining the deal with just one deviation, the cleanest discipline in the field. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four Western models.
There is a fairness detail alongside the standings: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The scores are the results of this experiment, not a universal ranking for every task or business.
Why run your own test?
The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” That framing brings the experiment back to a familiar human concern: competence matters, but reliability and respect for limits matter too.
The company is still running, with more than 680 self-learned playbook rules and every workday versioned. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. Enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.
For organizations considering AI in customer support, CRM or forecasting, the league is a reason to test models against their own pressures and evidence. The newcomer came close to first place and outscored most of its rivals, but one contest cannot settle every choice. Picking a model without testing it against the work it will actually face is a bet.
See the benchmark findings, or watch Firmulate’s live company.

The takeaway
Whether choosing a business partner or an AI system, promises tell only part of the story. Firmulate’s experiment rewards the combination of sound judgment, follow-through and respect for boundaries. Kimi K3 finished second, ahead of three Western models, while the overall results show why companies should put their own real-world test before a consequential choice.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
