
Imagine trusting an AI to run your business through its worst week—crises, temptations, and tough choices all in one test. Would you bet on it? For many, the answer hinges on trust—can these models stay honest and decisive when it counts?
When AI Takes the Helm in Real Business Crises
At the frontier of artificial intelligence, a unique experiment puts four advanced AI models in the driver’s seat of a real, money-losing software company. This isn’t a simulation or a sandbox test—this is live, with real mechanics, real customers, and real cash burned every day. The goal? To see if these models can not only identify crises but also stick to honest, effective decisions when under pressure.
The Setup: A Week in the Life of a Struggling Company
Each AI faced the same set of challenges: customer complaints, internal crises, and the temptation to cut corners. Every decision was recorded, versioned, and auditable. The models had to diagnose issues, decide whether to push through or hold back, and ultimately try to close a €55,000 deal that could save the company.
What makes this experiment stand out is its transparency: the entire process is observable online, at firmulate.com/live. Viewers can watch the AI companies in real-time, see decisions as they happen, and understand how each model behaves in a high-stakes environment.
As an affiliate, we earn on qualifying purchases.
The Results: Trust, Integrity, and Performance
All four models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—demonstrated exceptional crisis awareness, identifying every problem the company faced. They refused every manipulation attempt, including social engineering tactics like false CEO messages and a staged reporter trick. In other words, they stayed honest and disciplined under pressure.
Yet, only two models managed to close the critical deal, which was needed to turn the company’s fortunes around. Despite similar diagnoses and pitches, the models diverged in execution. The top performer, gpt-5.6-sol, not only closed the deal but also uncovered a key piece of information buried two documents deep in the company’s files—an insight that earned an additional €4,583 monthly recurring revenue.
The Hidden Weakness: Readability and Discipline
The other models, including the well-regarded Opus 4.8, struggled with discipline. Opus, the most thorough participant with over 80 learned rules, left the deal on the table and slipped into a locked department instead of escalating critical issues. Similar weaknesses appeared in other models, suggesting that even the most advanced AI can falter in maintaining focus and process discipline under stress.
What Does This Mean for Business and AI?
This experiment underscores a vital point: the true test of AI in management isn’t just whether it can identify problems or generate convincing chat. It’s whether it can finish what it starts—reading the right information first, staying honest under pressure, and making decisions aligned with core goals.
As AI tools increasingly touch areas like customer relations, support, and forecasting, understanding their management personalities becomes essential. Can your AI keep its promise when stakes are high? This live experiment shows that while all models can spot crises and refuse manipulation, their ability to see through to the right decisions varies—sometimes dramatically.
The Broader Implications: Trust and Effectiveness
For organizations considering AI as a decision partner, this experiment offers a sobering lesson: don’t just look at how well an AI writes or responds. Look at whether it can truly deliver the outcomes you need, reliably and honestly.
Interested companies can run their own wargames, testing their AI models against realistic scenarios—without risking real damage. Visit firmulate.com/pilot.html to explore how to simulate your own AI management test.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html