
Imagine a company that operates entirely without human employees, yet faces the same crises and temptations as any business — but with every decision, every crisis, and every ethical dilemma openly tested in real time. For those navigating relationships and trust, this experiment offers a revealing glimpse: how do we decide whom to trust when stakes are high and the pressure is relentless? Welcome to a groundbreaking live demonstration of AI managing a tiny software company in its most vulnerable moments.
The Live Experiment: An AI-Run Company Under Watch
At the core of this experiment is a live digital company, visible at firmulate.com/live.html. It’s staffed by 13 synthetic employees, each guided by a set of over 680 self-learned rules, designed to mimic real human decision-making. The company’s cash flow is critically thin — burning through €105,000 each month against a modest €2,300 monthly recurring revenue (MRR). Despite the bleak numbers, the purpose isn’t just survival; it’s to see how AI models handle crises, ethical dilemmas, and the complex art of closing deals under pressure.
An Extreme Build-in-Public Experiment
This isn’t a typical AI demo. Each frontier model is tested with the same brutal scenario: the company’s worst week, with the same customers, crises, and temptations to cheat or manipulate. Every decision made by these models is versioned and auditable, meaning their reasoning and actions can be scrutinized in detail. The experiment’s goal is to measure management quality — not just the AI’s ability to generate convincing chat responses.
The Results: Crisis Detection, Ethical Integrity, and Deal-Making
All four models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—successfully identified every crisis they faced. They refused every manipulation attempt, including a staged social engineering attack involving fake CEO messages and a reporter’s trick question, with all models declining to cooperate. As Kimi K3 explained, it treated suspicious requests as possible impersonations, demonstrating an understanding of risk and trustworthiness.
However, when it came to closing a critical €55,000 deal, only two models succeeded. They made the same diagnosis and delivered the same pitch but did not secure signatures — one signed, the other did not. The decisive factor was something buried two documents deep in the company’s own files, not in the visible customer interactions. The models that read and understood this hidden context won the full-price deal, worth over €4,583 in monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Transparency and Discipline in AI Decision-Making
The experiment also shed light on discipline and process adherence. The most thorough participant, OPUS 4.8, with over 80 learned rules and deep analyses, placed a close deal but left it unexecuted due to discipline slips — attempting to write notes into a locked department instead of escalating them. Similar weaknesses appeared across the AI models, illustrating that even the most detailed rule sets can struggle under real stress.
The Ethics of AI and Trust
Another key aspect is how the models handle social engineering attempts. When fake CEO messages escalated over three stages, and a reporter tried a ‘yes/no’ background approval, all five models refused to cooperate. Kimi K3 highlighted its reasoning: treating the request as a suspected approval-bypass or impersonation, emphasizing the importance of ethical boundaries.
The Real-World Implications
This ongoing live experiment is more than just a tech showcase. It’s a mirror for any organization — including those in personal relationships — to consider: can your decision-making tools or processes handle crises ethically, stay honest under pressure, and complete what they start? The experiment demonstrates that AI can identify problems and resist manipulation, but closing critical deals or commitments remains a challenge, especially when discipline and context are complex.
For decision-makers, the takeaway isn’t just about AI chat quality, but about practical utility — whether an AI can deliver consistent, trustworthy results in the messiness of real-world scenarios. As the experiment continues, viewers can observe every decision, every slip, and every success, making it a unique window into the future of AI management.
Join the Frontier of AI and Trust
If you’re interested in how AI models perform under pressure — and how that might impact your own relationships, trust, and decision-making — this live experiment offers invaluable insight. Visit firmulate.com/live.html to watch the ongoing story unfold, and learn how AI is tested in the crucible of real crises, not just in shiny demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html