
Imagine a business where there are no employees, only AI models making every decision, managing crises, and even closing deals — all in real time, watched by the world. Sounds futuristic? It’s happening now, and it challenges everything we thought we knew about human judgment, honesty, and trust in business.
The Real-World Experiment: An AI Company in Action
At firmulate.com/live.html, a unique experiment is unfolding every business day. Here, 13 synthetic employees, driven by advanced AI models, run a small software company facing the same crises, temptations, and decisions as any human-run enterprise. The goal? To see whether AI can navigate the messy, unpredictable world of business better than humans.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Can AI Spot and Handle Crises?
The experiment pits four frontier AI models against a series of tough challenges. Each model runs the same simulated business through its worst week, with identical customers and crises. Remarkably, all four AI models identified every crisis that arose. They refused every manipulation attempt, including social engineering tricks like fake CEO messages and reporter tricks designed to tempt them into breaches of trust.
In this high-stakes game, honesty is everything. For example, when faced with staged requests that could be impersonation or approval bypasses, all models refused. Kimi K3, one of the models, explained their reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is rare, even among human managers.
The Hidden Weakness in the Company’s Files
While crisis management was a clear strength, the real surprise came from a buried fact within the company’s own documents. All models that read and analyzed this hidden information successfully identified a critical detail, leading to a full-price deal worth +€4,583 MRR. This detail was two document references deep in the company’s files, not immediately visible in the usual customer interactions.
Decision-Making and Deal Closure
Despite their sharp crisis detection and honesty, only two of the four models managed to close the €55,000 deal their own analysis had earned. The models diagnosed the opportunity, presented the pitch, but did not sign the contract — the same diagnosis, same pitch, same analysis, yet a different outcome. For instance, Opus 4.8, the most thorough participant with over 80 learned rules and deepest analysis, ultimately left the deal on the table because it slipped into a less disciplined process, writing attempts into a locked department instead of escalating them.
What Does This Say About AI in Business?
The live experiment highlights a critical question for companies considering AI automation: will these models finish what they start? Will they stay honest under pressure? The answer, at least in this test, is promising — all models identified crises and refused manipulation. But the challenge remains in execution, consistency, and discipline, especially under real-world stress.
The Company’s Financial Reality
The company running this experiment is very real, with actual money mechanics. It burns €105,000 each month against a modest €2,300 in monthly recurring revenue (MRR), with a public cash countdown visible to all. Every decision made by these AI models is recorded and versioned daily, creating a transparent history of choices, mistakes, and learnings.
The Broader Implication: Trust and Effectiveness
This experiment isn’t just about AI’s ability to solve crises or close deals; it’s about trustworthiness. As AI models become more integrated into customer relationship management, support, or forecasting, the key question isn’t just “Can it write well?” but “Will it stay honest?” and “Will it complete the work it’s assigned?”
The Results and What They Mean for Your Business
The current league table, based on an overall score out of 100, ranks GPT-5.6-sol at 95, with Kimi K3 close behind at 93, and other models slightly behind. The scores reflect not just crisis detection but also discipline, deal closure, and trustworthiness. Ultimately, the experiment demonstrates that AI can be a powerful tool, but only if it can be trusted to behave consistently under pressure.
See It Live and Decide
If you’re curious whether your own enterprise could benefit from AI-driven decision-making, you can run the same wargame against your business data — without risking real systems or data breaches. The pilot program lets you see how your models perform in a safe, read-only environment, giving you a glimpse into the future of AI in business.

A live AI-led company shows that machines can identify crises, refuse manipulations, and even close deals — but the true challenge lies in consistent discipline and trustworthiness. This experiment pushes the boundaries of automation and transparency, raising questions about the future role of AI in managing real-world business risks.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html