
Imagine a high-stakes sports game where only scoring matters. But what if the real challenge was your team’s ability to stay honest, adapt under pressure, and see the big picture? In today’s business landscape, that’s exactly what AI management tools are being tested for — not just their cleverness in solving puzzles, but their resilience amid crises. Now, a groundbreaking live experiment by Firmulate exposes how AI models measure up in the real world, revealing critical gaps that traditional benchmarks miss.
Beyond the Scoreboard: Why Management Quality Is the New Benchmark
Most AI evaluations focus on answer scores—think of coding competitions or chatbots judged by correctness or fluency. But in actual business scenarios, success hinges on more than just getting the right answer. It’s about handling crises, resisting manipulation, maintaining honesty, and completing complex tasks under pressure. A new live experiment conducted by Firmulate puts AI models through a realistic simulation of running a small software company facing a week of disasters, confrontations, and temptations.
The Experiment: Putting AI Management to the Test
Four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—each managed the same company, encountering the same crises: customer churn, price hikes, down rounds, and PR scandals. Every decision was timestamped and auditable, mirroring real-world pressures. The models identified every crisis and refused to be manipulated—yet only two managed to close a lucrative deal worth €55,000.
The Hidden Weaknesses: Deep in the Files, Not the Crisis
The most decisive factor in winning the deal was not how the models responded to external crises but what they uncovered in the company’s own documents. Reading two layers deep into internal files, the models that did so secured the contract at full price, worth over €4,580 per month in recurring revenue. It highlights a fundamental flaw in benchmarks that only measure surface-level answer quality: the true test lies in depth, comprehension, and integrity.
Resisting Social Engineering and Manipulation
In a staged social engineering attack, fake CEO messages and a reporter trick, all models refused to cooperate, demonstrating an understanding of impersonation risks. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that models can recognize sophisticated attempts to manipulate management decisions under pressure, an essential quality for real-world AI managers.
The Reality of a Live Company
The experiment’s live company operates with 13 synthetic employees, burning €105,000 monthly against €2,300 in monthly revenue. It runs every business day, with over 680 learned rules, and its decisions are publicly viewable. Watching it in action at firmulate.com/live reveals how these models handle real-time crises, deadlines, and ethical dilemmas—far beyond what a chat demo can show.
What the Scores Mean for Business
The leaderboard is straightforward: gpt-5.6-sol leads with a score of 95, having uncovered the buried fact and closing the deal. Kimi K3 follows closely at 93, with the cleanest discipline and integrity. Sonnet 5 and Opus 4.8 trail behind, showing that thoroughness and discipline matter, but depth and honesty are decisive. The key insight: models that read and interpret internal documents perform significantly better in securing high-value deals.
Implications for Management and AI Adoption
For business leaders, the takeaway is clear: the real measure of an AI’s usefulness isn’t its chat quality or superficial answer correctness. It’s whether it can see through deception, read complex internal data, and stay honest under pressure. These qualities are essential when AI touches your CRM, support queues, or strategic planning. If these tools can’t finish what they start or recognize critical internal signals, they risk failing at the most crucial moments.
Tools to Prepare Your Business
Firmulate offers a platform where enterprises can run their own management wargames—simulating crises and decision-making scenarios—without risking real systems. This allows teams to evaluate how their AI workforce performs under the most demanding conditions, ensuring readiness before deployment. Learn more at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity and social engineering detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.