
Imagine coaching a sports team where every play is decided by an AI, and you get to see which AI genuinely makes the best strategic calls. The stakes? A real, money-losing software company, every decision visible in real time. This isn’t science fiction—it’s the groundbreaking experiment from Firmulate, where AI models are put through their managerial paces in a high-stakes environment.
Putting AI to the Test in a Real Business Crisis
At the heart of this experiment is a simple yet powerful question: how well can current frontier AI models handle complex management decisions under pressure? To find out, four different models — including GPT-5.6, Kimi K3, Sonnet 5, and Fable 5 — were each tasked with running a small software company through its worst week, mimicking every challenge a real business might face.
This setup was no simulation. The company, which operates with 13 synthetic employees and deals with real money mechanics, was faced with crises, customer issues, and ethical temptations, all in an environment where every decision was logged, auditable, and comparable across models.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Findings: AI Can Spot Crises and Reject Manipulation
Remarkably, all four models successfully identified every crisis and refused every manipulation attempt — including social engineering tactics like staged CEO messages or reporter tricks. In one scenario, a fake CEO message escalated over three stages, and every AI model declined to escalate or act on the fake request.
However, when it came to closing critical deals, only two models showed consistent success. Despite identical analysis and pitches, only GPT-5.6 and Kimi K3 signed the €55,000 contract they had earned through their own evaluations. The other models, including the most thorough contender Opus 4.8, left lucrative opportunities unclaimed, revealing a subtle weakness in discipline and follow-through.
internal document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files Matters
Digging deeper, the experiment uncovered a crucial insight: the decisive advantage belonged to the models that read and interpret documents in the company’s own files, not just external cues or customer interactions. In fact, the models that accessed two document references deep into internal files secured the full deal worth an extra €4,583 monthly recurring revenue (MRR). This suggests that the ability to read and analyze internal company data is essential for effective decision-making.
As an affiliate, we earn on qualifying purchases.
Personality Profiles and Management Styles
The models exhibited distinct management personalities. For example, Fable 5, despite being thorough with over 80 learned rules and deep analyses, finished last—leaving critical deals on the table and slipping into a less disciplined state by the end. Its decisions were meticulous but lacked the discipline to escalate issues appropriately. Kimi K3, on the other hand, ran without an effort parameter, making decisions more efficiently and fairly, and successfully closing deals without fuss.
This variability highlights that AI models are not just tools—they have measurable management personalities that influence their performance under pressure. Some are thorough but cautious; others are efficient but less disciplined. These traits can significantly impact their ability to navigate complex, real-world business scenarios.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business Leaders
Most importantly, the experiment underscores that the real value of AI in management isn’t just in generating nice-sounding chat or reports. It’s in whether the AI can finish what it starts—reading critical internal data, resisting manipulation, and closing deals. As one of the key findings states, “the gap is invisible in chat demos,” but it can be decisive in real operations.
For business leaders, this means that testing AI models in simulated, high-pressure environments can reveal their true management qualities before deployment. Firmulate’s live experiments, available at firmulate.com, demonstrate how AI companies can be evaluated as real corporate entities—facing real crises, managing real money, and making real decisions.
The Next Step: Wargaming Your AI Workforce
Businesses interested in future-proofing their operations can run their own wargames using Firmulate’s platform. These tests are safe, read-only exports of your business, allowing you to see how different AI models would perform without risking real systems. It’s a way to identify promising AI candidates and understand their management personalities—before hiring or trusting them with critical tasks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html