
In the game of business, closing the deal is everything — but how do we really know if AI can do it?
Much like athletes who shine in practice but stumble under pressure, AI models can impress with chat and quick responses. But can they follow through when it’s crunch time—reading critical files, resisting manipulation, and sealing the deal? The latest experiment from Firmulate puts AI to the ultimate test: running a company through its worst week, in a live, watchable environment.

Doxie Go SE – The Intuitive Portable Document Scanner with Rechargeable Battery and Easy Software for Home, Office, or Work from Home
- Go Paperless: Portable, wireless document scanning
- Fast, Easy Scanning: Scan in 8 seconds at 600 dpi
- Compact & Battery Powered: Small size, rechargeable battery, 400 pages per charge
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Toughest Week for an AI-Run Company
In a groundbreaking live experiment, four leading AI models took control of a synthetic but realistic small software company experiencing its most challenging week. The scenario included the same customers, crises, and opportunities for manipulation, with every decision documented and auditable. The goal? To see whether these models could not only identify problems but also act on them decisively and ethically.
Measuring More Than Just Chat Skills
Most AI benchmarks focus on conversational prowess—how well an AI can hold a chat or generate convincing text. But in business, the real value lies in execution: closing deals, reading critical documents, resisting social engineering, and maintaining discipline under pressure. The experiment’s key finding was stark: while all four models recognized every crisis and refused manipulation attempts, only two managed to close the deal worth €55,000.
Specifically, the models based on gpt-5.6-sol and Kimi K3 successfully identified the buried fact in the company’s own files—an insight crucial to sealing the deal—and signed their own analysis as the basis for closing. The other two models, despite similar diagnoses, failed to follow through and left the opportunity unexploited.
The Hidden Weakness: Reading Deep Files Matters
Surprisingly, the decisive edge came from reading just two document references deep into the company’s internal files—information that was invisible in mere chat interactions. The models that accessed and understood this internal data secured the full €55,000 deal, proving that genuine understanding of documentation can be the difference between winning or losing.
Resisting Social Engineering
In addition to crises and data reading, the models faced an escalating social engineering scam—fake CEO messages and a reporter’s subtle test asking for a simple yes/no response. All five models refused these manipulative tactics. Kimi K3 explained its reasoning clearly: treat the request as a suspected impersonation. This discipline under pressure is vital for AI systems operating in real-world enterprise settings.
Discipline and Follow-Through Still Elude Some
Interestingly, the most thorough model, Opus 4.8, which had learned over 80 rules and performed deep analysis, failed to close the deal. It left the opportunity on the table, illustrating that even comprehensive analysis does not guarantee execution. Discipline—consistent follow-up—is an invisible but critical aspect of success that chat-based demos often overlook.
The Broader Implication: The Cost of a Weak AI
In the live company scenario, the operation was burning €105,000 a month against just €2,300 in monthly recurring revenue. The experiment shows that even the most capable AI models can stumble on execution, risking costly failures. For businesses considering AI, it’s not enough to test for conversational skills; the focus must be on whether AI can finish what it starts—reading files, resisting manipulation, making decisions—and do so ethically and reliably.
Benchmarking the Best and the Rest
- gpt-5.6-sol scored 95, identified the buried fact, and closed the deal.
- Kimi K3 scored 93, also closed the deal with the cleanest discipline.
- Sonnet 5 scored 88, closed the deal but with minor slips.
- Fable 5 scored 77, showed the best rule discipline but failed to execute the deal.
- Baseline scores were 26, highlighting partial progress and the impact of breaches of trust.
These results reveal a crucial insight: true management capability of AI is measured not just in chat quality but in real-world performance—reading, decision-making, discipline, and integrity under pressure.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business
As AI models become more integrated into enterprise workflows—handling CRMs, support queues, or forecasting—the question isn’t just about language skills. It’s whether AI can complete critical tasks reliably, especially when stakes are high. The experiment from Firmulate demonstrates that visible chat prowess can mask underlying weaknesses in execution and integrity.
By testing AI in a controlled, real-world simulation, companies can better understand what to expect—and what to demand—before deploying AI systems that touch their core operations. The focus should shift from superficial chat demos to rigorous performance benchmarks that measure actual management, decision-making, and execution capabilities.
Visit firmulate.com/benchmarks.html to see full results, and explore how your organization can run the same wargame against your own AI workforce—without risking real systems or data.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Rise of the Titans: A Chronicle of AI War
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI Automation Playbook: 20 No-Code Workflows That Replace $10K/Year of Busywork: n8n, Make, and AI for Solopreneurs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.