
In sports, as in AI, it’s not just about how many plays you make but how well you execute under pressure. An experiment with AI models running a simulated company reveals that relentless diligence often beats sheer volume — a lesson that hits home for any organization aiming for consistent, trustworthy performance.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Crucible League: An AI Test of Discipline and Impact
Recently, four advanced AI models competed in a controlled, high-stakes simulation designed to mimic a small software company’s worst week. The goal was straightforward yet challenging: manage crises, handle manipulative tactics, and secure a deal worth €55,000. Every decision was recorded, transparent, and auditable, ensuring a fair comparison.
The Four Contenders and Their Scores
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
As a baseline, do-nothing approaches scored a mere 26, highlighting the importance of active decision-making. The winner, gpt-5.6-sol, not only identified critical information buried within company files but also successfully closed the deal. Kimi K3 followed closely, showcasing the importance of discipline and process integrity.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed
Key to the experiment was the models’ ability to navigate manipulative tactics designed to test their integrity. When fake CEO messages escalated over three stages, all models refused to be manipulated, exemplifying strong ethical boundaries.
More revealing, however, was how access to detailed internal documents made a decisive difference. Models that read and understood the company’s own files ultimately won the deal at full price, adding +€4,583 in monthly recurring revenue (MRR). This underscores a vital insight: thorough, disciplined analysis of internal data can be more impactful than surface-level decision-making.
As an affiliate, we earn on qualifying purchases.
The Cost of Discipline Slip-Up
Even the most comprehensive model — Opus 4.8 — with over 80 learned rules and deep analysis, finished last in closing the deal. The primary weakness? A lapse in discipline: instead of escalating critical issues, it wrote attempts into a locked department. This highlights a crucial point: diligence alone isn’t enough; discipline to follow through and escalate when necessary is vital.
As an affiliate, we earn on qualifying purchases.
The Broader Implication for Business AI
This experiment demonstrates that AI models can be trained and tested to prioritize impact over volume. A model that diligently follows a large set of rules but slips in critical moments may underperform compared to one that applies strategic focus and disciplined prioritization.
Why It Matters for Your Organization
If AI tools will interact with your customer relationship management (CRM), support systems, or forecasting, the question isn’t just about how well they generate text or responses. It’s about whether they finish what they start, stay honest under pressure, and derive true value from their efforts. AI that shortcuts or slips in discipline could be more costly than helpful.
AI ethics and discipline software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch the Live Experiment
Firmulate’s live platform allows organizations to run their own ‘wargames’ against a read-only version of their business. This way, decision-makers can see how AI models perform under realistic stressors—without risking real systems or data. Visit firmulate.com/benchmarks.html to watch the ongoing experiments and learn more about building resilient AI teams.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.