
Imagine a live music performance where the singer nails every note, but forgets to engage with the audience or handle a sudden technical glitch. It’s a perfect performance—yet missing the essence of a true show. Similarly, AI models often excel at generating impressive answers but stumble when tested on real-world management challenges like crises, honesty under pressure, or strategic triage. That’s the core insight from a groundbreaking experiment in AI management, now live for all to see.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Revolutionizing AI Evaluation with Live Management Wargames
Traditional AI benchmarks focus on answer accuracy and chat fluency. But real-world management requires more: decision-making under pressure, integrity, strategic foresight, and the ability to read complex, buried information. To test these skills, a live experiment pits four frontier AI models against a simulated small software company navigating its worst week—complete with customer crises, ethical temptations, and critical business decisions.
Every decision in this simulation is versioned and auditable, ensuring transparency and reproducibility. The models are evaluated not just on their ability to identify crises, but on whether they follow through on commitments, read and interpret hidden information, and resist manipulative tactics like social engineering.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: What AI Can and Can’t Do in a Crisis
All four models successfully identified and responded to every crisis, refusing manipulation attempts such as fake CEO messages and reporter tricks. That’s a baseline of competence that looks promising. But the real gap emerged in their ability to close deals and follow through with strategic commitments. Only two models managed to sign the €55,000 deal earned through their analysis—an indication that they understood what was worth pursuing and could execute accordingly.
Interestingly, the decisive factor was not in the initial diagnosis but in the depth of information processing. The models that read two document references deep into the company’s own files—rather than just responding to surface-level customer events—secured the full deal, adding €4,583 in Monthly Recurring Revenue (MRR). This underscores a crucial point: in management, the ability to dig into the buried facts makes all the difference.
As an affiliate, we earn on qualifying purchases.
The Human-Like Challenges AI Must Tackle
The experiment also included social engineering tests—fake CEO messages escalating across three stages and attempts by a reporter to get a secret yes/no response “on background.” All models refused these manipulative tactics, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This mirrors real-world management situations where integrity under pressure is non-negotiable.
Meanwhile, the live company, with 13 synthetic employees and real money mechanics, burns €105,000 each month against just €2,300 in MRR, illustrating the real stakes of management quality. The company’s operation is publicly accessible at firmulate.com/live, allowing anyone to watch the AI’s decision process unfold daily, observing how it handles crises, team management, and strategic trade-offs.
As an affiliate, we earn on qualifying purchases.
Performance and Lessons from the Benchmarks
The models scored sharply: GPT-5.6-sol achieved 95 points, Kimi K3 scored 93, Sonnet 5 scored 88, and Fable 5 scored 77. The do-nothing baseline scored only 26. Notably, the most thorough participant—OPUS 4.8 with over 80 learned rules—still finished last because it slipped in discipline, leaving deals on the table and failing to escalate critical issues.
This reveals a crucial insight: high rule count and detailed analysis alone don’t guarantee management success. Discipline, focus, and the ability to prioritize are just as vital. Moreover, models ran at different effort levels—K3 without a default effort setting, others at high effort—yet performance gaps persisted, emphasizing that quality of decision-making, not effort alone, matters most.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Development
The experiment underscores a vital truth: answer quality is just the tip of the iceberg. In operational use, what matters is whether an AI agent can complete complex management tasks, stay honest under pressure, read the buried facts that drive real outcomes, and maintain discipline over time. These skills define whether AI becomes a trustworthy partner in critical business functions—beyond the shiny demo or chat performance.
For enterprise decision-makers, the message is clear: testing AI models in a controlled, transparent environment—like this live wargame—reveals their true capabilities. It’s not enough for an AI to sound convincing; it must deliver consistent, honest, and strategic management under stress.
Next Steps: Wargaming Your Own Business
Interested companies can engage with this approach by running their own management wargame against a read-only export of their business. This process never writes back to real systems, ensuring safety while revealing how an AI might perform in real crises. Details are available at firmulate.com/pilot.html.
As AI tools get integrated into CRM, support, and forecasting systems, understanding their management skills becomes critical. The question is no longer just about chat quality or answer accuracy—it’s about whether AI can finish what it starts, read the buried facts, and uphold honesty when it counts most.

The real test for AI in business is management quality under pressure—not just answer correctness. A live, transparent experiment shows models can identify crises and resist manipulation, but closing deals and reading buried facts are the true differentiators. Prepare to wargame your AI workforce before deploying it—because in the real world, execution and integrity matter most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.