
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
AI’s Hidden Business Skills: Beyond Chat and Creativity
Imagine an AI that doesn’t just generate content or hold conversations but actually manages a real company through its toughest week — making critical decisions, resisting manipulations, and closing deals. That’s the bold experiment now underway, and the results are surprising even the most seasoned tech observers.
As an affiliate, we earn on qualifying purchases.
The Live Business Wargame: Testing AI in the Real World
At the forefront of this innovation is Firmulate, which has created a live, watchable experiment where leading AI models each run a small software company facing the same crises, temptations, and customer scenarios. Each model is tested in a controlled environment that mimics real-world pressures — from customer churn to security breaches — and their decisions are fully versioned and auditable.
What the Models Achieved
In the latest round, four frontier models competed in a crucial league, with their performances scored on a scale from 0 to 100. The results? The top performer, GPT-5.6-sol, scored 95, just ahead of the newcomer Kimi K3 with 93. Both closed a critical €55,000 deal, bringing in €4,583 in monthly recurring revenue, while the others — Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73) — also closed deals but with slightly weaker discipline and process adherence.
Decisive Strengths and Weaknesses
One of the experiment’s key insights was that all models successfully identified every crisis and refused manipulative tactics, such as fake CEO messages or reporter tricks. However, the real differentiator was in the reading depth. The winner, Kimi K3, uncovered a crucial buried fact within the company’s internal files — a piece of information critical to closing the deal at full price. Models that reviewed these documents thoroughly won the sale at maximum value, illustrating the importance of deep contextual understanding.
Discipline Under Pressure
Another interesting aspect was how models handled social engineering attempts. All five models refused to approve suspicious requests, treating them as potential impersonation or approval bypasses. Kimi K3 explicitly reasoned: ‘Treat the request as a suspected approval-bypass / possible impersonation,’ demonstrating cautious and disciplined decision-making under pressure.
The Real Company and Its Challenges
The live experiment runs on a real, functioning company with 13 synthetic employees, managing real money mechanics — burning €105k monthly against only €2.3k in MRR. Every workday, its decision processes are recorded and made transparent, making the experiment not just a theoretical test but a real-time showcase of AI-driven management.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Creators Alike
For those in the music, audio, and creator tech industries, this experiment signals a shift: AI models are moving beyond simple content generation to becoming active participants — capable of making strategic decisions, maintaining integrity under pressure, and ultimately, closing deals. The question is no longer whether AI can write well but whether it can finish what it starts, read the right information first, and stay honest when stakes are high.
The Fairness Note
It’s important to mention that Kimi K3 was run without an effort parameter (the API default), whereas the other models used an xhigh effort setting. This difference underscores the model’s innate discipline and decision-making quality without additional tuning.
As an affiliate, we earn on qualifying purchases.
What’s Next? Open League and Consumer Choice
The league standings highlight a competitive field where newcomers can beat established models by simply reading deeply and acting decisively. For companies considering AI as part of their management toolkit, this experiment emphasizes the importance of testing AI in real operational scenarios before making commitments. The platform allows enterprises to run their own wargames against copies of their business — with no risk to actual systems — available at firmulate.com.
Final Takeaway
AI models are proving they can do more than chat. They’re capable of managing, negotiating, and even closing deals — if they’re guided by disciplined, thorough decision-making. As the leaderboard shows, the best AI doesn’t just talk; it acts with integrity and precision, even under stress. For creators and businesses alike, embracing this new frontier could redefine what’s possible in the near future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
