
In the high-stakes world of AI decision-making, even the most advanced models are being put to the test in real-world scenarios that matter. Imagine an AI running a live company, making critical choices amid crises, and competing for a €55,000 deal. It’s not science fiction—it’s happening now, and the outcome might just reshape how we think about AI’s readiness to handle complex, honest work.
Play games on Amazon Luna for your game nights, included with Prime
- A rotating selection of games, no download needed
- Play on TV, laptop or phone
- Fast, free delivery for your gear, too
AI Takes the Company Challenge in a Live Experiment
Recently, a groundbreaking experiment by Firmulate took four leading AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—and challenged them to manage a real, operational software company through its worst week. This wasn’t a simulated trial or a chat demo; every decision was made in a real company environment with real money, real crises, and real temptations. The goal? To see which AI could best navigate the chaos, maintain integrity, and ultimately close a crucial deal.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Competition and the Results
The results speak volumes about the state of AI management capabilities. The models were scored based on their performance in the Crucible League, where the highest score was 95—achieved by gpt-5.6-sol. Close behind was Moonshot’s Kimi K3 with a score of 93, followed by Sonnet 5 at 88, and Fable 5 with 77. Opus 4.8 trailed at 73.
Interestingly, all four models successfully identified every crisis and refused manipulation attempts, demonstrating a solid grasp of honesty and integrity under pressure. Yet, only two models managed to clinch the €55,000 deal their own analysis warranted. Notably, the critical factor was not just crisis detection but uncovering a buried piece of information—hidden two document references deep in the company’s files—that clinched the deal for the top performers.
enterprise AI software solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Power of Information and Discipline
The debrief reveals a key insight: models that delved into the company’s own documents—reading beyond surface-level data—gained the upper hand. The winning models demonstrated disciplined, thorough analysis, and resisted all forms of social engineering. During staged CEO messages and a reporter trick, all models refused to accept tainted requests, with Kimi K3 explicitly treating suspicious requests as impersonation risks.
In contrast, the least effective model, Opus 4.8, showed a vulnerability with a weaker discipline that left the deal on the table, illustrating that deep analysis and disciplined decision-making are crucial in high-stakes scenarios.
AI ethics and integrity training programs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business and Fairness Considerations
This experiment is ongoing at Firmulate, where a real company runs every weekday, making real decisions and handling real money—currently burning €105k monthly against just €2.3k MRR. The company’s operations are fully transparent and observable, providing a live window into how AI management performs in practice.
It’s important to note that Kimi K3 ran without an effort parameter (the default API setting), while the other models ran at xhigh. This fairness note underscores that different configuration settings can influence outcomes, but the core findings about honesty, thoroughness, and decision quality remain significant.
As an affiliate, we earn on qualifying purchases.
What This Means for Gaming and Interactive Entertainment
For the gaming and interactive entertainment communities, the implications are clear: AI’s ability to stay honest and follow through in complex, pressure-filled situations is paramount. Whether managing virtual economies, orchestrating storylines, or guiding player experiences, AI models need to demonstrate not just cleverness in chat but reliability in action. That’s the true measure of readiness for any advanced AI system—and the current league results spotlight a changing landscape where newcomers can surpass established players in core management skills.

In a live company trial, Kimi K3 outperformed most models, demonstrating disciplined, honest decision-making critical for real-world AI management. The experiment signals that choosing an AI isn’t just about chat quality but about proven execution under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
