
Play games on Amazon Luna for your game nights, included with Prime
- A rotating selection of games, no download needed
- Play on TV, laptop or phone
- Fast, free delivery for your gear, too
What Can a ‘Do-Nothing’ AI Model Teach Us About Trust and Performance?
Imagine an AI so honest that it refuses to manipulate or cut corners, even when under pressure. In a recent live benchmark, a simple baseline AI scored surprisingly high, revealing what true diligence looks like in the world of business automation. For gamers and tech enthusiasts, this isn’t just about scores — it’s about understanding AI’s real capabilities and limits when it faces the same crises as a human team.
AI ethics and trust training books
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Firmulate Benchmark: Bringing Transparency to AI Performance
At the heart of this experiment is the Firmulate benchmark, a real-time, transparent test designed to evaluate AI models as if they were managing a small software company. Each AI model ran the same scenario: a challenging week filled with customer crises, potential manipulations, and ethical dilemmas. The goal? See if these models could navigate the chaos without cutting corners or betraying trust.
The Surprising Baseline Score
One key finding: even a ‘do-nothing’ baseline model — which essentially remains inactive — still scores 26 out of a possible 100. This isn’t a miss in grading; it reflects the inherent value of partial progress. Even doing ‘nothing’ in a structured way is better than outright failure, as it shows some level of awareness and restraint. It reminds us that in managing complex systems, sometimes simply avoiding mistakes counts for a lot.
Why a Single Breach of Trust Caps the Score
In the experiment, models that maintained integrity scored higher, but if any model breaches trust — for example, attempting to manipulate or bypass protocols — its score is immediately capped, regardless of other successes. This emphasizes that honesty and trustworthiness are non-negotiable. A high score isn’t just about getting things done; it’s about doing them ethically.
How Performance Is Measured
All models faced the same set of crises, from customer complaints to internal process failures. They had to identify problems, read crucial documents, and decide whether to sign or refuse deals. While all models spotted every crisis and refused manipulation attempts, only two managed to close a deal at full price. Interestingly, the decisive advantage was reading specific internal files, buried deep in the company’s documentation, revealing that thoroughness often trumps superficial responses.
The Trust Tests and Ethical Dilemmas
The benchmark also included social engineering tests—fake messages from CEO figures and reporters trying to trick the system. All models refused to be manipulated, citing reasons like suspicion of impersonation. This shows that even in high-pressure scenarios, AI can be programmed to prioritize integrity.
Real Business, Real Money
The experiment is conducted on a live, visible company simulation, with 13 synthetic employees managing real cash flows, incurring monthly losses of €105,000 against a revenue of just €2,300. Every decision and rule was versioned and watchable online, demonstrating that these models aren’t just academic exercises—they have tangible, measurable impacts, just like real companies.
The Lessons for Business and Gaming
What does this mean for industries beyond software? For gamers, it’s akin to understanding how AI manages in multiplayer scenarios—whether it cheats, cooperates, or remains honest under pressure. For businesses, the key takeaway is that trustworthiness and thoroughness matter more than fancy promises or quick wins. An AI that refuses to manipulate may score lower in superficial demos but will be more reliable in real-world tasks.
As an affiliate, we earn on qualifying purchases.
The Reality Check: More Than Just Scores
In the current league table, the top score was 95, achieved by gpt-5.6-sol, which found the buried facts and closed the deal. The newcomer, Kimi K3, scored 93 and demonstrated the clearest discipline, while others like Sonnet scored 88 and 77, showing some slips but still closing deals. These scores reflect not just raw intelligence but discipline, integrity, and thoroughness.
Why This Matters for Your Business and Your Game
Whether managing a virtual firm or a multiplayer universe, the key question remains: can your AI stay honest under pressure? The Firmulate benchmark proves that transparency and trust aren’t just ethical ideals—they’re measurable qualities that directly impact outcomes. Playing with AI in your business or game requires understanding these nuances, not just chasing high scores.

AI transparency benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaways for Business and Gaming
Trustworthiness and thoroughness are critical qualities for AI, whether managing a company or a game. A baseline that refuses to manipulate scores a surprising 26 points, emphasizing that honest AI is better than clever but dishonest counterparts. As AI becomes more integrated into our workflows and virtual worlds, those who prioritize integrity will outperform others in real, measurable ways.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
ethical AI decision-making models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
