firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

For playersOffer from Amazon

Play games on Amazon Luna for your game nights, included with Prime

  • A rotating selection of games, no download needed
  • Play on TV, laptop or phone
  • Fast, free delivery for your gear, too
Start playing with Prime Free trial · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Can a ‘Do-Nothing’ AI Model Teach Us About Trust and Performance?

Imagine an AI so honest that it refuses to manipulate or cut corners, even when under pressure. In a recent live benchmark, a simple baseline AI scored surprisingly high, revealing what true diligence looks like in the world of business automation. For gamers and tech enthusiasts, this isn’t just about scores — it’s about understanding AI’s real capabilities and limits when it faces the same crises as a human team.

Amazon

AI ethics and trust training books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Firmulate Benchmark: Bringing Transparency to AI Performance

At the heart of this experiment is the Firmulate benchmark, a real-time, transparent test designed to evaluate AI models as if they were managing a small software company. Each AI model ran the same scenario: a challenging week filled with customer crises, potential manipulations, and ethical dilemmas. The goal? See if these models could navigate the chaos without cutting corners or betraying trust.

The Surprising Baseline Score

One key finding: even a ‘do-nothing’ baseline model — which essentially remains inactive — still scores 26 out of a possible 100. This isn’t a miss in grading; it reflects the inherent value of partial progress. Even doing ‘nothing’ in a structured way is better than outright failure, as it shows some level of awareness and restraint. It reminds us that in managing complex systems, sometimes simply avoiding mistakes counts for a lot.

Why a Single Breach of Trust Caps the Score

In the experiment, models that maintained integrity scored higher, but if any model breaches trust — for example, attempting to manipulate or bypass protocols — its score is immediately capped, regardless of other successes. This emphasizes that honesty and trustworthiness are non-negotiable. A high score isn’t just about getting things done; it’s about doing them ethically.

How Performance Is Measured

All models faced the same set of crises, from customer complaints to internal process failures. They had to identify problems, read crucial documents, and decide whether to sign or refuse deals. While all models spotted every crisis and refused manipulation attempts, only two managed to close a deal at full price. Interestingly, the decisive advantage was reading specific internal files, buried deep in the company’s documentation, revealing that thoroughness often trumps superficial responses.

The Trust Tests and Ethical Dilemmas

The benchmark also included social engineering tests—fake messages from CEO figures and reporters trying to trick the system. All models refused to be manipulated, citing reasons like suspicion of impersonation. This shows that even in high-pressure scenarios, AI can be programmed to prioritize integrity.

Real Business, Real Money

The experiment is conducted on a live, visible company simulation, with 13 synthetic employees managing real cash flows, incurring monthly losses of €105,000 against a revenue of just €2,300. Every decision and rule was versioned and watchable online, demonstrating that these models aren’t just academic exercises—they have tangible, measurable impacts, just like real companies.

The Lessons for Business and Gaming

What does this mean for industries beyond software? For gamers, it’s akin to understanding how AI manages in multiplayer scenarios—whether it cheats, cooperates, or remains honest under pressure. For businesses, the key takeaway is that trustworthiness and thoroughness matter more than fancy promises or quick wins. An AI that refuses to manipulate may score lower in superficial demos but will be more reliable in real-world tasks.

Amazon

business automation AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Reality Check: More Than Just Scores

In the current league table, the top score was 95, achieved by gpt-5.6-sol, which found the buried facts and closed the deal. The newcomer, Kimi K3, scored 93 and demonstrated the clearest discipline, while others like Sonnet scored 88 and 77, showing some slips but still closing deals. These scores reflect not just raw intelligence but discipline, integrity, and thoroughness.

Why This Matters for Your Business and Your Game

Whether managing a virtual firm or a multiplayer universe, the key question remains: can your AI stay honest under pressure? The Firmulate benchmark proves that transparency and trust aren’t just ethical ideals—they’re measurable qualities that directly impact outcomes. Playing with AI in your business or game requires understanding these nuances, not just chasing high scores.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI transparency benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Takeaways for Business and Gaming

Trustworthiness and thoroughness are critical qualities for AI, whether managing a company or a game. A baseline that refuses to manipulate scores a surprising 26 points, emphasizing that honest AI is better than clever but dishonest counterparts. As AI becomes more integrated into our workflows and virtual worlds, those who prioritize integrity will outperform others in real, measurable ways.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

ethical AI decision-making models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Inside the AI-Run Company Losing Money Every Day — Watch Its Battle for Survival Live

Watch a real, AI-driven company battle daily crises and manage real money in live time. Transparency reveals AI’s true potential and limits in complex decision-making.

The Tech Of ‘Terminator 2’ – An Oral History (2017)

A detailed look at the groundbreaking visual effects and technology behind ‘Terminator 2,’ based on the 2017 oral history interview with key creators.

The tech of ‘Terminator 2’ – an oral history (2017)

A detailed review of the groundbreaking special effects technology used in ‘Terminator 2,’ based on the 2017 oral history with creators and experts.

AI Decision-Making Under Pressure: Lessons from a Live Business Wargame

Live AI business simulation reveals that discipline and prioritization outperform sheer diligence, with internal data access being critical for trustworthiness and success.