firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For playersOffer from Amazon

Play games on Amazon Luna for your game nights, included with Prime

  • A rotating selection of games, no download needed
  • Play on TV, laptop or phone
  • Fast, free delivery for your gear, too
Start playing with Prime Free trial · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Game masters know the trouble starts when the rules meet the players

A tabletop campaign can look airtight until a player finds the loophole nobody planned for. Companies face a similar test as AI moves from answering prompts to making decisions: what happens when an agent encounters pressure, conflicting priorities and a tempting shortcut? Firmulate has built a live experiment around that question, putting AI models in charge of a small software company through a simulated worst week.

One company, the same hard week

The experiment gave each frontier model the same customers, crises and temptations. Decisions were versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The headline result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The divide came at the finish: only two signed a €55,000 deal their own analysis had earned. As Firmulate puts it, “Same diagnosis, same pitch — no signature.” In a game, a character can identify the boss’s weakness and still fail to land the decisive move. In a business, that gap can mean a missed deal.

The clue was in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. It is a useful reminder for anyone evaluating AI agents: knowing how to respond to the visible prompt is only part of the challenge. The relevant clue may sit elsewhere in the company’s own information.

Integrity faced its own staged encounter. Fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of judgment a company needs to see under pressure, not just in a polished demo.

Thoroughness did not guarantee a win

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. More analysis, on its own, did not ensure follow-through.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. The exercise shifts the reader from leaderboard watching to judging individual choices.

A live company, then a company-specific pilot

Firmulate’s live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. Readers can watch the experiment at Firmulate.

The next step is to run a similar wargame against a company’s own business. Firmulate says an enterprise pilot can use a read-only export, test crisis scenarios against the company’s customers, pipeline and rules, and produce a board report with model rankings and weak points in existing playbooks. Nothing writes back to real systems. That boundary matters: leaders can examine how models handle consequential situations before giving them access to live workflows.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Watch the decisions before you delegate them

The live experiment makes AI behavior visible through choices, missed opportunities and attempts to cross a line. A company-specific pilot carries that scrutiny into an organization’s own scenarios, using its data in read-only form. Enterprises interested in running the wargame against their business can visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: Jev Plays Pokémon Red

Search and coverage interest is rising around a Show HN project title, but the reason for the spike and the project’s full details are unconfirmed.

The tech of ‘Terminator 2’ – an oral history (2017)

A detailed review of the groundbreaking special effects technology used in ‘Terminator 2,’ based on the 2017 oral history with creators and experts.

Can AI Managers Outsmart Crises? Watch Models Battle in Real Business Test

Watch AI models run a real company through crises and decision-making, revealing their management personalities—are they thorough, disciplined, or slip under pressure?

Inside the AI-Run Company Losing Money Every Day — Watch Its Battle for Survival Live

Watch a real, AI-driven company battle daily crises and manage real money in live time. Transparency reveals AI’s true potential and limits in complex decision-making.