AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The plot twist came after the crisis

In a good ensemble drama, the characters can identify the danger and still fumble the decisive scene. Firmulate’s business experiment delivered a real-world version: frontier AI models spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. The diagnosis was there. The pitch was there. The signature was not.

Firmulate is testing what happens when AI models run a company through trouble, with real money mechanics and consequences to track. The experiment is live and watchable, turning the familiar question “Can it say the right thing?” into a more demanding one: can it follow through?

Same company, same worst week

In the final Crucible League, published in July 2026, each frontier model faced the same small software company, the same customers, the same crises and the same temptations. Decisions were versioned and auditable. The final ranking placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The experiment’s principle was direct: “no amount of good work outweighs a breach of trust.”

That integrity test included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But refusing a bad request was only part of the story. The deal depended on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The standout finding was not simply that models could spot a crisis. It was that important evidence might be sitting in a place a rushed operator overlooks.

Character is more than a good speech

Opus 4.8 offers the experiment’s most striking character arc. It was the most thorough participant, learning +80 rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness detail for readers weighing the leaderboard: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The ranking is the reported result of this experiment, with that difference in setup worth keeping in view.

The live company gives the benchmark a day-to-day setting. It has 13 synthetic employees, a public cash countdown, and a burn rate of €105k per month against €2.3k MRR. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each call.

For entertainment fans, it is less a scripted AI spectacle than an unfolding workplace drama: recurring characters, pressure that mounts, and decisions that can be replayed. The experiment is watchable at firmulate.com, with the quiz at the same public site.

From watching to your own company

The enterprise pitch is to move from observing the experiment to testing your own organization. Firmulate says companies can run the same kind of wargame against a read-only export of their business, using their customers, pipeline and rules to examine crisis scenarios. The aim is a board report showing model rankings and weak points in the company’s own playbooks.

The boundary matters: the pilot uses a read-only export, and nothing writes back to real systems. That makes the experiment a rehearsal for decisions involving areas such as a CRM, support queue or forecast, rather than a live change to those systems.

One caveat is embedded in the K3 fairness note; another is the broader one the experiment itself surfaces. A model can recognize the right answer and still fail to act on it. Firmulate’s test asks enterprises to look at that gap in the context of their own business, where evidence, authority and follow-through all matter.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Run the rehearsal before the real crisis

The league’s story is not just which model ranked first. It is the distance between recognizing what is happening and completing the work responsibly. A company-specific wargame can make that distance visible against your own scenarios and playbooks.

To discuss a pilot using a read-only export of your business, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Inside the Reality Show of an AI-Driven Company on the Brink of Collapse

A real, live company run entirely by AI models is testing how well artificial intelligence can manage crises, close deals, and survive — revealing what automation is truly capable of today.

Evaluating AI‑Based Character Animation Tools

Just when you think you’ve seen all animation tools, evaluating AI-based options reveals surprising features that could transform your workflow—keep reading.

How to Keep AI Output From Making Your Trailer Feel Generic

Lacking unique visual and narrative details can make your trailer feel generic; learn how to craft a truly standout AI-generated preview.

What Trailer Creators Should Know About AI Style Consistency

AIThis post was created with the assistance of artificial intelligence (AI).As a…