firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Parents know the difference between a child explaining what they would do and actually handling a hard moment. The same gap matters when a business puts AI to work: a model may spot trouble and recommend the right move, then fail to follow through. Firmulate’s experiment puts that gap on display—and offers companies a way to rehearse with their own business data before trusting AI with real operations.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A company’s worst week, on repeat

Firmulate ran frontier AI models through the same small software company’s worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The live company has 13 synthetic employees and real money mechanics, including a burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown and more than 680 self-learned playbook rules make the experiment watchable at Firmulate.

The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league’s standard is blunt: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Knowing the answer did not guarantee the sale

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing the right move and carrying it out is captured in the experiment’s line: “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a reminder that an agent can miss an opportunity even when the clue belongs to the business’s own information.

The pressure tests included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusing a risky request matters; so does knowing when to escalate a legitimate action.

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail to keep in view when comparing the standings.

From watching to rehearsing

For a family, a rehearsal can expose what a child remembers under pressure. For an enterprise, the equivalent is testing AI against the company’s own customers, rules and failure scenarios. Firmulate’s proposed pilot starts from a read-only business data export and runs crisis scenarios against that company. The output is a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

That boundary makes the exercise easier to evaluate: leaders can see how a model handles a crisis, a manipulation attempt or a valuable lead without giving it permission to alter live records. Firmulate’s public site also offers a “guess the model” quiz built from 242 real, unedited management decisions—a way to test whether readers can tell which model made a call.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the judgment to a test

AI readiness is more than producing a convincing answer. It means finding the evidence, making the call, respecting boundaries and following through. Firmulate’s live experiment shows how those behaviors can diverge. To wargame your own business using a read-only export, explore the Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Prince George Eton College

Prince George begins a new educational chapter at Eton College’s preparatory program, marking a significant step in his schooling journey.

The Montessori Method at Home: Boosting Independence in Toddlers

Inevitably, implementing Montessori principles at home can transform your toddler’s independence—discover how to foster confidence and self-reliance today.

Nanit Smart Baby Monitor Review: Worth It for September Nights?

An honest hands-on review of the Nanit Smart Baby Monitor with floor stand and 8-inch display — sleep tracking, HD video, and the real trade-offs.

‘Ketchup Kids’ Help Principal Find True Calling

A group of students known as ‘Ketchup Kids’ played a key role in guiding their principal toward a new career path, sparking widespread interest.