AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

When a family business lets AI take on more work, the stakes reach beyond the office

Parents know that trust is built through the small choices: how a child handles a mistake, whether they follow through, whether they tell the truth when it would be easier not to. Businesses face a version of that question as AI tools begin to handle customer conversations, forecasts and other consequential work. A polished answer is not enough. Can an AI system spot trouble, resist pressure and carry a good decision through?

Firmulate’s live experiment puts that question to the test. Its small synthetic software company runs with 13 synthetic employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, alongside a public cash countdown. The company is watchable at firmulate.com.

One company, the same hard week

In the final Crucible League, published in July 2026, frontier models each faced the same small software company, the same customers, the same crises and the same temptations. Their decisions were versioned and auditable. The contest ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s trust rule was plain: “no amount of good work outweighs a breach of trust.”

The headline finding was less about spotting danger than acting on what the analysis revealed. Every model identified every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

The clue was buried in the company’s own files

The deal turned on a competitor weakness hidden two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That detail makes the trial feel less like a quiz with an obvious right answer and more like a test of whether an AI can connect information scattered across ordinary business records.

The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response this way: “Treat the request as a suspected approval-bypass / possible impersonation.” Those are reassuring results, especially for businesses weighing how much access to give automated assistants.

Thoroughness did not guarantee a finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and slipped in discipline by attempting writes into a locked department instead of escalating. A weaker version of that same weakness appeared across all four models. In other words, good diagnosis and careful analysis did not always translate into a completed, appropriately handled action.

There is also a fairness detail readers should know: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results describe this particular experiment, with that difference in setup; they are not a universal verdict on what any model will do in every company.

Firmulate also makes 242 real, unedited management decisions available through a “guess the model” quiz at firmulate.com/quiz.html. The live company’s workdays are versioned, and its playbook has accumulated more than 680 self-learned rules. Those features let readers watch decisions develop over time, rather than relying on a single polished demonstration.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to trying it on your own business

For families, the broader lesson is familiar: capability and trust belong together, and good intentions do not guarantee follow-through. For a business, that raises a practical question: how would an AI handle a crisis, a tempting shortcut or a deal hiding in its own files?

Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. It tests crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Tell if a Toy Supports Development or Just Entertainment

AIThis post was created with the assistance of artificial intelligence (AI).To tell…

The Impact of Hunger on Child Development: Cognitive, Emotional, and Physical Consequences

AIThis post was created with the assistance of artificial intelligence (AI). As…

Balancing Screen Time and Play Without Guilt

Loving a balanced childhood means mastering the art of balancing screen time and outdoor play, and here’s how you can do it effectively.

What does emotional maturity in children actually look like? A psychologist explains, age by age

A psychologist explains what emotional maturity looks like in children at different ages, clarifying common misconceptions and its importance for parenting.