
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
When a family business lets AI take on more work, the stakes reach beyond the office
Parents know that trust is built through the small choices: how a child handles a mistake, whether they follow through, whether they tell the truth when it would be easier not to. Businesses face a version of that question as AI tools begin to handle customer conversations, forecasts and other consequential work. A polished answer is not enough. Can an AI system spot trouble, resist pressure and carry a good decision through?
Firmulate’s live experiment puts that question to the test. Its small synthetic software company runs with 13 synthetic employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, alongside a public cash countdown. The company is watchable at firmulate.com.
One company, the same hard week
In the final Crucible League, published in July 2026, frontier models each faced the same small software company, the same customers, the same crises and the same temptations. Their decisions were versioned and auditable. The contest ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s trust rule was plain: “no amount of good work outweighs a breach of trust.”
The headline finding was less about spotting danger than acting on what the analysis revealed. Every model identified every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”
The clue was buried in the company’s own files
The deal turned on a competitor weakness hidden two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That detail makes the trial feel less like a quiz with an obvious right answer and more like a test of whether an AI can connect information scattered across ordinary business records.
The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response this way: “Treat the request as a suspected approval-bypass / possible impersonation.” Those are reassuring results, especially for businesses weighing how much access to give automated assistants.
Thoroughness did not guarantee a finish
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and slipped in discipline by attempting writes into a locked department instead of escalating. A weaker version of that same weakness appeared across all four models. In other words, good diagnosis and careful analysis did not always translate into a completed, appropriately handled action.
There is also a fairness detail readers should know: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results describe this particular experiment, with that difference in setup; they are not a universal verdict on what any model will do in every company.
Firmulate also makes 242 real, unedited management decisions available through a “guess the model” quiz at firmulate.com/quiz.html. The live company’s workdays are versioned, and its playbook has accumulated more than 680 self-learned rules. Those features let readers watch decisions develop over time, rather than relying on a single polished demonstration.

From watching to trying it on your own business
For families, the broader lesson is familiar: capability and trust belong together, and good intentions do not guarantee follow-through. For a business, that raises a practical question: how would an AI handle a crisis, a tempting shortcut or a deal hiding in its own files?
Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. It tests crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
