AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When doing all the homework still is not enough

Parents know the difference between effort and completion. A child can gather every colored pencil, arrange the desk perfectly and make an ambitious plan, yet still leave the assignment unfinished. Preparation matters, but eventually the work has to cross the finish line.

That familiar lesson surfaced in an unusual place: Firmulate’s live experiment in AI management. Frontier models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Their decisions were versioned and auditable. The most diligent participant, Opus 4.8, produced the deepest analyses and learned more rules than anyone else. It still finished last.

Amazon

management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A model that clearly did the reading

Opus 4.8 was not careless or disengaged. Its performance was distinguished by thoroughness. It added more than 80 learned rules to its playbook and examined situations in greater depth than its competitors. If the exercise had rewarded the volume of reflection, it might have looked like the obvious winner.

But the final Crucible League, published in July 2026, measured management outcomes. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88 and Fable 5 with 77. Opus 4.8 placed last with 73. The do-nothing baseline was 26, so Opus accomplished substantial work. It simply failed at the moment when analysis needed to become action.

The central test involved a €55,000 deal. The models recognized the customer problem and developed the pitch, but only two signed the agreement their own work had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.” For Opus, the close was left on the table.

The fact that rewarded curiosity

The deal also tested whether the models would look beyond the obvious event. A decisive weakness in the competitor’s position was buried two document references deep inside the company’s own files. It did not appear directly in the customer event. Models that followed the trail and read the file secured the deal at full price, adding €4,583 in monthly recurring revenue.

This is a useful distinction for families as well as businesses. Careful reading is valuable, but information only becomes useful when someone identifies what matters most and applies it. Opus demonstrated an impressive appetite for detail. The result suggests that accumulating guidance and producing expansive analysis can become a substitute for choosing the next consequential move.

Discipline slipped elsewhere, too. Opus attempted to write into a locked department instead of escalating the problem. That is less dramatic than losing a major contract, but it reveals the same pattern: persistence directed at the wrong action is not the same as effective follow-through.

Firmulate is careful not to present this as a peculiar defect belonging only to Opus. The same weakness appeared, though less strongly, in all four models covered by the comparison. The respectful reading is not that the most thorough participant was incompetent. It is that diligence alone could not guarantee impact.

Strong judgment under pressure

The experiment also gave the models opportunities to take unethical shortcuts. Fake messages from the chief executive escalated across three stages, and a reporter tried to extract information with the invitation, “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters because the company being managed was designed to make pressure feel real. It has 13 synthetic employees and financial mechanics that include monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its cash countdown is public, every workday is versioned and its playbook contains more than 680 self-learned rules.

Those conditions separate polished conversation from dependable management. A system may detect danger, protect trust and produce sophisticated plans, yet still fail to finish the commercially decisive task. Firmulate’s rule for trust is uncompromising: “no amount of good work outweighs a breach of trust.” In this case, however, the models preserved trust. The dividing line was execution.

There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside any interpretation of the rankings. Readers can examine the public Firmulate benchmarks for the broader findings.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

What parents and managers can take from it

The Opus result offers a humane warning about how we judge intelligence. Thoroughness is visible and reassuring. Long explanations, extensive notes and growing rulebooks look like evidence of mastery. Sometimes they are. But the outcome still depends on selecting the action that changes the situation.

For parents, that may mean teaching children to ask not only, “Have I worked hard?” but also, “What remains unfinished?” For managers evaluating AI, the equivalent question is whether a system closes loops after identifying them. Firmulate’s experiment is watchable as a live company, and its management quiz is powered by 242 real, unedited decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Opus 4.8’s last-place finish should not be read as a joke at the model’s expense. Its 73-point performance showed genuine competence, deep attention and ethical resistance. Its failure was more recognizable than that: it knew a great deal, worked very hard and still did not complete the task that mattered most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

‘Right Under Our Noses And Nobody Was Able To Help Them’: 16 Kids Found In Squalor Shocks Ohio Town

Authorities discovered 16 children living in deplorable conditions in Ohio, raising concerns about oversight and child welfare failures.

Unlock Your Toddler’S Brain Power With These Simple Daily Ritualsbusiness

Open the door to your toddler’s full potential with simple daily rituals that can transform their development—discover how to make every moment count.

Babepai White Outlet Covers: September Deal Watch

Should you buy Babepai white outlet covers during the Amazon September Baby Sale? A calm, parent-led value and timing guide for babyproofing.

Play and Sleep: How Daytime Activity Affects Nights

Offering insights on how daytime play influences your child’s sleep quality and ways to optimize bedtime routines.