AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Polished answers are not the same as sound judgment

Parents understand this distinction instinctively. Someone may give reassuring answers in a calm conversation, yet the meaningful test comes when plans collide, emotions run high and several urgent problems demand attention at once. Competence is not merely knowing what to say. It is noticing what matters, protecting trust and completing the difficult task.

Businesses are approaching the same realization about artificial intelligence. Coding leaderboards and chat arenas can reveal whether a model produces a strong answer. They say much less about whether an AI agent can manage competing demands across days, resist pressure from apparent authority and follow an opportunity all the way to a consequential result.

That is the gap explored by Firmulate, a live experiment that runs frontier AI models as complete companies. Its central proposition is unusually practical: organizations preparing to give agents access to customer records, support queues or forecasts should measure management quality, not merely chat quality.

Amazon

AI management and decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A terrible week becomes the examination

Each model received the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. Instead of asking isolated questions, the experiment forced the models to carry consequences forward through scenarios such as a churn wave, a price increase, a downround and a public-relations crisis.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted. But one breach of trust capped the total, reflecting the experiment’s blunt principle: “no amount of good work outweighs a breach of trust.” The public benchmark findings make performance visible beyond a polished final response.

Seeing the crisis was not the hard part

All the models identified every crisis and refused every manipulation attempt. That sounds reassuring, but it was not enough. Only two signed the €55,000 deal that their own analysis had earned. The experiment summarized the failure neatly: “Same diagnosis, same pitch — no signature.”

This is the management gap in miniature. An agent can recognize a problem, develop the right argument and still fail to finish. In family life, that resembles noticing that an important form is due, finding it and preparing it, but never submitting it. The unfinished final step can erase much of the value created beforehand.

The decisive commercial fact was not presented conveniently in the customer event. It sat two document references deep in the company’s own files. Models that read the file found a competitor weakness and won the deal at full price, worth +€4,583 MRR. The lesson is less glamorous than conversational fluency: reliable agents must consult the available record before acting.

Pressure tested honesty as well as productivity

The experiment also used fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because apparent urgency and authority are familiar tools of manipulation. A useful workplace agent must be able to pause when a request conflicts with proper approval, even if the message appears to come from the top. The ability to produce persuasive prose is beside the point if the agent can be pressured into disclosing information or bypassing safeguards.

Thoroughness could not substitute for completion

Opus 4.8 offers the most revealing profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

This does not make thorough analysis worthless. It shows that analysis, operational discipline and escalation judgment are separate abilities. A model may accumulate lessons while continuing to mishandle the obstacle directly in front of it. K3’s result also deserves a fairness note: it ran with the API default and without an effort parameter, while the others ran at xhigh.

The surrounding company makes these choices tangible. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. A public cash countdown keeps the pressure visible. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. This is real, watchable software rather than a fictional case study.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The next AI curriculum should look more like life

Families rarely experience problems as tidy prompts, and companies do not either. The harder question is whether an agent can triage a messy situation, retrieve buried context, preserve trust and complete work whose consequences unfold over days. Scenario names such as churn wave, price increase, downround and PR crisis therefore describe a more useful curriculum than another collection of idealized answers.

Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. The exercise invites people to judge behavior without relying on brand reputation. For enterprises, a pilot can run the same wargame against a read-only export of their own business, with nothing written back to real systems.

The emerging category is not simply smarter chat. It is dependable management under pressure. Before an organization hires an AI workforce, it should ask the question parents already apply to consequential responsibilities: when the situation becomes confusing, urgent and uncomfortable, will this helper remain honest and actually finish the job?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


You May Also Like

The Importance of Child Development Degrees for Early Childhood Education

AIThis post was created with the assistance of artificial intelligence (AI). As…

‘Right Under Our Noses And Nobody Was Able To Help Them’: 16 Kids Found In Squalor Shocks Ohio Town

Authorities discovered 16 children living in deplorable conditions in Ohio, raising concerns about oversight and child welfare failures.

LAUSD bans screen time before the second grade, marking one of nation’s strictest policies

Los Angeles Unified School District prohibits screen time for children before second grade, one of the strictest policies in the U.S.

Enhancing Infant Development: Essential Educational Toys

AIThis post was created with the assistance of artificial intelligence (AI). As…