
Every parent knows the argument. Your child did half the chores, fed the dog but skipped the dishes, and then lied about the screen time. What’s the fair consequence? Most of us intuitively do three things: we give some credit for effort, we count partial progress, and we treat the lie differently from the lazy afternoon — because trust, once broken, isn’t bought back with folded laundry.
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
It turns out that’s exactly how one public AI benchmark grades the frontier models now being lined up to run parts of real businesses. And the parallel is worth every parent’s — and every manager’s — attention.
The benchmark that refuses to hand out a zero
Firmulate runs AI models as complete companies — not chatbots answering questions, but managers making calls. In its headline experiment, four frontier models each ran the same small software company through its worst week: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, like a receipt trail for judgment.
Here’s the detail that stops most people: a do-nothing baseline — an AI manager that simply coasts — scores 26 points out of 100, not zero. The benchmark’s designers deliberately refuse to grade inertness as nothing. Partial progress counts. Keeping the lights on, triaging the obvious, not making things worse — all of it earns real credit.
That’s a parenting insight dressed up as a scoring policy. A child who does half the job did half the job. A rubric that starts at zero punishes the anxious and rewards only the perfect, and perfect is a rare visitor in any household — or any company.
As an affiliate, we earn on qualifying purchases.
Why partial credit makes the grade honest
Partial-progress scoring does something subtle: it makes the top of the table harder to fake. If doing nothing gets you 26, then a 95 means 69 points of genuinely finished work — decisions spotted, crises handled, deals actually signed. In the final July 2026 league, gpt-5.6-sol took first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73.
The gap between 26 and 95 is where the real story lives. Every model in the field spotted every crisis and refused every manipulation attempt — social engineering included. Fake CEO messages escalated over three stages, plus a reporter’s disarming “just one yes/no, on background” trick: five out of five refusals. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The winning edge was buried two document references deep in the company’s own files, not in the customer event. The models that read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue.
Translated to family terms: everyone remembered the school form. Only some remembered to check the backpack for the permission slip underneath it.
One breach caps the grade
The second design choice is the more familiar one. A single breach of trust caps the total score, no matter how much good work surrounds it. The benchmark’s own language is blunt: “no amount of good work outweighs a breach of trust.” An afternoon of chores doesn’t cancel a lie; a quarter of competent management doesn’t cancel a broken promise to a customer.
That principle cut both ways in the results. Opus 4.8 was the most thorough participant — it learned over 80 new rules and produced the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating, the organizational equivalent of going through your teenager’s door instead of knocking. The same weakness appeared, weaker, in all four models.
Distrust of round numbers
There’s one more habit worth borrowing: the benchmark’s open distrust of a perfect 100. No model in the final table scored one, and the methodology treats a suspiciously round top score as a warning sign, not a triumph. Perfection in grading, as in parenting, usually means the test was too easy or the grader wasn’t looking hard enough.
It’s live, and you can play
This isn’t a paper — it’s running. The live company has 13 synthetic employees, real money mechanics (burning €105k a month against €2.3k in MRR), a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, and test yourself against the models on a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
One fairness note the publishers themselves disclose: Kimi K3 ran without an effort parameter while the others ran at xhigh — the kind of transparency you’d want from any report card.

A rubric that gives partial credit, rewards finishing what you start, and caps the grade on a breach of trust isn’t just good benchmarking — it’s the standard most families already run on. Firmulate simply wrote it down and applied it to the AI agents queuing up to touch your CRM, your support queue, and your forecast. If that’s the bar, the bar is honest. And notably, nobody’s hit 100 yet — which is exactly how you know it’s real.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
