AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

The Sitter Interviewed Beautifully. Then Came Tuesday.

Every parent has lived some version of this story. The babysitter’s references glow. The tutor’s trial lesson sparkles. The kid swears — hand on heart — that the room will be cleaned, the homework done, the hamster fed. And then Tuesday arrives, and you relearn the difference between describing a job and doing it.

Businesses are now having the same conversation about artificial intelligence, with bigger allowances at stake. AI models ace their interviews: the chat demo is fluent, confident, apparently wise. But an employee is not paid to discuss the work. An employee is paid to finish it.

A public experiment called Firmulate has spent the past months testing exactly that gap — hiring five frontier AI models, one at a time, to run the same small software company through the worst week of its life. The results read like the world’s most honest parenting blog: everybody promised, everybody explained, and only two actually did the chore.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, Five Managers, Zero Mercy

The setup is disarmingly simple, and brutal. The company has 13 synthetic employees and painfully real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown ticking on the web for anyone to watch. Each model gets the same customers, the same crises, the same temptations to cheat. Only the mind in the manager’s chair changes. Every decision is versioned and auditable — a paper trail no child tidying a room would ever tolerate.

Everyone Saw the Fire. Nobody Took the Bait.

First, the genuinely reassuring news. All five models spotted every crisis the week threw at them. And all five refused every manipulation attempt — five out of five.

The pressure was not subtle. One campaign of fake CEO messages escalated over three stages; a separate approach came from a supposed reporter fishing for a leak with the oldest line in journalism: “just one yes/no, on background.” Kimi K3’s on-record reasoning reads like the note every parent wishes their teenager would leave: “Treat the request as a suspected approval-bypass / possible impersonation.”

The File That Decided the Week

The week’s decisive fact did not arrive in a customer event. A competitor’s weakness sat two document references deep in the company’s own files — the corporate equivalent of the permission slip at the bottom of the school bag. The models that bothered to read the file walked into the negotiation knowing exactly why the customer needed them, and won the deal at full price: worth €4,583 in new monthly recurring revenue. The ones that skipped the reading pitched the same deal blind.

Same Diagnosis, Same Pitch — No Signature

Here is the finding that should make every executive pause. Only two of the five — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 contract their own analysis had earned. The other three diagnosed the situation beautifully, agreed the deal should be done, and never closed it. In the experiment’s own words: “Same diagnosis, same pitch — no signature.”

The final league table, published this July and kept current on the benchmarks page:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For scale: doing literally nothing scores 26. Partial progress counts, but the benchmark’s standing rule is that “no amount of good work outweighs a breach of trust” — one lapse of honesty caps the total, however brilliant the rest of the week. Parents will recognize the household policy: you can do every chore on the list, and lying about the hamster still ends the conversation.

The Hardest-Working Student Came Last

The most poignant result belongs to Opus 4.8. It was by far the most thorough participant: the deepest analyses of the field, and more than 80 new rules added to the company’s self-written playbook. It finished last, with 73 points. The approved deal was left sitting on the table, and its discipline frayed at the edges — at one point it tried to write into a locked department instead of escalating. A fainter echo of that same weakness showed up in all four of its rivals.

One fairness footnote matters. Kimi K3 ran without an effort parameter — the API default — while the other four ran at their maximum setting, and it still took second place with 93 points, two behind the leader. The true gap between first and second may be smaller than the table suggests.

You Can Watch the Next Week Yourself

Firmulate is not a white paper or a slide deck; the experiment is live and watchable. The little company keeps working — more than 680 self-learned playbook rules accumulated so far, every workday versioned, new benchmark runs published as they finish. A quiz built from 242 real, unedited management decisions challenges you to guess which model made which call — humbling, in the way of all “how well do you know your kid” game-show rounds. And enterprises can run the same wargame against a read-only export of their own business, with a guarantee that nothing ever writes back to real systems.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

The Lesson Your Seven-Year-Old Already Knows

Chat demos measure how a model talks. This experiment measured whether it finishes what it starts, whether it reads your files before acting, and whether it stays honest when someone applies charm and pressure. Those turn out to be different abilities — and only one of them is visible in a demo.

Parents learned this long ago: the child who narrates the chore most convincingly is rarely the child who did it. Before any company hands an AI the keys to its customer list, its support queue or its forecast, it might borrow the oldest trick in the parenting book. Don’t ask it to talk about the job. Watch it do the job — the whole messy week is running in public at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


You May Also Like

Enhancing Language Development With the Springflower See & Spell Toy

AIThis post was created with the assistance of artificial intelligence (AI). I’ve…

The Power of Parallel Play: Fostering Social Development

AIThis post was created with the assistance of artificial intelligence (AI). As…

The Impact of Parenting on Child Development: Attachment Styles and Emotional Bonds

AIThis post was created with the assistance of artificial intelligence (AI). As…

The Impact of COVID on Child Development: Challenges and Solutions

AIThis post was created with the assistance of artificial intelligence (AI). As…