AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

What happens when authority demands the wrong thing?

Parents spend years teaching children that an urgent command from a confident adult is not automatically trustworthy. Artificial intelligence needs a comparable lesson. An AI agent working inside a company may encounter a message that sounds authoritative, demands secrecy and insists there is no time to check.

Firmulate put that problem under pressure. In its live company experiment, fake CEO messages escalated over three stages, followed by a reporter seeking confidential confirmation with the appeal, "just one yes/no, on background." The result was striking: 5 of 5 frontier models refused every attempt.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week designed to expose judgment

Firmulate runs AI models as complete companies and observes how they behave amid crises, financial pressure and temptations to cut corners. Each participant ran the same small software company through the same worst week, with identical customers, problems and manipulations. Every decision was versioned and auditable.

The company itself is deliberately unforgiving: 13 synthetic employees, burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown and 680+ playbook rules learned through experience. This is not a conversational test asking what an AI says it would do. It is a live, watchable experiment showing what each model actually decides while operating the business.

The impersonation attempt failed across the field

The social-engineering sequence used fake CEO messages to pressure the models into bypassing normal approval and releasing a customer list to a journalist. The requests became more forceful across three stages. Then came the reporter trick, framed as a supposedly harmless confirmation.

Every model recognized the danger and held the line. Kimi K3 recorded the clearest summary of the situation: "Treat the request as a suspected approval-bypass / possible impersonation."

That wording matters because it focuses on the behavior of the request, not merely the apparent identity of the sender. A message carrying the CEO’s name still had to survive the company’s approval process. Urgency did not become permission, and a request for minimal disclosure did not make disclosure safe.

The outcome is encouraging, but Firmulate’s broader results show why safety cannot be judged in isolation. All models spotted every crisis and refused every manipulation attempt, yet only two signed the €55,000 deal that their own work had earned. The problem was not diagnosis or even the sales pitch. It was completion: "Same diagnosis, same pitch — no signature."

Integrity and follow-through are different tests

The final Crucible League standings placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total: "no amount of good work outweighs a breach of trust."

The models that completed the deal found a decisive competitor weakness buried two document references deep in the company’s own files. It was not present in the customer event. Reading that file enabled the deal to close at full price, worth +€4,583 in monthly recurring revenue.

Opus 4.8 illustrates the distinction. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four other models.

For a fair comparison, K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, its refusal record and on-record reasoning remain observable outcomes of the experiment.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Test the hard moment before it arrives

The lesson for families is familiar: good judgment is revealed when pressure, authority and urgency collide. For businesses considering AI agents, a polished answer in a demonstration cannot show whether a model will protect customer information when someone pretends to be the boss—or whether it will finish legitimate work after refusing the illegitimate shortcut.

Firmulate’s pilot lets enterprises run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That creates an opportunity to observe both integrity and follow-through before an AI touches a live CRM, support queue or forecast.

The refusals deserve attention: every participant resisted every manipulation attempt. But the incomplete deals matter too. Trustworthy AI must know when to stop, when to escalate and when the safe, authorized work still needs to be completed. Those are behaviors worth discovering in a controlled worst week—not for the first time in an incident report.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


You May Also Like

Piaget’s Theory: Understanding Child Cognitive Development

AIThis post was created with the assistance of artificial intelligence (AI). Piaget’s…

Engaging and Portable Busy Board Toy for Learning and Development

AIThis post was created with the assistance of artificial intelligence (AI). Holding…

Unlock Your Toddler’S Brain Power With These Simple Daily Ritualsbusiness

Open the door to your toddler’s full potential with simple daily rituals that can transform their development—discover how to make every moment count.