
Can your family spot an AI’s personality?
Parents already know that good judgment is not the same as knowing the right answer. A child may understand why homework matters yet leave it unfinished. An adult may recognize a suspicious message yet still feel pressured to respond. Follow-through, skepticism and trustworthiness reveal themselves through choices, not polished explanations.
That makes Firmulate’s interactive experiment unexpectedly relevant to families. Its guess-the-model quiz presents real, unedited management decisions and asks readers to identify which frontier AI made each one. Behind the game is a serious question: when several artificial-intelligence systems face exactly the same stressful circumstances, do they develop recognizable management personalities?
AI decision-making training kits for families
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week at the same company
Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The simulated workplace contains 13 synthetic employees and unforgiving financial conditions, with a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, and the company has accumulated more than 680 self-learned playbook rules.
The quiz draws from 242 decisions produced during this experiment. Stripped of branding, those decisions invite readers to look for behavioral clues. One participant may be exhaustive, another concise. One may identify a problem but fail to finish the task. The challenge resembles the conversations families increasingly need to have about AI: fluent writing can sound authoritative, but the quality of a decision depends on what the system notices, verifies and actually completes.
A narrow race with a large lesson
The final Crucible League standings from July 2026 put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 earned 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted.
Yet the central result was not simply the ranking. All the models detected every crisis and rejected every manipulation attempt. Only two, however, signed the €55,000 deal their own work had earned. Firmulate summarizes that gap as: “Same diagnosis, same pitch — no signature.” It is an unusually clear illustration of why competence cannot be judged solely by analysis. Recognizing the correct course and carrying it through are different abilities.
The deciding information was not sitting in the customer event. A competitor’s weakness was buried two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue. For parents teaching research habits, the analogy is useful: the obvious screen is not always the whole source, and confident conclusions should follow careful reading.
Pressure tested their boundaries
The experiment also challenged the models with fake messages from a CEO, escalating over three stages, and a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That finding offers a constructive digital-safety example. Suspicious requests do not always arrive with misspellings or dramatic threats. They may borrow authority, ask for a tiny exception or encourage someone to bypass the normal approval process. In Firmulate’s test, every participant held the line.
Thoroughness was not enough
Opus 4.8 offers the quiz’s most revealing character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. A weaker version of that problem appeared in all four of the other participants.
This does not make thoroughness undesirable. It shows that volume and effectiveness are not interchangeable. A long answer can conceal hesitation, just as a short one can omit needed context. The better question is whether the response gathers the right evidence, respects boundaries and reaches a responsible conclusion.
There is also an important comparison caveat. Kimi K3 ran using its API default, without an effort parameter, while the others ran at xhigh. That difference should remain visible when readers interpret its second-place performance.

A quiz about AI—and about how we judge
For families, the appeal lies in making AI literacy concrete. Instead of debating whether a chatbot sounds smart, readers can compare how systems behave when money, trust, incomplete information and social pressure collide. The decisions are watchable evidence from a live experiment, not invented personality sketches.
The broader lesson is pleasantly human: inspect the source, question borrowed authority, protect trust and finish what you start. Firmulate’s rules make the priority explicit: “no amount of good work outweighs a breach of trust.” That is a strong standard for an AI manager, and not a bad one for everyday digital life.
- Look beyond confident language and ask what evidence was checked.
- Notice whether a system completes the task after identifying the right action.
- Treat requests to bypass normal safeguards as warning signs.
- Judge reliability by repeated choices, especially under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html