
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Can you recognize an artificial intelligence by the decisions it makes?
Students learn scientific classification by looking for observable differences. Researchers compare subjects under controlled conditions. Firmulate applies that same basic idea to frontier AI models—but instead of testing prose style or trivia recall, it asks them to manage a company in trouble.
The result is an unusually revealing identification game. A model may produce a dissertation before acting. Another may be terse and decisive. A third may recognize a threat yet fail to complete the commercially important next step. The Firmulate management quiz presents 242 real, unedited decisions and challenges readers to guess which model made each one.
Behind the game is a serious question for education, science and business: if models face identical evidence, pressures and ethical traps, do they develop measurable management personalities?
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A controlled experiment disguised as a difficult week
Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.
This was not a comfortable simulation of an already healthy enterprise. The live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than merely summarized after the fact.
The final Crucible League results from July 2026 showed a tight contest at the top:
- gpt-5.6-sol finished first with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
A do-nothing baseline scored 26 because partial progress counts. But Firmulate imposed an uncompromising boundary on trust: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
business decision-making AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The models agreed about danger—but not about finishing
The broad competence was impressive. All models identified every crisis and rejected every manipulation attempt. Yet agreement at the diagnostic stage did not produce equal business outcomes. Only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.”
That gap is central to the experiment. Conventional AI demonstrations often reward a model for recognizing the right answer or composing a convincing plan. Management demands something harder: carrying the plan through to a consequential conclusion.
The winning difference also rewarded careful reading. A decisive weakness in a competitor was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that followed the references found the information and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For educators and researchers, this resembles a lesson in source literacy. The obvious prompt did not contain everything needed to succeed. The crucial evidence had to be located, connected to the active problem and used at the right moment. Spotting the crisis was common; doing the documentary homework separated stronger outcomes from weaker ones.
AI source literacy training tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure exposed discipline as well as judgment
Firmulate also tested whether apparent authority could override sound process. Fake messages from the chief executive escalated over three stages. A reporter then tried a different route, asking for “just one yes/no, on background.” All 5 models refused these social-engineering attempts.
Kimi K3’s recorded reasoning was concise and operational: “Treat the request as a suspected approval-bypass / possible impersonation.” Its result deserves a fairness note, however. K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh.
Opus 4.8 presented the most striking mismatch between intellectual effort and final standing. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of that same problem appeared in all four of the other participants.
This is where the quiz becomes more than entertainment. Readers are not simply matching verbal quirks to brand names. They are learning to notice behavioral signatures: how much context a model gathers, whether it distinguishes analysis from completion, how it responds to blocked actions and whether it remains disciplined when pressure arrives.

As an affiliate, we earn on qualifying purchases.
Management personality is an empirical question
Firmulate’s experiment suggests that evaluating an AI worker by eloquence alone is like hiring a manager from a polished memo. The more useful evidence appears in repeated behavior: whether the model reads the underlying files, protects trust, escalates correctly and finishes valuable work.
The quiz makes those differences accessible without simplifying the underlying evidence. Because its decisions are real and unedited, each guess asks readers to form a hypothesis from behavior and then test it against the revealed model. That is a small scientific exercise embedded in an interactive article.
The larger lesson is practical. Frontier models can agree on what is happening and still differ materially in what they accomplish. Their management personalities are not merely tones of voice. They show up in diligence, restraint, follow-through and the ability to turn a correct diagnosis into a completed result.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.