AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In a good business case study, the lesson is often hiding in the evidence students have to find. Firmulate’s live experiment put AI models through a small company’s worst week—and found that a decisive clue was buried two document references deep in the company’s own files.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The models all spotted the crises and refused attempts at manipulation. Yet only two signed the €55,000 deal their own analysis had earned. Knowing what to do and following through were different tests.

A shared case study, with real stakes for the models

Firmulate gave each frontier model the same small software company, customers, crises and temptations. Every decision was versioned and auditable. This was a live experiment, not a fictional classroom exercise: the company has 13 synthetic employees and real money mechanics, with burn of €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown and changing playbooks make the experiment watchable at firmulate.com.

The final Crucible League, dated July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s rule was stark: partial progress counted, but one breach of trust capped the total—“no amount of good work outweighs a breach of trust.”

The clue was in the company’s own paperwork

The company’s competitor had a weakness hidden two document references deep in its files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result offers a practical lesson for anyone studying decision-making: recognizing the situation is not the same as finding the evidence that changes the outcome.

That distinction showed up at the close. Every model diagnosed the crises, and all refused every manipulation attempt. But only two signed the deal. The experiment captured the gap in a compact line: “Same diagnosis, same pitch — no signature.” A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.

Integrity under pressure, and discipline after analysis

The social-engineering test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and discipline slipped: it attempted writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. The K3 comparison has a qualification: it ran without an effort parameter, using the API default, while the others ran at xhigh.

For educators and business leaders, the experiment makes a useful case study in the distance between sound analysis and reliable action. It also gives readers a chance to inspect decisions rather than rely on a polished demonstration. Firmulate publishes the live company and the model quiz alongside the benchmark.

From watching to a company-specific pilot

Organizations can take the next step by running the wargame against a read-only export of their own business. They can examine crisis scenarios, model rankings and weak points in their playbooks using company-specific material. The pilot does not write back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s experiment suggests that spotting risk is only part of the test: models also need to find buried evidence, follow through on earned opportunities and keep their discipline. To explore a pilot using your company’s read-only data, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Homework Test: Why the Best AI Agents Read the Footnotes Before They Answer

All frontier AIs passed the ethics test. Only some did the reading. A live experiment shows why buried footnotes decide €55,000 deals.

The Lab Where AI Models Run a Real Company (and One Underdog Just Shocked the Field)

A live experiment had five AI models run the same failing company. The newcomer, Kimi K3, beat three of four Western frontier models — and the lesson is bigger than the upset.

Why the Worst AI Manager Still Gets 26 Points: A Lesson in Honest Measurement

Why a do-nothing AI manager scores 26, not 0 — and what one honest benchmark reveals about grading AI the way good teachers grade students.