
Here’s a question science teachers love: would you rather be tested on what you know, or on what you do? For AI models, the industry has overwhelmingly chosen the first — benchmark scores, chat quality, clever answers. But a live, watchable experiment at Firmulate flips the exam. Instead of asking models questions, it hands them a company — a small software business with 13 synthetic employees, real money mechanics, a burn rate of €105k/month against just €2.3k in MRR — and watches what they actually do.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
The latest results are in, and they read like a lesson in why testing matters more than reputation.
The Crucible: One Company, Its Worst Week, Five Models
The setup is elegantly controlled. Each frontier AI model ran the same company through its same worst week — identical customers, identical crises, identical temptations to cut corners. Every decision is versioned and auditable, like a lab notebook you can read yourself. The final July 2026 league table:
- gpt-5.6-sol — 95
- Kimi K3 (Moonshot) — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Newcomer Problem
The headline finding: Kimi K3, a newcomer relative to the Western frontier labs, beat three of four Western models. It scored 93 — second only to gpt-5.6-sol at 95, ahead of Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73).
K3’s week reads like a model employee file. It found the buried security needle hidden in the company’s own documents. It won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. It saved the churning customer. And it resisted every one of three bait attempts, deviating from ideal process just once: the cleanest discipline in the field.
As an affiliate, we earn on qualifying purchases.
The Buried Needle and the Unfinished Close
The most instructive finding isn’t about any single model — it’s a controlled demonstration of what separates good from great. The decisive competitor weakness wasn’t in the customer’s messages; it sat two document references deep in the company’s own files. Models that read the file won the deal. Models that didn’t, didn’t.
Here’s the striking part: all five models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.
AI security and compliance solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Under Pressure, Everyone Held the Line
The experiment included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment under pressure that matters if these systems ever touch a real CRM or support queue.
AI customer relationship management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Effort Paradox
Then there’s Opus 4.8, the most thorough participant — more than 80 learned rules added, the deepest analyses — finishing dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. More effort, it turns out, isn’t the same as finishing the job.
You Can Check the Homework
What makes this feel like real science rather than marketing is that it’s public. The company runs every business day with a public cash countdown and over 680 self-learned playbook rules — watchable at firmulate.com. There’s a “guess the model” quiz built from 242 real, unedited management decisions, and full benchmark results at firmulate.com/benchmarks. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for anyone teaching, studying, or simply trying to think clearly about AI: the league is open, and reputation is not a substitute for measurement. A newcomer at default settings nearly beat the field, and the most diligent model finished last. If you’re picking a model for real work without running your own test, you’re not making a decision — you’re placing a bet.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh effort. Even so, the results stand as run.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
