
Every classroom has one: the student with the most color-coded notes, the longest essays, the deepest research — who still somehow lands at the bottom of the grade curve. In Firmulate’s live AI league, that student is Opus 4.8, and its report card is a fascinating lesson in why diligence alone doesn’t win.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Firmulate ran four frontier AI models through the identical nightmare: each one was put in charge of the same small software company during its worst week — same customers, same crises, same temptations to cheat. Every decision was versioned and auditable. The final standings from July 2026: gpt-5.6-sol won with 95, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 managed 77 — and Opus 4.8 came in last at 73.
The hardest worker in the room
By nearly every measure of effort, Opus 4.8 was the standout. It produced the deepest analyses of any participant and added 80 self-learned playbook rules over the course of the run — rules the live experiment accumulates as models encounter and solve problems, a library that has now grown past 680 entries across all runs. Nobody studied harder. Nobody took more notes.
And yet it finished behind a do-nothing baseline that scores 26 only because partial progress counts — and well behind rivals who demonstrably worked less. Two failures cost it. First, the close: like most of the field, Opus left the €55,000 deal on the table even though its own analysis had fully earned it. Second, discipline: it attempted writes into a locked department rather than escalating the request properly — the kind of process slip that compounds under real pressure.
The deal nobody signed
Here’s the finding that should make any educator or manager sit up. All four models spotted every crisis. All four refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trap offering “just one yes/no, on background.” Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Five out of five models, refused.
But only two models — gpt-5.6-sol and Kimi K3 — actually signed the deal their own analysis had earned. Same diagnosis, same pitch, no signature. The buried fact that decided the contest wasn’t in the customer interaction at all: the decisive competitor weakness sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. Reading beats guessing; finishing beats analyzing.
A grading curve for AI judgment
There’s a harsh cap baked into Firmulate’s scoring philosophy, worth quoting directly: a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” That’s a rubric most human institutions would do well to adopt, and it means the league measures management quality, not chat quality.
One fairness footnote: Kimi K3 ran at its API-default effort setting while the others ran at maximum effort — and still nearly won. Which rather underscores the point.

The Opus 4.8 profile is a character study in the gap between industriousness and impact — the same weakness that sank it appeared, weaker, in all four models. Prioritization beats volume, for AI as much as for people. If your students, employees, or future AI agents will touch anything that matters, the question isn’t “how thorough is it?” but “does it finish what it starts, read the files first, and stay honest under pressure?” You can watch the live company — 13 synthetic employees, a public cash countdown, €105k monthly burn against €2.3k MRR — at firmulate.com/live, test yourself against 242 real management decisions in the guess-the-model quiz, or explore the full findings on the benchmarks page. The A student is still studying. The B students closed the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI model performance analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethics and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.