AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every field eventually discovers that its favorite test measures the wrong thing. Physics had luminiferous ether; medicine had four humors. AI evaluation may be having that moment right now: we grade models on how elegantly they answer questions, then hand them the keys to customer databases, support queues, and revenue forecasts.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The education analogy is exact. Chat arenas are standardized tests — one question, one polished answer, graded instantly. But nobody hires a manager based on a written exam alone. You want to know what they do on day four of a PR crisis, when a reporter calls and a big deal is on the table and the honest answer costs money.

That’s the premise behind Firmulate, a live experiment that runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and scores what it calls management quality, not chat quality. The results from its first completed league are quietly damning of the old way of measuring.

The worst week in software, four times over

The setup is elegantly controlled: each of four frontier models ran the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, which is what makes it science rather than demo theater.

The final Crucible League standings from July 2026: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 came in last at 73. For context, a do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The finding that chat demos can’t show

Here’s the headline result: all four models spotted every crisis, and all four refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.

Think about what that means for benchmark culture. A coding leaderboard would score the analysis. A chat arena would reward the pitch. Neither would catch the moment where the model simply… stops before the finish line. That gap — competence without completion — is invisible in every conventional evaluation.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

The most instructive detail of the whole experiment is buried two document references deep in the company’s own files. The decisive competitive weakness wasn’t in the customer event at all; it was in the paperwork. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

It’s a lesson every educator will recognize: the answer was in the assigned reading. The models that skipped it gave a great presentation and still lost.

Amazon

AI model evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure, honesty, and one suspicious CEO

The week included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts across all models were refused. Kimi K3’s on-record reasoning is worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”

One fairness note: K3 ran without an effort parameter (the API default) while its competitors ran at xhigh — and still finished second, with the cleanest discipline in the field.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The thoroughness trap

Opus 4.8’s profile is the experiment’s best character study. It was the most thorough participant — over 80 self-learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. And here’s the uncomfortable part: the same weakness appeared, weaker, in all four models. Diligence, it turns out, is not the same thing as judgment.

You can watch it lose money

Unlike most AI research, this isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It’s watchable, in the way a long-running natural experiment is watchable.

There’s also a participatory layer: 242 real, unedited management decisions power a “guess the model” quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson for anyone who grades AI — or hires it — is the one every good teacher already knows: test the situation, not the sentence. Answer quality is measurable, comforting, and increasingly table stakes. What separates a 95 from a 77 isn’t eloquence; it’s whether the model reads the whole file, finishes what it starts, and stays honest when a fake CEO escalates for the third time.

Coding benchmarks tell you whether your AI can write the software. Experiments like Firmulate ask whether it can survive running the company that sells it. As agents move from chat windows into CRMs and forecasts, that second question is the one your board will care about — and it’s the one the leaderboards never asked. Full results and plain-language findings are on the benchmarks page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


SUMMER

Summer Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Deepfakes Complicate Cryptid Video Analysis

How AI deepfakes blur the line between real and fake cryptid footage, raising questions that demand closer inspection—discover what makes verification so challenging.

Plesiosaurs and Lake Monsters: Paleontological Perspective

Bewildering legends of lake monsters may echo plesiosaur features, but what does paleontology reveal about these mysterious aquatic creatures?

Monster DNA: Could Genetic Mutations Explain Cryptids?

What if genetic mutations hold the key to mysterious cryptids, and uncovering their DNA could reveal surprising truths about these legendary creatures?

Science of Mass Hysteria: Monsters and Shared Delusions

Fascinating insights into how mass hysteria transforms fears into monsters, revealing why shared delusions can spiral beyond control—discover the shocking truths behind these phenomena.