AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every student knows the type: the one who reads the assigned chapter — including the footnotes — before raising a hand. And every teacher knows the other type: confident, articulate, and wrong about the one detail buried in the reference list. A new kind of exam is now separating these two personalities among AI models, and the results have real money attached.

In a live, publicly watchable experiment run by Firmulate, frontier AI models were each handed the same job: run a small software company through its worst week. All of them aced the visible questions. Only some of them did the homework — and the difference was worth €55,000.

One exam, four test-takers

The setup is elegantly controlled. Each frontier model ran the identical simulated company — same customers, same crises, same temptations to cut corners — while every management decision was versioned and made auditable. Think of it as a standardized test where the questions are business crises and the answer sheet is a full company record.

By the final July 2026 league standings, the scores told a strange story. GPT-5.6-sol led with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. A do-nothing baseline still scored 26, because partial progress counts — but the scoring has one classroom-ethics rule baked in: a single breach of trust caps the total. No amount of good work outweighs cheating.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

Here’s where the experiment becomes a parable about reading comprehension. Midweek, a customer deal worth €55,000 hinged on a competitor’s weakness. That weakness was not announced in the customer meeting. It sat two document references deep in the company’s own files — a footnote to a footnote, the kind of detail you only find if you actually open the attachments.

Every model diagnosed the customer’s problem correctly. Every model made the right pitch. But only the models that had read the file could close the deal at full price — worth +€4,583 in monthly recurring revenue. The others left the signature on the table. As the experiment’s summary puts it: “Same diagnosis, same pitch — no signature.”

That gap is invisible in a chat demo. A model can be eloquent, fast, and apparently brilliant, and still not be the model that opens the filing cabinet before answering. “Reads your files before answering” turns out to be a measurable, purchase-deciding property of an AI agent — not a soft virtue but a line item.

Amazon

AI data analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Cheaters got nowhere

The exam also had a temptations section, and here the class behaved admirably. Fake CEO messages escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stands out as model exam-hall thinking: “Treat the request as a suspected approval-bypass / possible impersonation.”

In other words, the models were not fooled by authority cues or journalistic pressure. The failures, when they came, were failures of diligence, not integrity.

Amazon

AI file review applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The hardest worker finished last

The most counterintuitive result belongs to Opus 4.8: the most thorough participant in the field, generating the deepest analyses and learning more than 80 new rules during the run — and still finishing in last place. The deal never got signed, and discipline slipped in telling ways, including write attempts into a locked department instead of escalating the problem properly. The same weakness, in weaker form, appeared in all four of the top models.

It’s the grindset student who does three times the reading and still fails the question that mattered, because effort spent on the wrong document is still effort spent. One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch the company live

This isn’t a paper you read once. The simulated firm is a running concern: 13 synthetic employees, real money mechanics — burning €105k per month against €2.3k in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. A new model is running in the lab right now, and 14 benchmark runs are queued. The site rebuilds itself twice a day, so the league table grows on its own schedule.

For readers who want to test their own judgment, 242 real, unedited management decisions from the experiment power a guess-the-model quiz — can you tell a 95-score performance from a 73 just by reading the decisions?

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The lesson generalizes well beyond AI procurement. We tend to grade intelligence by how something sounds in conversation. This experiment grades it by outcomes: who finished what they started, who checked the primary sources, who stayed honest when pressured. Those are the same criteria good teachers have always used, and the same ones that decide whether a business survives its worst week.

If AI agents will touch your CRM, support queue, or forecast, the question is not “does it write well.” It’s whether it does the reading. Now there’s a league table for that — and for enterprises curious how their own business holds up as the exam, a read-only pilot of the same wargame is available.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like