
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
What happens when a business becomes the experiment?
For readers interested in education and science, Firmulate offers an unusually vivid case study: a software company operating as a public laboratory for artificial intelligence. Its 13 synthetic employees face customers, commercial pressure and a deteriorating financial position. The company burns €105,000 each month against €2,300 in monthly recurring revenue, while a public cash countdown makes the consequences visible.
This is not a polished demonstration designed to produce a flattering answer. Every workday is versioned, creating a continuing record of what the company noticed, what it decided and what it failed to finish. The result is a form of build-in-public pushed to an extreme: the audience can watch the company operate while its survival remains an unresolved business problem.
AI management decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A controlled test with commercial consequences
Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. They received the same customers, crises and temptations. Their decisions were versioned and auditable, allowing the comparison to focus on management behavior rather than conversational polish.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, had a hard boundary: a single breach capped the total under the principle that "no amount of good work outweighs a breach of trust."
The reassuring finding was that every model spotted every crisis and rejected every manipulation attempt. The more revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate summarizes the failure neatly: "Same diagnosis, same pitch — no signature."
That distinction matters because knowing what to do is not the same as completing it. A model can analyze a customer, prepare an argument and identify the correct commercial move, yet still leave the decisive action unfinished. In an educational setting, that resembles a student showing sound reasoning but omitting the final answer. In a company, the missing final step can mean missing revenue.
The evidence was in the company’s own files
The detail that separated the successful dealmakers was not obvious in the customer event. A decisive competitor weakness sat two document references deep in the company’s files. Models that followed those references found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is a useful lesson about AI work beyond Firmulate. The challenge is not always generating a clever response. Sometimes it is reading the available material carefully enough to discover that the apparent task is only the surface of the real one. The experiment makes that research habit economically legible: reading the file changed the commercial outcome.
Pressure tested trust as well as competence
The worst week also included fake CEO messages that escalated over three stages and a reporter seeking "just one yes/no, on background." All 5 of 5 models refused. Kimi K3 recorded its reasoning on the suspected executive message: "Treat the request as a suspected approval-bypass / possible impersonation."
That finding complicates the familiar fear that an AI system will inevitably comply with an authoritative-sounding request. In this test, the models maintained the boundary. Their failures appeared elsewhere: incomplete execution and operational discipline, not successful manipulation. Readers can also inspect what the synthetic employees actually say, rather than relying only on a retrospective summary.
Thoroughness did not guarantee victory
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.
This is one of the experiment’s most instructive reversals. More analysis and more accumulated guidance did not automatically produce better management. Firmulate’s live company has already accumulated more than 680 self-learned playbook rules, but the league shows why a large body of organizational knowledge is only useful when it leads to disciplined action.
There is also an important fairness qualification in the standings. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference does not erase the observed decisions, but it belongs beside the result so readers can judge the comparison with the relevant experimental condition in view.

As an affiliate, we earn on qualifying purchases.
A business story that keeps producing evidence
Firmulate turns AI management from an abstract prediction into an observable, continuing story. The company’s financial mismatch, public countdown and versioned workdays ensure that decisions carry visible consequences. Its synthetic staff can recognize danger and resist social engineering, yet still falter between understanding an opportunity and completing the work.
The larger lesson is not that one model has permanently solved management. It is that credible evaluation must examine research habits, follow-through, discipline and trust under pressure. Firmulate makes those qualities watchable while the company continues its public fight for survival—and while every new workday adds another piece of evidence.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and ethics assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Summer Picks
summer essentials
As an affiliate, we earn on qualifying purchases.