AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

What happens when a business becomes the experiment?

For readers interested in education and science, Firmulate offers an unusually vivid case study: a software company operating as a public laboratory for artificial intelligence. Its 13 synthetic employees face customers, commercial pressure and a deteriorating financial position. The company burns €105,000 each month against €2,300 in monthly recurring revenue, while a public cash countdown makes the consequences visible.

This is not a polished demonstration designed to produce a flattering answer. Every workday is versioned, creating a continuing record of what the company noticed, what it decided and what it failed to finish. The result is a form of build-in-public pushed to an extreme: the audience can watch the company operate while its survival remains an unresolved business problem.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A controlled test with commercial consequences

Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. They received the same customers, crises and temptations. Their decisions were versioned and auditable, allowing the comparison to focus on management behavior rather than conversational polish.

The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, had a hard boundary: a single breach capped the total under the principle that "no amount of good work outweighs a breach of trust."

The reassuring finding was that every model spotted every crisis and rejected every manipulation attempt. The more revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate summarizes the failure neatly: "Same diagnosis, same pitch — no signature."

That distinction matters because knowing what to do is not the same as completing it. A model can analyze a customer, prepare an argument and identify the correct commercial move, yet still leave the decisive action unfinished. In an educational setting, that resembles a student showing sound reasoning but omitting the final answer. In a company, the missing final step can mean missing revenue.

The evidence was in the company’s own files

The detail that separated the successful dealmakers was not obvious in the customer event. A decisive competitor weakness sat two document references deep in the company’s files. Models that followed those references found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is a useful lesson about AI work beyond Firmulate. The challenge is not always generating a clever response. Sometimes it is reading the available material carefully enough to discover that the apparent task is only the surface of the real one. The experiment makes that research habit economically legible: reading the file changed the commercial outcome.

Pressure tested trust as well as competence

The worst week also included fake CEO messages that escalated over three stages and a reporter seeking "just one yes/no, on background." All 5 of 5 models refused. Kimi K3 recorded its reasoning on the suspected executive message: "Treat the request as a suspected approval-bypass / possible impersonation."

That finding complicates the familiar fear that an AI system will inevitably comply with an authoritative-sounding request. In this test, the models maintained the boundary. Their failures appeared elsewhere: incomplete execution and operational discipline, not successful manipulation. Readers can also inspect what the synthetic employees actually say, rather than relying only on a retrospective summary.

Thoroughness did not guarantee victory

Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.

This is one of the experiment’s most instructive reversals. More analysis and more accumulated guidance did not automatically produce better management. Firmulate’s live company has already accumulated more than 680 self-learned playbook rules, but the league shows why a large body of organizational knowledge is only useful when it leads to disciplined action.

There is also an important fairness qualification in the standings. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference does not erase the observed decisions, but it belongs beside the result so readers can judge the comparison with the relevant experimental condition in view.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI model version control tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business story that keeps producing evidence

Firmulate turns AI management from an abstract prediction into an observable, continuing story. The company’s financial mismatch, public countdown and versioned workdays ensure that decisions carry visible consequences. Its synthetic staff can recognize danger and resist social engineering, yet still falter between understanding an opportunity and completing the work.

The larger lesson is not that one model has permanently solved management. It is that credible evaluation must examine research habits, follow-through, discipline and trust under pressure. Firmulate makes those qualities watchable while the company continues its public fight for survival—and while every new workday adds another piece of evidence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and ethics assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business analytics software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

SUMMER

Summer Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Women in Cryptozoology: Overlooked Explorers

I invite you to discover how women in cryptozoology have shaped mysterious legends and challenged biases, revealing stories that demand our attention.

The Monster’s Lair: Exploring Real Caves Linked to Myths

Jump into the depths of real caves linked to ancient myths and uncover secrets that could change everything you thought you knew.

Investigating Cryptid Evidence: Footprints, Photos, and DNA

Just how can we tell real cryptid evidence from hoaxes and natural causes? Discover the key methods to uncover the truth.

The Rational Reconstruction of Fairy Tales

Providing a window into societal values, the rational reconstruction of fairy tales reveals hidden meanings that will make you see these stories differently.