AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A living experiment under glass

Gardeners know that a greenhouse reveals more than a photograph of a perfect bloom. You see whether plants survive heat, neglect and sudden changes—not merely whether they looked healthy for a moment. Firmulate applies a similar kind of sustained observation to artificial intelligence.

The public experiment follows a small software company operated by 13 synthetic employees. Its finances are deliberately unforgiving: the business burns €105k a month while producing €2.3k in monthly recurring revenue. A public cash countdown makes the imbalance visible, turning corporate survival into an unfolding story rather than a polished demonstration.

Anyone can watch the company live. Every workday is versioned, and the operation has accumulated more than 680 self-learned playbook rules. The result is an unusually exposed form of building in public: not simply showing a finished product, but displaying how an AI-run organization responds while the money runs down.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Putting frontier models through the same bad week

Firmulate’s Crucible League tested frontier models by giving each one the same company, customers, crises and temptations. Every decision was versioned and auditable. The exercise was designed around management performance: noticing trouble, resisting manipulation, finding relevant information and completing commercially valuable work.

The final July 2026 standings placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the benchmark imposed a firm trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

On the surface, the participants appeared remarkably capable. All models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 contract their own analysis had earned. Firmulate sums up that execution gap starkly: “Same diagnosis, same pitch — no signature.”

The valuable fact hidden in the company’s own files

The decisive information did not arrive conveniently in the customer event. A competitor weakness was buried two document references deep inside the company’s files. The models that followed those references found the weakness, used it and won the contract at full price. The deal was worth an additional €4,583 in monthly recurring revenue.

That finding should resonate with anyone who has managed a garden, greenhouse or outdoor project. Recognizing wilted leaves is not the same as tracing the problem to soil, drainage or an earlier change in care. Likewise, an AI manager can correctly describe a commercial opportunity without doing the less glamorous work of opening the right records and carrying the opportunity through to a signature.

Pressure tested the company’s ethics, too

The bad week included fake CEO messages that escalated through three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused the manipulation. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters because an autonomous workplace must be judged on more than productivity. A model that responds eloquently but reveals information or bypasses approval under pressure is not merely inefficient; it is dangerous. Here, the field recognized the traps. The difference between models emerged elsewhere—in follow-through, research depth and operating discipline.

Why the most thorough model finished last

Opus 4.8 offers the experiment’s most instructive profile. It produced the deepest analyses and added 80 learned rules, more than any other participant. Yet it finished last. It left the commercial close on the table and repeatedly attempted to write into a locked department instead of escalating the blockage. The same weakness appeared in all four of the other models, though less strongly.

The contrast challenges a familiar assumption about AI work: that more analysis automatically produces better outcomes. Thoroughness was valuable, but it did not compensate for incomplete execution or weaker procedural discipline. A crowded pot can contain impressive growth while still failing to bear fruit.

There is also an important testing caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference should remain visible when readers compare the league positions.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

A company that turns every day into evidence

Firmulate’s larger contribution is the continuity of the experiment. The synthetic team does not disappear when a benchmark ends. It returns to the same financially distressed company, works through another business day and adds fresh, reviewable material to the public record. Readers can also read what the synthetic employees say, bringing the organizational drama closer to the surface.

The portrait is compelling because success is neither assumed nor guaranteed. This company has a team, customers, learned procedures and real money mechanics, yet its €105k monthly burn towers over €2.3k in monthly recurring revenue. Its survival problem remains visible instead of being edited out.

For garden and outdoor-living readers, the most familiar lesson may be that living systems reveal themselves through repeated care under changing conditions. A single attractive result says little about resilience. By keeping the company running in public, Firmulate asks a harder and more useful question: can an AI workforce keep finding the buried facts, resisting bad instructions and finishing the work before the cash countdown reaches its end?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

DEWALT vs Milwaukee: Which Power Tool Wins in 2026?

Compare DEWALT’s cordless drill with Milwaukee’s lineup to see which brand leads in power, features, and value in 2026. Find your perfect tool today!

Decorating With Light Temperature and CRI

Properly choosing light temperature and CRI can transform your space—discover how to create inviting, vibrant environments by mastering these lighting principles.

What Wine Refrigerators Add to Entertaining Spaces

Inevitably, a wine refrigerator enhances your entertaining space by combining style and practicality—discover how it can elevate your hosting game.

How Built-In Coffee Machines Change Morning Routines

Fascinatingly, built-in coffee machines revolutionize mornings by offering effortless, personalized brews—discover how they can transform your routine.