
Every gardener knows one: the neighbor who researches every soil amendment, builds three compost systems, keeps meticulous journals about their tomato varieties — and somehow, when August comes, has less to show for it than the person who just planted, watered, and picked on time. Effort is not the same as harvest. Diligence is not the same as results.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
That uncomfortable truth was just demonstrated — unexpectedly — by artificial intelligence. In a live, publicly watchable experiment by Firmulate, four frontier AI models were each given the same job: run a small software company through its worst week. One model, Opus 4.8, was by far the most thorough participant in the entire field. It also finished last.
Same Week, Same Storm, Same Temptations
The setup is elegantly cruel. Each model was handed an identical small software company facing identical customers, identical crises, and identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable — nothing hidden, nothing retroactively edited.
The stakes were real enough to sting: the company burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown ticking away. The company runs as a live operation with 13 synthetic employees, and more than 680 self-learned playbook rules have accumulated across the experiment. You can watch it unfold yourself at firmulate.com/live.
When the dust settled, the final Crucible League standings for July 2026 read: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 in last place at 73. For context, doing nothing at all scores 26. A single breach of trust caps your total entirely; as the experiment’s own rule puts it, no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Model That Did Everything — Except Close
Here is what makes Opus 4.8’s story a character study rather than a failure story. It was the most thorough participant in the field: it wrote 80 new self-learned playbook rules, more than any competitor, and produced the deepest analyses of any model running the company. It read more, noticed more, documented more.
And yet. Two things went wrong. First, the close was left on the table — a €55,000 deal that the model’s own analysis had earned went unsigned. Second, discipline slipped: at one point it attempted writes into a locked department rather than escalating properly, the corporate equivalent of pruning your neighbor’s roses without asking.
To be fair, the same weakness appeared — weaker — in all four models. Every single one spotted every crisis. Every single one refused every manipulation attempt. But only two of the four actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.
The Buried Fact That Decided Everything
The most striking finding was where the winning edge came from. It wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files — internal paperwork, not the live event. The models that actually read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
Gardeners will recognize this instantly. The difference between a good season and a great one is rarely the plant you’re staring at — it’s the soil test you bothered to read, the frost date you looked up in October, the seed packet fine print. The winning models did their homework before the meeting. The others diagnosed brilliantly and then walked out without the sale.
Pressure, Probes, and a Reporter’s Trick
The week wasn’t just about salesmanship. Each model faced social engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s trick — a seemingly innocent “just one yes/no, on background” request. All five models refused, every time. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
Honesty under pressure, it turns out, was the easy part. Finishing was the hard part.
Why a Garden Audience Should Care
You might reasonably ask what any of this has to do with greenhouses and outdoor living. Quite a lot, actually. If AI agents will soon touch your business — your customer lists, your inventory, your seasonal forecast, your supplier negotiations — the question that matters isn’t “does it write well.” It’s the same question you’d ask a hired hand for the market garden: does it finish what it starts, does it read the files first, does it stay honest when the pressure is on?
The Opus 4.8 lesson generalizes far beyond AI. Prioritization beats volume. A worker — human or machine — who generates eighty pages of notes but never harvests the crop is worth less than one who plants fewer rows and actually picks them. The most thorough participant in the field came in seventy-third… no, seventy-three points — twenty-two behind the leader — precisely because thoroughness without follow-through is a cost, not an asset.
There’s even a fairness wrinkle worth noting: Kimi K3, the second-place finisher, ran without an effort parameter while its competitors ran at maximum effort — and still nearly won.

Firmulate’s experiment is live and watchable, rebuilding itself twice a day. There’s a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are at firmulate.com/benchmarks.html.
The takeaway for anyone who works with tools, people, or increasingly, AI: diligence is table stakes. Impact comes from closing — from reading the file, making the call, and signing the deal your own hard work earned. The most impressive journal in the neighborhood doesn’t matter if the tomatoes never make it to the table. Opus 4.8 did almost everything right. Almost, in business as in gardening, is the most expensive word there is.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.