
Every gardener knows the type: the neighbor who plants tomatoes before checking the soil pH, waters on a schedule instead of checking the moisture, and wonders come August why the harvest disappointed. The information was all there — in the soil test, in the weather station, in the extension office’s free pamphlet. They just didn’t read it.
It turns out the most advanced AI systems in the world have exactly the same failing, and a remarkable live experiment has now measured it — in dollars, not vibes. A public project called Firmulate ran four frontier AI models through the same simulated business crisis, and the difference between winning and losing a €55,000 deal came down to one simple habit: whether the AI read the company’s own files before answering.
The worst week in business, run four times
Firmulate’s setup is elegantly cruel. Each frontier AI model was handed the same small software company and the same seven-day gauntlet: angry customers, cascading crises, and a series of temptations to cut corners. Every decision was versioned and auditable, so nothing could be fudged after the fact. Only the model changed between runs.
The Crucible League’s final July 2026 standings tell the story. GPT-5.6-Sol took first place with a score of 95, Kimi K3 — a newcomer from Moonshot — scored 93, Sonnet 5 came in at 88, Fable 5 at 77, and Opus 4.8 finished last at 73. For context, a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — as the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Everyone passed the obvious test. Almost everyone failed the quiet one.
Here’s what should encourage anyone worried about AI agents going rogue: all five models spotted every crisis and refused every manipulation attempt. The social engineering was genuinely nasty — fake CEO messages escalating over three stages, plus a reporter’s trick request framed as “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But the deal itself was another matter. Only two of the models signed the €55,000 contract that their own analysis had earned. The rest delivered the same diagnosis, made the same pitch — and never closed. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”
The fact buried two documents deep
Why did two models close at full price while the others left the money on the table? The decisive competitor weakness — the fact that made the customer’s decision obvious — wasn’t in the customer meeting at all. It sat two document references deep in the company’s own internal files. A model had to follow a citation from one document to another to find it.
The models that did the follow-the-thread work won the €55,000 deal at full price, worth an added €4,583 in monthly recurring revenue. The models that didn’t simply never got there. This is the AI equivalent of the gardener who never turns to page two of the soil report — where the nitrogen deficiency was hiding all along.
Thoroughness isn’t the same as follow-through
The most counterintuitive finding concerns Opus 4.8. It was the most thorough participant in the entire experiment — it learned the most new rules, over 80, and produced the deepest analyses. Yet it finished dead last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the problem properly. The same weakness showed up, more mildly, in all four of the other models.
One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.
You can watch the company, and test yourself
Firmulate isn’t a one-off lab report. It’s a live, watchable operation: 13 synthetic employees, real money mechanics, a burn rate of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day as new benchmark runs finish.
There’s also a genuinely fun twist: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. If you’ve ever wondered whether you can tell one AI’s judgment from another’s, you can try. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Why a gardener should care
You may never run a software company, AI or otherwise. But if you use any AI tool — to plan a greenhouse layout, diagnose blight from photos, or price a landscaping job — the lesson transfers directly. The question is no longer “does it write well?” Every frontier model now writes beautifully. The questions that separate a good harvest from a disappointing one are: does it finish what it starts, does it read your files before answering, and does it stay honest under pressure?
The gap between a 95 and a 73 isn’t eloquence. It’s whether the AI bothered to open the second document — the one with the answer in it.

Firmulate’s experiment turned a vague worry — “will AI agents actually do their homework?” — into a hard number: a €55,000 deal, lost by models that stopped reading one reference short. Refusing manipulation is now table stakes; the frontier models all passed. Reading the file before answering is what separates the winners, and it’s a measurable, purchase-deciding property you can check before you trust any AI with your business — or your garden. The full league table and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html