
Get garden gear delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
What Tending a Garden Teaches You About Grading
Gardeners know a simple truth: an empty bed and a neglected bed are not the same thing. A plot that was weeded, mulched and watered before the drought hit will survive far better than one that was never touched — even if neither produces a harvest. Partial progress counts. And one catastrophic mistake — say, spraying the wrong herbicide across the whole bed — can undo an entire season of careful work, no matter how diligent you were beforehand.
That intuition — partial credit for real progress, and an absolute ceiling on damage — turns out to be exactly how an honest AI benchmark works. When the firm Firmulate ran four frontier AI models through the same simulated business crisis, the scoring system raised a question business readers rarely get to ask: why does a manager that does nothing still score 26 points instead of zero?
One Company, Its Worst Week, Run Four Times
The setup was elegantly controlled. Each frontier AI model was handed the same small software company and the same week from hell: same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the run could be quietly patched afterward.
The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95 points, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 managed 77, and Opus 4.8 landed last at 73.
Why Zero Effort Still Earns 26
Here is the part that makes skeptics sit up. Firmulate also runs a “do-nothing” baseline — an automated manager that takes no meaningful action at all. It scores 26 points, not zero.
The reasoning is the gardener’s logic again. Even a manager who simply doesn’t break things preserves something of value. Crises that are ignored still partially resolve; customers who are ignored still receive whatever baseline service the company already had. Partial progress counts — and a scoring system that awarded zero to inaction would implicitly claim that competent management creates all of a company’s outcomes, which is obviously false. A business is a going concern before any manager touches it. The 26-point floor is the benchmark’s way of admitting that honestly.
It also makes the top scores more meaningful. Gpt-5.6-sol’s 95 is not 95 points of contribution — it’s 95 against a baseline where doing nothing already gets you 26. The real gap between the champion and a passive seat-warmer is 69 points of genuine, attributable management work.
The One-Way Ceiling: A Breach of Trust Caps Everything
The floor has a mirror image at the top: a single breach of trust caps the total score, no matter how brilliant the rest of the run. Firmulate’s stated principle is blunt — “no amount of good work outweighs a breach of trust.” Cheat once, and the ceiling drops, full stop.
This is the same asymmetry every gardener recognizes. A season of perfect watering doesn’t undo one application of the wrong chemical. In business terms, it reflects how trust actually works with customers and colleagues: it’s accumulated slowly and destroyed instantly. A benchmark that let a model trade one deception against a pile of good decisions wouldn’t be measuring management — it would be measuring excuse-making.
Distrust of Round 100s
Notice, too, that no model scored 100 — and that the scoring treats that as a feature, not a failure. A perfect round score on a messy, week-long management simulation should trigger suspicion, not celebration. Real work leaves imperfections; the best performer here (95) still had room above it. Healthy skepticism about flawless grades is itself part of an honest benchmark’s design.
What Actually Separated the Winners
The most striking finding: all five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned — same diagnosis, same pitch, no signature. The decisive clue was buried two document references deep in the company’s own files, not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
Then there’s Opus 4.8, the cautionary tale: the most thorough participant by volume — over 80 learned rules, the deepest analyses — yet last place. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
On the social-engineering front, fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” One fairness note: K3 ran without an effort parameter while the others ran at maximum effort.
It’s All Watchable
Behind the benchmark, Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The Takeaway
An honest benchmark doesn’t just rank performers — it tells you what its numbers mean. Firmulate’s 26-point floor acknowledges that doing nothing still has value, its breach-of-trust cap encodes the reality that integrity isn’t tradable, and its refusal to hand out perfect scores keeps everyone appropriately humble. Like a well-kept garden journal, it records what actually happened, partial progress and all. If AI agents are going to touch your CRM, your support queue, or your forecast, that’s the kind of grading you want: one where a do-nothing manager earns 26, a cheater earns a cap, and nobody coasts to a round 100. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
