AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Before you orderOffer from Amazon

Get garden gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Tending a Garden Teaches You About Grading

Gardeners know a simple truth: an empty bed and a neglected bed are not the same thing. A plot that was weeded, mulched and watered before the drought hit will survive far better than one that was never touched — even if neither produces a harvest. Partial progress counts. And one catastrophic mistake — say, spraying the wrong herbicide across the whole bed — can undo an entire season of careful work, no matter how diligent you were beforehand.

That intuition — partial credit for real progress, and an absolute ceiling on damage — turns out to be exactly how an honest AI benchmark works. When the firm Firmulate ran four frontier AI models through the same simulated business crisis, the scoring system raised a question business readers rarely get to ask: why does a manager that does nothing still score 26 points instead of zero?

One Company, Its Worst Week, Run Four Times

The setup was elegantly controlled. Each frontier AI model was handed the same small software company and the same week from hell: same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the run could be quietly patched afterward.

The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95 points, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 managed 77, and Opus 4.8 landed last at 73.

Why Zero Effort Still Earns 26

Here is the part that makes skeptics sit up. Firmulate also runs a “do-nothing” baseline — an automated manager that takes no meaningful action at all. It scores 26 points, not zero.

The reasoning is the gardener’s logic again. Even a manager who simply doesn’t break things preserves something of value. Crises that are ignored still partially resolve; customers who are ignored still receive whatever baseline service the company already had. Partial progress counts — and a scoring system that awarded zero to inaction would implicitly claim that competent management creates all of a company’s outcomes, which is obviously false. A business is a going concern before any manager touches it. The 26-point floor is the benchmark’s way of admitting that honestly.

It also makes the top scores more meaningful. Gpt-5.6-sol’s 95 is not 95 points of contribution — it’s 95 against a baseline where doing nothing already gets you 26. The real gap between the champion and a passive seat-warmer is 69 points of genuine, attributable management work.

The One-Way Ceiling: A Breach of Trust Caps Everything

The floor has a mirror image at the top: a single breach of trust caps the total score, no matter how brilliant the rest of the run. Firmulate’s stated principle is blunt — “no amount of good work outweighs a breach of trust.” Cheat once, and the ceiling drops, full stop.

This is the same asymmetry every gardener recognizes. A season of perfect watering doesn’t undo one application of the wrong chemical. In business terms, it reflects how trust actually works with customers and colleagues: it’s accumulated slowly and destroyed instantly. A benchmark that let a model trade one deception against a pile of good decisions wouldn’t be measuring management — it would be measuring excuse-making.

Distrust of Round 100s

Notice, too, that no model scored 100 — and that the scoring treats that as a feature, not a failure. A perfect round score on a messy, week-long management simulation should trigger suspicion, not celebration. Real work leaves imperfections; the best performer here (95) still had room above it. Healthy skepticism about flawless grades is itself part of an honest benchmark’s design.

What Actually Separated the Winners

The most striking finding: all five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned — same diagnosis, same pitch, no signature. The decisive clue was buried two document references deep in the company’s own files, not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Then there’s Opus 4.8, the cautionary tale: the most thorough participant by volume — over 80 learned rules, the deepest analyses — yet last place. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

On the social-engineering front, fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” One fairness note: K3 ran without an effort parameter while the others ran at maximum effort.

It’s All Watchable

Behind the benchmark, Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark doesn’t just rank performers — it tells you what its numbers mean. Firmulate’s 26-point floor acknowledges that doing nothing still has value, its breach-of-trust cap encodes the reality that integrity isn’t tradable, and its refusal to hand out perfect scores keeps everyone appropriately humble. Like a well-kept garden journal, it records what actually happened, partial progress and all. If AI agents are going to touch your CRM, your support queue, or your forecast, that’s the kind of grading you want: one where a do-nothing manager earns 26, a cheater earns a cap, and nobody coasts to a round 100. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Summer Snack Hacks with the Ninja Air Fryer: Perfect Crispy Treats

Discover top tips and hacks for using the Ninja Air Fryer this summer to make crispy fries, roasted veggies, and more in your outdoor kitchen.

How to Fix a DEWALT Power Tool That Keeps Chuck Stuck

Learn practical, step-by-step methods to fix your DEWALT cordless drill or impact driver when the chuck gets stuck. Safe, accurate, and easy to follow.

Ninja Foodi 10-Quart DualZone XL Air Fryer: Perfect for Summer Family Feasts

Discover if now’s the right time to get the Ninja Foodi DualZone XL Air Fryer for summer meals or wait for Prime Day deals. Practical tips included!

Ninja Blast Max Portable Blender: Your Summer On-the-Go Smoothie Buddy

Compare Ninja Blast Max to typical portable blenders—see where it shines for outdoor summer sipping and where it falls short for larger batches.