AIThis post was created with the assistance of artificial intelligence (AI).

Every gardener knows the rule: you don’t plant a new variety across the whole plot in year one. You run a trial bed first — same soil, same weather, same pests — and you watch what happens before you commit. A bad season in a test corner costs you a few rows. A bad season everywhere costs the harvest.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Businesses are now hiring AI agents to tend their customer pipelines, support queues and forecasts the way a gardener tends tomatoes — daily, unsupervised, with real stakes. But almost nobody runs the trial bed first. A live experiment called Firmulate has been doing exactly that: running frontier AI models as the management of a real small software company through its worst week, and scoring what actually happened — not how charmingly the models chatted about it.

Same company, same crises, only the model changes

The setup is elegantly controlled, like a greenhouse with identical beds. Four frontier AI models were each handed the same small software company, the same customers, the same sequence of crises and temptations. Every decision was versioned and auditable — think of it as a gardening journal where every single action is logged and replayable.

The final league table from July 2026 tells a story no demo would reveal:

  • gpt-5.6-sol — 95 points
  • Kimi K3 — 93 (with a fairness caveat: K3 ran without an effort parameter, at API default, while the others ran at xhigh)
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

The do-nothing baseline scores 26. Partial progress counts, but a single breach of trust caps the total — in Firmulate’s words, “no amount of good work outweighs a breach of trust.”

Everyone passed the obvious tests. Almost nobody closed.

Here’s the finding that should stop any executive mid-sentence: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap is invisible in a chat demo. It only shows up when the model has to actually finish the job.

The buried fact

The decisive competitive weakness in that €55k deal wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. Knowing your own soil, it turns out, is what closes the harvest.

Pressure-tested against tricks

The experiment also threw social engineering at the models: fake CEO messages escalating over three stages, plus a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of guard rail you want verified before an agent touches your real systems, not after.

The most thorough model came last

Opus 4.8 is the cautionary tale of the batch: the most thorough participant, with over 80 learned rules and the deepest analyses — and a last-place score. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence without follow-through is a very human failure — and the wargame catches it.

You can watch it live — or read the tea leaves yourself

The lab company runs for real: 13 synthetic employees, real money mechanics — €105k monthly burn against just €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com. And if you think you could out-manage the machines, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The gardener’s logic applies cleanly here: trial bed first, full field later. If AI agents will ever touch your CRM, your support queue or your forecast, you want to know — before deployment — whether they spot crises, refuse impersonators, read their own documentation, and actually close. A polished demo tells you none of that.

Enterprises can now run this same wargame against a read-only export of their own business: your customers, your pipeline, your rules, hit with churn waves, price wars and social-engineering pressure. Nothing ever writes back to real systems — it’s the trial bed, not the field. If that’s a season you want to test before planting, start at firmulate.com/pilot.html or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ninja Blast Max Portable Blender: Your Summer On-the-Go Smoothie Buddy

Compare Ninja Blast Max to typical portable blenders—see where it shines for outdoor summer sipping and where it falls short for larger batches.

Summer-Ready Drinks: Must-Have Accessories for Your Ninja SLUSHi

Enhance your Ninja SLUSHi experience with accessories that boost flavor, convenience, and fun for summer frozen drinks at home.

Ninja DualBrew Pro Coffee Maker: Perfect Summer Brews

An honest review of the Ninja DualBrew Pro, ideal for summer mornings. Discover who it’s for, its strengths, and trade-offs for outdoor coffee lovers.

Ninja Foodi 8-Quart DualZone Air Fryer: Summer Meal Magic

Discover why the Ninja Foodi 8-Quart DualZone Air Fryer is perfect for quick, versatile summer meals with its dual zones and six cooking functions.