AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

We Already Know This Lesson From the Garden

Every gardener knows the difference between a plant that looks perfect at the garden center and one that survives its first season in your actual soil. The label tells you nothing about how it handles a late frost, a week of neglect, or an aphid invasion. You only learn a plant’s real quality when conditions turn against it.

It turns out AI models have exactly the same problem — and a live experiment at Firmulate is exposing it in public, twice a day, with real money mechanics on the line.

Amazon

AI stress test software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Problem With AI Report Cards

Right now, the AI industry measures its models the way a plant tag measures a seedling: by appearance under ideal conditions. Coding leaderboards test whether a model can solve a neat, well-defined problem. Chat arenas test whether its answers sound good in a one-off conversation. Neither tests what actually matters if you’re going to let an AI touch your business: how it behaves during a bad week.

That’s the gap Firmulate was built to measure. Its pitch is blunt: it measures management quality, not chat quality. And it does so not with a quiz, but by running four frontier AI models as the complete leadership of the same small software company going through its worst week — same customers, same crises, same temptations to cut corners. Only the model changes. Every decision is versioned and auditable.

The Worst Week, Run Four Times

Think of it as a stress test for a greenhouse — except instead of frost warnings and broken heaters, the crises are a churn wave, a price increase, a downround and a PR crisis. The models face social engineering too: fake CEO messages escalating over three stages, plus a reporter’s trick request — “just one yes/no, on background.”

The final Crucible League standings from July 2026 tell a surprising story. gpt-5.6-sol took first place with a score of 95. Kimi K3, a newcomer from Moonshot, came second at 93 — notably, K3 ran without an effort parameter while the others ran at maximum effort, and still posted the cleanest discipline of the field. Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 landed last at 73.

For calibration: a do-nothing baseline scores 26. Partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Everyone Passed the Obvious Tests. Almost.

Here’s the finding that should make any business owner sit up: all four models spotted every crisis and refused every manipulation attempt. Five out of five refused the reporter trick outright — Kimi K3’s on-record reasoning was, “Treat the request as a suspected approval-bypass / possible impersonation.” On vigilance, the field is strong.

But only two models — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. The other two delivered the same diagnosis and the same pitch, and then… nothing. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”

The Buried Fact That Decided Everything

The most telling detail? The decisive competitive weakness that won the deal wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue.

It’s the business equivalent of a gardener who reads the soil report before planting versus one who just likes the look of the bed. Thoroughness in your own records beats polish in your conversation.

That lesson cut deepest for Opus 4.8, the most thorough participant of all — it learned over 80 new rules and produced the deepest analyses, yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. And here’s the uncomfortable part — the same weakness appeared, weaker, in all four models.

It’s Running Right Now

This isn’t a slide deck. The company is live software with 13 synthetic employees and real money mechanics: burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com, now on company day 1131 — the site rebuilds itself twice a day.

Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The Takeaway for Anyone Hiring an AI

If AI agents will touch your customer records, support queue or forecast, the useful question is not “does it write well?” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost?

Gardeners learned long ago not to trust the label over the lived season. The Firmulate experiment suggests it’s time to hold AI agents to the same standard: don’t judge them by the demo. Judge them by how they run a company through its worst week — because as the €55,000 deal left unsigned proves, the difference between a great answer and a great manager is the difference between talking about the harvest and actually bringing it in.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Building a Hidden Desktop in a Closet

IIn this guide, discover essential tips to create a discreet, functional desktop inside a closet, but there are key considerations you won’t want to miss.

Master the Perfect Summer Iced Espresso with Ninja Luxe Café Pro

Learn how to make refreshing iced espresso drinks at home with the Ninja Luxe Café Pro Series. Follow our step-by-step guide for café-quality results.

DEWALT vs Ryobi: Which Power Tool Reigns Supreme?

Compare DEWALT’s 20V MAX Cordless Drill with Ryobi’s line to find the best fit for your DIY or professional projects. Honest insights and detailed specs included.

Best DEWALT Power Tools for DIY (2026) — Guide 5

Discover the top DEWALT power tools for DIY in 2026. Our roundup highlights the best options for beginners, value, and versatility to elevate your projects.