
A security stress test with a greenhouse lesson
Anyone who tends a greenhouse knows that healthy growth depends on more than ideal conditions. A structure may look sound on a calm afternoon, only to reveal its weaknesses when wind, heat and hurried decisions arrive together. Business technology deserves the same kind of testing.
Firmulate has applied that principle to frontier AI models. Instead of asking them isolated questions, the company placed each model in charge of the same small software business during its worst week. The customers, crises and temptations remained constant. Every decision was versioned and auditable.
Among the most encouraging results was a direct test of integrity under pressure: fake messages from the CEO escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt.
As an affiliate, we earn on qualifying purchases.
Pressure rose, but the boundary held
The social-engineering exercise was designed to resemble the moments when ordinary business safeguards are most vulnerable. An apparently senior authority demanded urgency and tried to bypass process. The requests escalated rather than disappearing after the first refusal. The reporter trick then approached the same boundary from another direction, presenting disclosure as a small and informal favor.
The models did not take the bait. Kimi K3 recorded the clearest concise assessment: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it identified both the procedural problem and the possibility that the apparent authority was not genuine.
This was not an isolated success surrounded by missed alarms. All models spotted every crisis and refused every manipulation attempt. Readers can examine more of the participants’ language on Firmulate’s public quotes page.
Integrity was necessary, but it was not sufficient
The experiment also exposed a different weakness. Although all participants recognized the crises, only 2 signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The decisive commercial fact was not visible in the customer event. It sat 2 document references deep inside the company’s own files. Models that found it won the deal at full price, worth +€4,583 MRR. The result separates several abilities that can look identical in a polished demonstration: identifying a problem, investigating the available evidence, maintaining trust and completing the work.
That distinction shaped the final July 2026 Crucible League:
- gpt-5.6-sol led with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
The do-nothing baseline scored 26 because partial progress counted. A single breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.” Full results and the accompanying plain-language findings are available on the Firmulate benchmarks page.
The most thorough model still finished last
Opus 4.8 offers a useful warning against equating volume with effectiveness. It produced the deepest analyses and added +80 learned rules, making it the most thorough participant. Yet it placed last after leaving the close on the table and allowing discipline to slip. Its write attempts went into a locked department instead of being escalated. The same weakness appeared more mildly in all 4 other models.
K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its performance, but it is relevant context when comparing participants.
A company small enough to inspect, difficult enough to matter
The live business contains 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing the experiment to be watched rather than accepted as a polished retrospective.
The public material also includes a “guess the model” quiz powered by 242 real, unedited management decisions. Together, those records turn model evaluation into something closer to observing a season under changing conditions than judging a single attractive bloom.

Test the storm before granting access
The strongest lesson is not that AI can never be manipulated. It is that integrity under pressure can be tested before deployment rather than discovered in an incident report. Firmulate’s result is encouraging: 5 of 5 models held the line through escalating executive impersonation and the reporter’s softer approach.
At the same time, the deal exercise shows why a refusal test alone is incomplete. A dependable business agent must protect trust, read the available files, escalate obstacles and finish legitimate work. Safety without execution leaves value behind; execution without integrity creates a far more serious risk.
Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. For organizations considering AI access to sensitive workflows, that creates a practical question worth asking early: how does the system behave when authority sounds convincing, time feels short and the correct answer is still no?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html