
In crypto, an AI agent that can move money or answer customers can turn a small mistake into a very public one. The harder question is not whether it can explain a crisis. It is whether it can handle pressure, follow the rules and finish the job. Firmulate’s live experiment puts AI models through a company’s worst week to find out.
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A shared crisis, different outcomes
In the final Crucible League, published in July 2026, five models faced the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable. The results ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s rule is pointed: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
Recognizing the crisis was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap, captured in the experiment’s phrase “Same diagnosis, same pitch — no signature,” is a reminder that sound analysis and effective action are different tests.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That is a practical lesson for any business testing AI against its own operations: useful context may sit in internal records, while the visible crisis is unfolding elsewhere.
Trust under pressure
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.”
But refusal alone did not guarantee strong management. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and discipline slipped when it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. In business settings, restraint matters; so does knowing when to act and when to raise a problem.
Watch the company; then test your own
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k each month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The live experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call at firmulate.com.
For enterprises, the next step is a pilot using a read-only export of their own business. Teams can run crisis scenarios against their company and receive a board report with model rankings and weaknesses in their playbooks. The pilot does not write back to real systems. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh—a qualification to keep in mind when comparing the league results.

Put the playbook under pressure
A model can identify a threat and still miss the moment to act. For businesses considering AI in customer operations, finance or other sensitive work, a company-specific wargame offers a way to examine decisions before deployment. Explore a Firmulate pilot using a read-only export of your business, and contact contact@firmulate.com to discuss it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
