
A new power tool earns its place in a workshop by doing the job safely and reliably, not by sounding confident in the manual. AI agents deserve the same kind of practical trial. Before they touch a company’s customer records or support queue, can they handle a rough week, follow the rules and finish the work?
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
One company, the same hard week
Firmulate puts AI models through a live, watchable business experiment. Each model was given the same small software company, the same customers, the same crises and the same temptations. Its decisions were versioned and auditable, so the results can be followed rather than taken on trust. The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue. Its public cash countdown makes the pressure visible.
The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s rule is plain: “no amount of good work outweighs a breach of trust.”
Knowing what to do is only half the job
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. They reached the same diagnosis and made the same pitch, but some never closed. It is a familiar workshop lesson: identifying the right repair does not help if the job is left unfinished.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The company’s 680+ self-learned playbook rules and versioned workdays offer a record of how decisions played out under pressure.
Trust held up even against staged deception. Five of five models refused fake CEO messages that escalated over three stages, as well as a reporter’s request for “just one yes/no, on background.” Kimi K3’s on-record reasoning called it “a suspected approval-bypass / possible impersonation.”
Thorough work still needs follow-through
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the close on the table and discipline slipped when it tried to write into a locked department instead of escalating. That weakness appeared, less strongly, in all four models. A careful plan matters; so does respecting boundaries and knowing when to ask for help.
One comparison deserves context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The league gives readers a result to inspect, not a reason to ignore the conditions behind it. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.
From watching to trying it on your business
For a company considering AI agents, the useful next step is to test them against the work and pressure they may actually face. Firmulate’s enterprise pilot uses a read-only export of a company’s data to create a digital twin, then runs crisis scenarios against it. The resulting board report can show model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems.
That makes the pilot a practical rehearsal: expose gaps in judgment, follow-through or escalation while the exercise is still separate from live operations. Readers can watch the public experiment at Firmulate and explore what a company-specific trial involves on the pilot page.

Put the plan through its paces
AI can recognize trouble and resist manipulation, yet still fail to complete the work. Test how it handles your company’s own crises before giving it a role in live operations. To discuss an enterprise pilot using a read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
