🔍 Read the full analysis: Try These Business Stress Tests Before Deploying AI Agents on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five frontier models completed a simulated company’s difficult week in its Crucible League, with scores ranging from 73 to 95. The results highlight gaps between recognizing a crisis and acting on available evidence, while the company is offering read-only pilots to test models against businesses’ own data.
Firmulate has published results from a simulated business crisis in which five AI models managed the same small software company, and says it is offering companies a way to test agents against their own data using read-only exports. The July 2026 Crucible League found that models identified the crises and rejected manipulation attempts, but varied in whether they acted on internal evidence to close a justified deal.
In the final Crucible League, the models faced a difficult week at a simulated software company. Firmulate says decisions were versioned and auditable. Its reported scores were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The scoring system counted partial progress but capped a model’s total if it breached trust.
Firmulate reports that all five models spotted every crisis and refused each manipulation attempt. The distinction came in a €55,000 deal: only two signed after their own analysis supported the opportunity. The company says the deciding evidence was a competitor weakness referenced two documents deep in the simulated company’s files. Models that found it closed the deal at full price, adding €4,583 in monthly recurring revenue in the experiment.
The exercise also tested authority boundaries. Fake messages purporting to come from the CEO escalated over three stages, followed by a reporter’s request for an informal yes-or-no answer “on background.” Firmulate says all five models refused. It also reports that Opus 4.8 generated the most learned rules, at 80, and produced the deepest analyses, but left the deal unsigned and tried to write to a locked department rather than escalating. The company says weaker versions of that boundary problem appeared in the other four models.
From Crisis Recognition to Execution
The results focus attention on a practical distinction for businesses considering AI agents: identifying a problem is not the same as completing the work safely. An agent may recognize an emergency and make a persuasive recommendation, yet fail to retrieve relevant information from company files or carry an approved action through to completion. In Firmulate’s account, document retrieval and follow-through separated some deal outcomes even when the models had diagnosed the situation similarly.
That distinction matters if companies plan to connect agents to customer records, sales pipelines or internal procedures. A model’s performance in a polished demonstration may not reveal how it behaves when evidence is scattered, an apparent authority figure requests an exception, or its preferred route is blocked. Firmulate’s proposed pilot aims to surface such weaknesses before an agent can affect live operations. Its results are from one simulated company and one set of scenarios, however; they do not establish how a model will perform in every business or deployment.
A Simulated Company and Pilot
Firmulate’s live experiment uses a company with 13 synthetic employees and financial pressure built into the scenario. The company describes monthly costs of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and workdays tracked by version. These are features of the simulation, not reported financial results from a real operating business.
The company also offers a quiz based on 242 management decisions, asking readers to guess which model made each choice. Its enterprise pilot shifts from the synthetic setup to a company’s own information: Firmulate says it takes a read-only data export, runs crisis scenarios and produces a board report with model rankings and weaknesses in the company’s playbooks. The stated design does not write changes back to operational systems.
“No amount of good work outweighs a breach of trust.”
— Firmulate, describing its scoring rule
Limits of the League Results
The standings describe one experiment; the published account does not establish that the same ranking would hold across different companies, tasks or operating conditions. A comparison caveat also applies: Firmulate says Kimi K3 ran without an effort parameter and used the API default, while the other models ran at xhigh. The effect of that difference on the results is not specified.
Details needed to independently assess the broader applicability of the scores—including how each scenario was weighted and how the models’ outputs were evaluated—are not given in the account summarized here. Firmulate’s pilot description also does not specify the full data-handling terms, scenario design or evaluation method for a particular customer. Those details would matter when interpreting a company-specific report.
Company-Specific Pilots Ahead
Firmulate is inviting companies to discuss pilots using read-only exports of their data. The proposed next step is to run scenarios against a company’s customers, pipeline, rules and pressure points, then review a board report on model rankings and playbook weaknesses. The company says no results are written back to real systems. Businesses considering a pilot would need to review its process and the resulting findings in their own operational context.
The public experiment remains available at firmulate.com/live, with the full standings at firmulate.com/benchmarks.html. Firmulate directs businesses interested in a pilot to its pilot page or to contact@firmulate.com.
Source: ThorstenMeyerAI.com
Key Questions
What did the Firmulate Crucible League test?
It tested five AI models managing the same simulated software company through a difficult week, including crisis response, information retrieval, a sales opportunity and attempts to bypass trust or authority boundaries.
Which model scored highest?
Firmulate reported that gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
What was the main difference in the deal scenario?
Firmulate says only two models signed a €55,000 deal supported by their analysis. The deciding competitor weakness was in the simulated company’s files, two document references deep; models that found it won at full price.
How does the proposed enterprise pilot work?
Firmulate says it uses a read-only export of a company’s data to run crisis scenarios and prepare a report on model rankings and weaknesses in the company’s playbooks. The described pilot does not write back to live systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
