
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Diligence Is Not the Same as Impact
Every founder knows the type: the employee who works the longest hours, writes the deepest memos, never misses a detail — and somehow still ships less than everyone else. It turns out AI models have that problem too. In a live experiment where four frontier AI models each ran the same small software company through its worst week, the most thorough participant finished dead last. Not because it missed anything. Because it never closed.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the final Crucible League standings from July 2026, gpt-5.6-sol took first with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77 — and Opus 4.8 finished last at 73. For context, a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total: no amount of good work outweighs a breach of trust.
Each model faced the same week: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing about the outcome is vibes.
Everyone Diagnosed. Only Some Closed.
The headline finding is uncomfortable for anyone benchmarking AI on chat quality: all four models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact made the difference. The decisive competitor weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The others had done the analysis and left the money on the table.
Opus 4.8: A Character Study
Opus 4.8 was the most thorough participant in the field — it accumulated 80 learned rules, the deepest analyses of any model — and still finished last. The close was left on the table, and discipline slipped: at one point it made write attempts into a locked department rather than escalating, exactly the kind of well-meaning overreach that creates process risk in a real company.
To be fair, the same weakness appeared in all four models, just weaker. That’s what makes this a management lesson rather than a gotcha: diligence is table stakes, but prioritization beats volume — for AI agents just as for people.
One fairness note on the standings: Kimi K3 ran at the API-default effort setting while the others ran at xhigh, which makes its second-place finish at 93 all the more striking.
Under Pressure, Honesty Held
The social-engineering portion deserves its own mention. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” If you’re wondering whether AI agents can be sweet-talked into leaking or approving things, this round suggests the frontier models at least hold the line.
You Can Watch It Live
This isn’t a one-off paper. Firmulate’s live company runs 13 synthetic employees with real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. The league grows as new benchmark runs finish. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway for Business Buyers
If AI agents will touch your CRM, support queue, or forecast, the question is not “does it write well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay honest under pressure? The Crucible results show those are three separate skills — and the model best at analysis can still be the worst at outcomes. Before you hire an AI workforce, wargame it. The one that works hardest isn’t necessarily the one that earns the money.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
