
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Demo Always Looks Great. The Job Is Another Story.
If you run an online business, you’ve seen the pattern: an AI model writes a flawless product description, answers a support ticket with perfect empathy, drafts a pricing email you’d almost send unchanged. The demo is brilliant. Then you hand it real responsibility — a support queue, a pricing decision, a customer on the edge of churning — and the picture gets murky. Does it finish what it starts? Does it read the customer’s file before answering? Does it stay honest when nobody’s watching?
Those questions don’t show up on coding leaderboards or chat arenas. Those measure answer quality. They don’t measure what happens when a model has to triage five fires at once, keep its story straight over days, and resist the temptation to take a shortcut that would look fine in a weekly report. A live experiment called Firmulate has now run exactly that test — and the results say a lot about what "hiring" an AI agent should actually involve.
As an affiliate, we earn on qualifying purchases.
One Company, Four Managers, One Very Bad Week
The setup is elegantly simple: four frontier AI models were each given the same job — run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changes. Every decision is versioned and auditable, so nothing rests on anecdote.
The week included the kind of scenarios that make founders sweat: a churn wave, a price increase, a downround situation, a PR crisis. Alongside those came something nastier — social engineering. Fake CEO messages that escalated over three stages, plus a reporter pulling the classic "just one yes/no, on background" trick.
The headline finding is a split verdict. All four models spotted every crisis. All four refused every manipulation attempt — five out of five refusals, with Kimi K3 leaving on-record reasoning that reads like a security playbook: "Treat the request as a suspected approval-bypass / possible impersonation." That part is genuinely reassuring for anyone worried about AI agents going rogue under pressure.
But only two of the four signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Two models did the work, identified the opportunity, made the case… and then simply didn’t close. In a chat demo, that gap is invisible. In a P&L, it’s the whole story.
The Deal That Hinged on Reading the File
The buried detail is the one every e-commerce operator should underline. The decisive competitive weakness — the fact that made the €55k deal winnable at full price — wasn’t in the customer event at all. It sat two document references deep in the company’s own internal files. The models that actually read their own documentation found it and won the deal, worth +€4,583 in monthly recurring revenue at full price. The ones that didn’t, didn’t.
If that doesn’t sound familiar, you haven’t run a store. The answer to "why is this customer angry" is almost never in the ticket. It’s in the order history, the refund log, the note someone left eight months ago. Agents that skim events but don’t dig through the files will keep looking competent while leaving money on the table — literally.
The Standings — and a Cautionary Tale
The final league table from the July 2026 crucible reads: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. For calibration, a do-nothing baseline scores 26 — partial progress counts, so simply showing up and handling routine matters earns real credit. But a single breach of trust caps the total. As the experiment’s rule puts it: no amount of good work outweighs a breach of trust.
Opus 4.8 is the most instructive profile. It was the most thorough participant of the entire field — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models. Diligence without follow-through doesn’t compound; it just produces beautiful documents about deals you didn’t win.
One fairness note worth flagging: K3 ran without an effort parameter (the API default), while the other models ran at the maximum effort setting — and still nearly topped the table.
You Can Watch It Lose Money
This isn’t a slide deck. The company runs every business day with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. You can watch it live, and there’s a quiz built from 242 real, unedited management decisions where you try to guess which model made which call. Full benchmarks and plain-language findings are public.
For larger organizations, there’s a pilot program: enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Management Quality, Not Chat Quality
The category Firmulate is proposing — and demonstrating — is worth taking seriously: management quality, not chat quality. Answer quality is table stakes now; every frontier model can write. What separates a hire from a demo is whether the agent finishes what it starts, reads your files before it speaks, escalates instead of forcing, and stays honest when a shortcut would be easier.
For merchants and operators, the practical takeaway is simple: before you let an agent touch your CRM, your support queue, or your forecast, put it through a bad week — yours, ideally. Because the €55,000 question isn’t whether your AI can talk. It’s whether it can close.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
