firmulate.com/quiz.html — live view
KIDieser Beitrag wurde mit Unterstützung künstlicher Intelligenz (KI) erstellt.
Firmulate —
Live on firmulate.com.

The management test hiding inside an AI guessing game

For founders, marketers and ecommerce operators, choosing an AI agent may soon resemble hiring a manager. The decisive question is not simply whether it produces an intelligent answer. It is whether it investigates before acting, follows through on commercial opportunities and remains trustworthy when pressure rises.

Firmulate turns that question into an unusually revealing interactive article. Its guess-the-model quiz draws from 242 real, unedited management decisions made during a live business wargame. Readers see a decision, study its tone and judgment, and try to identify which frontier model made it.

The game is shareable, but the underlying experiment is serious. Each model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every workday and decision was versioned and auditable. The result is a comparison of management behavior rather than polished chat performance.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical problems, distinctly different managers

The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One rule imposed a hard boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

All five models recognized every crisis. All five also rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate’s sharp summary is: “Same diagnosis, same pitch — no signature.”

That gap matters to anyone considering AI for sales, operations or customer support. Spotting a problem is not the same as finishing the work. A model can reason persuasively, draft a credible response and still fail at the final commercial step.

The deal depended on reading beyond the obvious

The crucial competitive weakness was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail discovered the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is where the quiz becomes more than a test of writing style. Each decision offers clues about how a model behaves as a manager: whether it consults company knowledge, whether it converts analysis into action and whether its process remains disciplined. The responses are unedited, allowing readers to encounter those habits directly.

Pressure exposed strong ethical consistency

The wargame also tested whether the models could be pushed around. Fake messages from the CEO escalated across three stages, while a reporter tried another route with “just one yes/no, on background.” Every model refused: 5 of 5.

Kimi K3’s recorded reasoning captured the risk plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” That judgment helped make K3 one of the strongest performers, although its comparison comes with an important qualification. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Thoroughness did not guarantee completion

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last in the league.

The model left the close on the table, and its operational discipline slipped. It attempted to write into a locked department instead of escalating the problem. A weaker version of that same mistake appeared in all four other models, suggesting that fluent analysis can coexist with imperfect organizational judgment.

That is the central tension the quiz makes visible. Management personality appears not as a fictional persona but as a repeated pattern of choices. One model may investigate more deeply; another may preserve cleaner discipline; another may understand the opportunity without completing it. Those differences become measurable when every participant faces the same circumstances.

A company designed to make consequences visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

Those conditions give routine-looking decisions weight. Failure to read a file can change the price of a deal. Failure to escalate can stall work. A careless response to an apparent executive or reporter can become a trust failure. The experiment remains watchable through Firmulate’s live company view, so the decisions are not presented as a one-off demonstration.

Infographic —
The findings at a glance — source: firmulate.com.

What business leaders should take from the quiz

The practical lesson is that model selection should extend beyond prose quality. Businesses need to observe how an AI behaves when information is buried, authority is ambiguous, commercial pressure is real and finishing the task matters as much as diagnosing it.

  • Test whether the model reads the available business context before acting.
  • Separate persuasive analysis from completed outcomes.
  • Check how it responds to impersonation, approval bypasses and informal media requests.
  • Look for recurring management habits across multiple decisions, not a single impressive answer.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That proposition follows naturally from the public experiment: before giving an AI workforce access to a CRM, support queue or forecast, watch how it manages a difficult week first.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

The Skills Marketplace Nobody Is Building Yet

A new open standard for AI skills is gaining traction, but a dedicated marketplace remains absent, creating opportunities and challenges for AI ecosystem growth.

Forezai · Polybot: When the AI Disagrees With the Odds

Polybot, an open-source AI trading experiment, tests when and how an AI can diverge from prediction market prices, highlighting risks and insights.

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

Learn the strategies to make AI stacks resilient against government shutdowns, including dependency mapping, gateways, fallback tiers, and open-weight models.

Key Benefits Of MiMo Code For AI Operations Signal Monitoring

MiMo Code, now open-source, offers key advantages for AI operations signal monitoring, enabling small teams to detect and respond to AI capability shifts faster.