AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

When polished AI advice meets an actual revenue decision

For business, marketing and ecommerce leaders, generative AI’s writing ability is becoming the least interesting part of the story. A model can draft a campaign, diagnose a customer problem and produce an impressive sales pitch. The harder question is whether it will finish the work, protect trust and make the commercially important decision when the pressure is real.

Firmulate turns that question into something readers can inspect—and play. Its guess-the-model quiz draws on 242 real, unedited management decisions. Each decision came from a live experiment in which frontier models ran the same small software company through its worst week, facing the same customers, crises and temptations.

The quiz is entertaining because the models do not merely sound different. Their decisions reveal distinct management personalities. Some responses are exhaustive; others are sharply concise. Some models identify the correct opportunity but fail to complete it. The result feels less like a writing comparison and more like reviewing a set of executives under pressure.

Amazon

AI decision-making simulation game

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same problems produced very different companies

Every decision in the experiment was versioned and auditable. That matters because Firmulate is not asking readers to judge a cherry-picked chatbot answer. The models encountered identical business conditions, making their differences easier to see.

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26, with partial progress counting toward that result.

The broad competence was impressive: every model spotted every crisis and refused every manipulation attempt. Yet that common awareness concealed the experiment’s most commercially revealing result. Only two models signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”

That gap should feel familiar to anyone running sales, marketing or ecommerce operations. Recognizing a high-value action is not equivalent to taking it. A model may produce convincing reasoning while still leaving the decisive step unfinished. In a business workflow, completion is not a stylistic preference; it is part of the outcome.

The winning clue was buried in the company’s own files

The decisive competitor weakness did not appear in the customer event. It sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, worth +€4,583 MRR.

This finding shifts the practical question for companies adopting AI. The challenge is not simply whether a model understands an incoming request. It is whether the model reads the available business context before responding. Customer histories, internal notes and product records may contain the fact that changes a negotiation. A capable model that stops at the visible event can still miss the strongest commercial move.

Pressure revealed discipline as well as intelligence

The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 described its response in unusually direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean refusal helps explain why the experiment is relevant beyond sales. AI systems operating near customer records, forecasts or support queues must handle persuasion without quietly abandoning organizational boundaries. In this test, the entire field recognized the manipulation attempts.

K3’s result does require a fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it finished second with 93 and showed the cleanest discipline of the field.

Thoroughness did not guarantee execution

Opus 4.8 offers the clearest warning against equating depth with management quality. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last with 73. The deal close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.

This is precisely what makes the quiz more than a novelty. Readers begin to recognize behavioral signatures: expansive analysis, concise escalation, cautious refusal or an unfinished commercial action. These are not fictional characters written for entertainment. They are patterns visible across real, unedited decisions from the same operating test.

A company built to make behavior visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing the experiment to be watched as an operating business rather than presented afterward as a polished demonstration.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business management AI training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What business leaders should take from the quiz

The leaderboard is useful, but the deeper lesson is that model selection is a management decision. A model may be excellent at diagnosis and still fail to close. It may be highly thorough while losing operational discipline. It may also remain dependable when an apparent executive or reporter pressures it to bypass normal approval.

Firmulate’s enterprise pilot extends the same approach to a read-only export of a company’s own business. Nothing writes back to real systems. That offers a practical way to examine model behavior against familiar customers, documents and commercial situations before granting an AI workforce meaningful responsibility.

For readers, the quiz provides the immediate test: can you identify a model from the way it manages? The harder question comes afterward. If these models have measurable management personalities, which one would you trust with your customers, your company knowledge and the final step of a deal?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI workflow decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Agent Test Succeeds In Finding Hidden Data

An AI agent successfully located concealed information in company files, enabling a €55,000 deal and demonstrating advanced file-reading capabilities.

Why Hardworking AI Sometimes Misses Its Targets

An analysis of recent AI automation experiments reveals that thorough problem recognition doesn’t guarantee successful outcomes, highlighting the gap between understanding and action.

How Close We Were To Missing An Essential AI Alert

An in-depth look at how AI agents nearly bypassed an essential security alert, revealing vulnerabilities in AI training and oversight.

The Financial Benefits Of Using Claude Opus 5.5 In AI Development

Anthropic’s Claude Opus 5.5 reduces AI development costs by up to 40%, improves efficiency, and enhances performance in knowledge work and coding tasks.