AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

When an AI knows how to sell, will it actually close?

For business owners, a polished pitch is not the same as revenue. The more consequential test is whether an AI can recognize a customer opportunity, follow the evidence and make the call. Firmulate’s live company experiment puts that question inside a small software business facing a week of crises, commercial pressure and tempting shortcuts.

A company under pressure

Firmulate runs a synthetic business with real money mechanics: 13 employees, monthly burn of €105,000 against €2,300 in monthly recurring revenue, and a public cash countdown. Its playbooks contain more than 680 self-learned rules, and each workday is versioned. The company is live and watchable at firmulate.com.

In the Crucible League’s final, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Every decision was versioned and auditable. The final ranking put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The gap between advice and action

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s shorthand for the gap: “Same diagnosis, same pitch — no signature.” It is a useful distinction for anyone considering AI in sales or operations: recognizing the right move does not guarantee the system will carry it through.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The useful evidence was there; the challenge was finding and acting on it.

Trust and discipline under pressure

The social-engineering test escalated through three fake CEO messages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unclosed and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. The lesson for a business buyer is concrete: extensive analysis is not enough if follow-through and boundaries falter.

There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can guess which model made each choice at firmulate.com.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to testing your own business

A benchmark can show how models behave in a shared scenario. A pilot asks what happens against the details and playbooks of your own company. Firmulate says enterprises can run the same kind of wargame using a read-only export, test crisis scenarios and receive a board report ranking models and identifying weak points in their playbooks. Nothing writes back to real systems.

For businesses weighing AI in sales, support or operations, that offers a way to examine decisions before putting an AI workforce to work. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

The Labor Displacement Data: What Q1-Q2 2026 Actually Shows

New data from Q1-Q2 2026 shows significant AI-driven layoffs in tech, with material impacts on specific worker cohorts but limited overall employment decline.

Indusface Introduces SwyftComply AI And Launches A New Era Of Application Security With Autonomous Vulnerability Remediation

Indusface launches SwyftComply AI, enabling autonomous vulnerability remediation to enhance application security. The development marks a new era in cybersecurity.

Beyond The Demo: The AI Leaderboard That Defines The Market

Firmulate’s live AI benchmark tests models in real company scenarios, revealing management capabilities beyond traditional chat and coding assessments.

Évian and the Fallout: What Europe Actually Wants From Amodei, Hassabis, and Altman

Europe pushes for reliable AI access, sovereignty, and safety at G7 summit with Amodei, Hassabis, and Altman, amid U.S. export controls and geopolitical tensions.