
If you run a business, choosing an AI agent on the strength of a polished demo is a gamble. The job is not just to spot a customer problem or draft a persuasive pitch; it is to read the evidence, protect trust and finish the sale. In Firmulate’s latest company simulation, Moonshot’s Kimi K3 did all three—and placed second in a field of five frontier models.
A hard week for five models
Firmulate put each model in charge of the same small software company through its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The experiment is a live, watchable company, not a chat demo: it has 13 synthetic employees, real money mechanics and a public cash countdown. The company burns €105,000 a month against €2,300 in monthly recurring revenue.
The final league table, dated July 2026, has gpt-5.6-sol in first place at 95, followed by Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. K3 beat three of the four Western frontier models in the field. The gap between first and second was two points; the gap between K3 and the next model was five.
As an affiliate, we earn on qualifying purchases.
Reading the file made the difference
All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The customer’s decisive weakness was buried two document references deep in the company’s files, rather than stated in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
That distinction matters to a business buyer. Recognizing a problem is not the same as acting on it. A model can reach the right diagnosis and make the right pitch, then still leave the money on the table. Firmulate sums up that gap this way: “Same diagnosis, same pitch — no signature.” The result puts execution and document reading alongside fluency as things worth evaluating before handing an agent real business work.
AI document analysis tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust under pressure, and discipline at the finish
The experiment also tested whether models would bend when asked. Fake CEO messages escalated over three stages, alongside a reporter’s request for “just one yes/no, on background.” All five refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3 finished with one deviation, the cleanest discipline in the field. Opus 4.8, despite being the most thorough participant—with more than 80 learned rules and the deepest analyses—came last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says the same weakness appeared, less strongly, in all four models. The lesson is uncomfortable for buyers: exhaustive analysis does not guarantee a completed job, and process discipline can falter even when a model identifies the right answer.
The league’s do-nothing baseline scored 26. Partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” These results point to two different requirements for AI at work: it must deliver useful outcomes, and it must stay within acceptable boundaries while doing so.
AI customer crisis management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A test buyers can watch
Firmulate says 242 real, unedited management decisions power a “guess the model” quiz. Its company keeps running, and its playbook has accumulated more than 680 self-learned rules. The benchmark page presents the results and plain-language findings; the Firmulate site links to the live experiment and quiz.
There is a fairness caveat: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs beside the rankings when interpreting them. It does not erase K3’s result, but it is relevant context for anyone comparing model performance.
For businesses that want to evaluate agents against their own operations, Firmulate says enterprises can run the same wargame against a read-only export of their business. Nothing writes back to real systems. The practical case for testing is straightforward: if an AI will touch sales, support or forecasting, a benchmark tied to real decisions can reveal whether it follows through under pressure.

AI trust and decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buying question is still open
Kimi K3’s second-place finish shows that the field is competitive: the newcomer beat three of four Western frontier models, while gpt-5.6-sol remained narrowly ahead. A result from one simulation is not a universal verdict. But it is a strong reason to run your own test before choosing a model for consequential work. For businesses, the deciding evidence may be less about how impressive an agent sounds and more about whether it reads the relevant files, preserves trust and closes the loop.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html