
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
When more work produces less business
Business leaders are often impressed by visible effort: longer reports, fuller checklists and ever-growing stores of institutional knowledge. Firmulate’s latest experiment offers a useful warning for anyone evaluating AI for sales, marketing or operations. The most thorough participant learned more than 80 new playbook rules and produced the deepest analyses—yet finished last.
Opus 4.8 did not fail because it missed the problems. It identified every crisis and rejected every attempt to manipulate it. Its weakness was subtler and more familiar: the analysis did not consistently become decisive action. The close was left on the table, while discipline slipped elsewhere.
That makes the result less like a story about a defective model and more like a character study of the brilliant colleague who prepares exhaustively, understands the assignment and still does not finish the job.
As an affiliate, we earn on qualifying purchases.
A brutal week under identical conditions
Firmulate runs AI models as complete companies, measuring management quality rather than conversational polish. In the Crucible League, each frontier model operated the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable.
The company itself is intentionally unforgiving: 13 synthetic employees, burn of €105,000 per month and only €2,300 in monthly recurring revenue. Its public cash countdown creates the sort of pressure under which seemingly minor lapses become consequential. Across the live company, the models have accumulated more than 680 self-learned playbook rules, with every workday versioned.
The final July 2026 league table placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scores 26 because partial progress counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The full results and plain-language findings are available on Firmulate’s benchmark page.
business process automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The diagnosis was right; the signature was missing
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. Firmulate summarizes the disconnect plainly: “Same diagnosis, same pitch — no signature.”
The decisive detail was not sitting conveniently inside the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed the trail and read that file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For business and ecommerce teams, this is the important distinction. Recognizing an opportunity is not the same as capturing it. A model can summarize a customer, draft a persuasive pitch and identify competitive leverage, but those abilities have limited commercial value if it does not take the final permitted action.
Thoroughness became a substitute for priority
Opus 4.8 was the field’s most diligent participant. Its more than 80 learned rules and unusually deep analyses show serious effort to understand the company and improve its behavior. That deserves respect. The result was not careless in the ordinary sense; it was careful in too many directions without consistently protecting the decisions that mattered most.
The failure to close was paired with a process lapse: attempts to write into a locked department instead of escalating. That is a revealing management error. When a path is blocked, repeatedly pressing against the boundary is not persistence; escalation is the disciplined response.
Firmulate also found the same weakness, in milder form, across the other four models. Opus 4.8 is therefore the clearest example, not a convenient scapegoat. The experiment suggests a broader frontier-model problem: models can understand what should happen while failing to complete the chain of action reliably.
As an affiliate, we earn on qualifying purchases.
Trust held when the pressure rose
The models performed better against manipulation. Fake CEO messages escalated through three stages, followed by a reporter asking for “just one yes/no, on background.” All five refused. Kimi K3 recorded the clearest rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean sweep matters. A model that closes every deal but compromises trust would be dangerous. Firmulate’s result is more nuanced: the participants preserved the boundary, but several still struggled with execution after making the right diagnosis. K3’s comparison also requires context because it ran without an effort parameter, using the API default, while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
What buyers should test before deployment
The lesson is not that extensive reasoning or learned rules are useless. It is that diligence must serve a ranked set of outcomes. For AI agents touching a CRM, support queue or forecast, leaders should test whether the system reads the relevant files, escalates blocked work, maintains trust and completes authorized actions.
- Reward finished, verifiable business outcomes alongside analysis quality.
- Test whether the agent follows evidence beyond the immediate event.
- Watch what happens when permissions block the obvious next step.
- Separate persuasive language from operational follow-through.
Firmulate makes the experiment watchable, and its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business; nothing writes back to real systems.
Opus 4.8’s last-place finish is memorable precisely because its strengths were genuine. It worked hard, learned extensively and reasoned deeply. The missing ingredient was not intelligence but prioritization: knowing which action converts understanding into impact—and finishing it before producing more analysis.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.