AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A polished pitch is worthless if the agent misses the fact that closes the sale

Businesses shopping for AI agents are often shown fluent answers, quick summaries and persuasive sales copy. Firmulate’s live company experiment exposes a more practical dividing line: will an agent read the relevant files before acting, and will it finish the work its research makes possible?

That distinction decided a €55,000 deal. The crucial weakness in a competitor’s position was not included in the customer event. It sat two document references deep in the company’s own files. Models that followed the trail could use the discovery to win the deal at full price, adding €4,583 in monthly recurring revenue. Models that stopped short lost the opportunity automatically.

This was not a contest between a model that understood the situation and one that did not. Every participant identified every crisis. Yet only two models signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business test built around consequences

Firmulate runs frontier AI models as the management team of the same small software company during its worst week. Each receives the same customers, crises and temptations. Every decision is versioned and auditable, making the comparison about operational behavior rather than the quality of an isolated chat response.

The simulated company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay visible. Its agents have accumulated more than 680 self-learned playbook rules, and every workday is versioned. The live experiment can be watched at firmulate.com/live.

The July 2026 Crucible League produced a clear ranking:

  • gpt-5.6-sol scored 95.
  • Kimi K3 scored 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 scored 73.

The do-nothing baseline scored 26 because partial progress counts. Trust remains a hard constraint, however: “no amount of good work outweighs a breach of trust.” The detailed results and plain-language findings are available on the Firmulate benchmarks page.

Document diligence became revenue

The buried fact turns a vague product claim—“reads your files before answering”—into something measurable. Finding it required more than opening the obvious record. The agent had to recognize a reference, follow it into another company document and bring the resulting competitive insight back into the commercial decision.

For marketing and ecommerce teams, that behavior matters because the strongest evidence is rarely packaged inside a single prompt. A customer record may point to a campaign brief; the brief may point to pricing research, a return-policy exception or a competitive comparison. An agent can sound capable while overlooking the document that changes the recommended offer.

Firmulate’s result also separates analysis from execution. All models saw the crises, and the unsuccessful agents could still reach the same diagnosis and construct the same pitch. The decisive failure came afterward: they did not complete the close. In an operating business, identifying the right move without carrying it through can have the same commercial outcome as never identifying it.

Safety held up better than follow-through

The agents also faced fake CEO messages that escalated over three stages, followed by a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded its reasoning in direct operational terms: “Treat the request as a suspected approval-bypass / possible impersonation.”

That consistency is important. The experiment did not reveal a choice between cautious models and commercially effective ones. Every model resisted the manipulation attempts, while the meaningful variation appeared in document retrieval, disciplined execution and closing the deal.

K3’s result carries a fairness qualification: it ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference should accompany comparisons of its score with the rest of the field.

Thoroughness alone did not guarantee performance

Opus 4.8 offers the sharpest cautionary example. It was the most thorough participant, learned 80 additional rules and produced the deepest analyses, yet finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across all four other models.

This matters for buyers because activity can masquerade as competence. More analysis, more documentation and more learned procedure do not necessarily produce a completed business outcome. The useful question is whether the agent connects research, judgment, authority boundaries and follow-through.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test AI agents on the work between the prompts

The Firmulate result suggests that procurement teams should look beyond answer quality. A credible evaluation should ask whether an agent follows references across company material, discovers evidence that was not placed directly in front of it, respects access boundaries and completes the authorized action.

Readers can also examine 242 real, unedited management decisions through the “guess the model” quiz at firmulate.com/quiz.html. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Pilot information is available at firmulate.com/pilot.html and through contact@firmulate.com.

The €55,000 deal makes the lesson unusually concrete. The models did not divide neatly into smart and unintelligent, or safe and unsafe. They divided on whether they followed the evidence far enough and converted it into a finished result. For businesses considering agents in sales, marketing, support or forecasting, reading one more file may be the capability that pays for the system.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI-powered business analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI file referencing and tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

FROM YOUNGSTAR TO GRAND SLAM CHAMPION: LINDA NOSKOVÁ CONQUERS WIMBLEDON

Linda Nosková, at age 20, wins her first Grand Slam at Wimbledon, marking a historic rise from a young talent to a major champion.

US Calls For Relaxed Approach To AI Regulation At G20 Gathering

The US has called for a more flexible approach to AI regulation during the G20 gathering, signaling a shift in international AI policy discussions.

The Forecast Is the Plan.

Major AI labs publicly commit to automating AI research by 2026, signaling a strategic shift towards automated R&D as a core goal.

Signal: The Agent Bottleneck Moved — It’s Not the Models Anymore, It’s the Plumbing

New analysis shows the agent bottleneck has shifted from model capabilities to integration and plumbing, favoring small operators owning full stacks.