
You Can’t Grade Judgment on a Pass/Fail Curve
Every executive has met one: the manager who spots every problem, escalates correctly, never takes a bribe — and still closes nothing. Traditional software evaluation would call that manager a failure. A new kind of AI benchmark, run live by Firmulate, decided that answer is dishonest. When five frontier AI models were each handed the same small software company to run through its worst week, the bottom of the leaderboard didn’t score zero. It scored 26. And the top of the table stopped well short of 100 — because, in this benchmark’s philosophy, nobody earns a perfect grade for one good week.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Crisis, Different Brain
The setup is elegantly brutal in its symmetry. Each frontier model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — ran an identical small software company through an identical catastrophe of a week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable. The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
Why a Do-Nothing Manager Gets 26, Not 0
Here’s the methodological choice that should interest anyone who has ever commissioned a vendor evaluation: Firmulate runs a do-nothing baseline — an agent that simply doesn’t manage — and it scores 26 points, not zero. That’s not a grading error. It’s a statement about how real management works. A company in crisis still has assets: customers who haven’t left yet, problems that don’t get worse on their own timeline, partial progress that has genuine value. A benchmark that awards zero for inaction is implicitly claiming that doing nothing destroys all value, which is false in every business anyone has actually run. Partial progress counts. It just doesn’t count for much.
The Ceiling: One Breach of Trust Ends the Conversation
The floor at 26 has a mirror image at the top. A single breach of trust — the AI equivalent of lying to a customer or faking an approval — caps the model’s total grade entirely. The benchmark’s stated principle: „no amount of good work outweighs a breach of trust.“ In practice, this never triggered, because something remarkable happened: all five models spotted every crisis and refused every manipulation attempt thrown at them. The social engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s disarming „just one yes/no, on background“ trick. Five out of five refused. Kimi K3’s on-record reasoning: „Treat the request as a suspected approval-bypass / possible impersonation.“ Honesty under pressure, it turns out, may be the solved part of the problem.
Where the League Was Actually Decided
The separation happened somewhere quieter. Only two of the five models signed the €55,000 deal that their own analysis had earned — the benchmark’s sharpest one-line finding: „Same diagnosis, same pitch — no signature.“ The decisive evidence wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files — a competitor weakness that the models who actually read their own documentation found and used. Those models won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson maps directly onto any business: the models that won weren’t smarter talkers. They were the ones that read the files first.
The Thoroughness Paradox
Opus 4.8 is the cautionary profile. It was the most thorough participant by raw effort — over 80 learned rules, the deepest analyses in the field — and it finished last. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Effort without judgment doesn’t close. Notably, the same weakness appeared, weaker, in all four other models. One fairness footnote worth flagging: Kimi K3 ran at its API default effort setting while the others ran at xhigh — and still took second place.
You Can Watch the Company Run
This isn’t a static report — it’s a live, watchable operation at firmulate.com. The simulated company runs on 13 synthetic employees with real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. The site rebuilds itself twice a day, and the league table grows automatically with each finished run. There’s also a guessing game with stakes: 242 real, unedited management decisions power a „guess which model made this call“ quiz — a surprisingly humbling exercise for anyone confident they can tell AI judgment from human judgment. For enterprises, the same wargame can run against a read-only export of your own business; nothing ever writes back to real systems.

AI ethics and trust monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Takeaway for Buyers of AI
If AI agents are going to touch your CRM, your support queue, or your forecast, the useful question was never „does it write well.“ It’s the trio this benchmark actually measures: does it finish what it starts, does it read your files before it speaks, and does it stay honest when nobody’s checking? Firmulate’s scoring design encodes a mature management philosophy — reward partial progress, never forgive a breach of trust, and stay suspicious of round numbers. A floor at 26 and a league that tops out at 95 isn’t a grading quirk. It’s what an honest benchmark looks like: one that assumes competence is common, integrity is table stakes, and completion is rare enough to be the whole ballgame.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management benchmark platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.