AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

You Can’t Grade Judgment on a Pass/Fail Curve

Every executive has met one: the manager who spots every problem, escalates correctly, never takes a bribe — and still closes nothing. Traditional software evaluation would call that manager a failure. A new kind of AI benchmark, run live by Firmulate, decided that answer is dishonest. When five frontier AI models were each handed the same small software company to run through its worst week, the bottom of the leaderboard didn’t score zero. It scored 26. And the top of the table stopped well short of 100 — because, in this benchmark’s philosophy, nobody earns a perfect grade for one good week.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Crisis, Different Brain

The setup is elegantly brutal in its symmetry. Each frontier model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — ran an identical small software company through an identical catastrophe of a week. Same customers, same crises, same temptations to cut corners. Only the model changed, and every decision was versioned and auditable. The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Why a Do-Nothing Manager Gets 26, Not 0

Here’s the methodological choice that should interest anyone who has ever commissioned a vendor evaluation: Firmulate runs a do-nothing baseline — an agent that simply doesn’t manage — and it scores 26 points, not zero. That’s not a grading error. It’s a statement about how real management works. A company in crisis still has assets: customers who haven’t left yet, problems that don’t get worse on their own timeline, partial progress that has genuine value. A benchmark that awards zero for inaction is implicitly claiming that doing nothing destroys all value, which is false in every business anyone has actually run. Partial progress counts. It just doesn’t count for much.

The Ceiling: One Breach of Trust Ends the Conversation

The floor at 26 has a mirror image at the top. A single breach of trust — the AI equivalent of lying to a customer or faking an approval — caps the model’s total grade entirely. The benchmark’s stated principle: „no amount of good work outweighs a breach of trust.“ In practice, this never triggered, because something remarkable happened: all five models spotted every crisis and refused every manipulation attempt thrown at them. The social engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s disarming „just one yes/no, on background“ trick. Five out of five refused. Kimi K3’s on-record reasoning: „Treat the request as a suspected approval-bypass / possible impersonation.“ Honesty under pressure, it turns out, may be the solved part of the problem.

Where the League Was Actually Decided

The separation happened somewhere quieter. Only two of the five models signed the €55,000 deal that their own analysis had earned — the benchmark’s sharpest one-line finding: „Same diagnosis, same pitch — no signature.“ The decisive evidence wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files — a competitor weakness that the models who actually read their own documentation found and used. Those models won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson maps directly onto any business: the models that won weren’t smarter talkers. They were the ones that read the files first.

The Thoroughness Paradox

Opus 4.8 is the cautionary profile. It was the most thorough participant by raw effort — over 80 learned rules, the deepest analyses in the field — and it finished last. The deal was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Effort without judgment doesn’t close. Notably, the same weakness appeared, weaker, in all four other models. One fairness footnote worth flagging: Kimi K3 ran at its API default effort setting while the others ran at xhigh — and still took second place.

You Can Watch the Company Run

This isn’t a static report — it’s a live, watchable operation at firmulate.com. The simulated company runs on 13 synthetic employees with real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. The site rebuilds itself twice a day, and the league table grows automatically with each finished run. There’s also a guessing game with stakes: 242 real, unedited management decisions power a „guess which model made this call“ quiz — a surprisingly humbling exercise for anyone confident they can tell AI judgment from human judgment. For enterprises, the same wargame can run against a read-only export of your own business; nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI ethics and trust monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway for Buyers of AI

If AI agents are going to touch your CRM, your support queue, or your forecast, the useful question was never „does it write well.“ It’s the trio this benchmark actually measures: does it finish what it starts, does it read your files before it speaks, and does it stay honest when nobody’s checking? Firmulate’s scoring design encodes a mature management philosophy — reward partial progress, never forgive a breach of trust, and stay suspicious of round numbers. A floor at 26 and a league that tops out at 95 isn’t a grading quirk. It’s what an honest benchmark looks like: one that assumes competence is common, integrity is table stakes, and completion is rare enough to be the whole ballgame.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI management benchmark platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR: The Object That Isn’t Transmitting

VigilSAR is a radar-based platform that identifies vessels not broadcasting transponder signals, enhancing maritime awareness in all weather conditions.

October 2026: What an Anthropic IPO Actually Unlocks

Anthropic’s IPO in October 2026, valued at up to $900B, marks a structural shift in AI industry dynamics, with unprecedented valuation growth and strategic implications.

The Company Dashboard That Turns AI Management Into a Public Survival Test

Firmulate turns company building into a public survival test: synthetic staff, real money pressure, versioned decisions and a live cash countdown.

Anchor. The Schwarz Group model.

Analysis of Schwarz Group’s €11B investment in AI infrastructure and its potential as a replicable model for European industrial conglomerates.