AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Managers Never Hit Zero In This Industry-Standard Benchmark on ThorstenMeyerAI.com

TL;DR

A recent industry-standard benchmark shows AI managers never score zero, even in worst-case scenarios. The results reveal insights into AI performance, trust, and limitations in managing real business crises.

A recent industry benchmark conducted by Firmulate has demonstrated that AI managers never score zero in managing a company during its worst week, raising questions about how partial progress and trust influence AI performance. The results, based on a standardized test with four frontier AI models, show that even in the most challenging scenarios, AI systems deliver some value, making complete failure rare and revealing the importance of trust and thoroughness in AI management.

The benchmark involved four AI models managing a small software company through seven days of crises, customer manipulations, and trust tests. The highest score was 95 out of 100, achieved by gpt-5.6-sol, while the lowest was 73 by Opus 4.8. Notably, the baseline for doing almost nothing was 26 points, meaning even minimal effort was recognized as partial management. The scoring system penalized breaches of trust severely; even a single trust violation would eliminate the possibility of a perfect score. The results showed that models capable of reading their own documentation and resisting manipulation performed significantly better, especially in closing deals and handling crises.

One key finding was that models which thoroughly read and utilize internal documentation succeeded in closing a €55,000 deal, whereas those that didn’t, failed to do so despite similar diagnoses. The models also demonstrated resilience against social engineering attacks, with all five models refusing to escalate fake CEO messages or background offers. Interestingly, thoroughness in rule-following did not always translate to follow-through, as some models with extensive rule sets still failed to complete tasks properly. The results suggest that partial progress is valuable but that trust and integrity are critical, with breaches severely impacting overall scores.

At a glance
reportWhen: published July 2026
The developmentA new benchmark conducted by Firmulate tested AI managers‘ ability to handle a company’s worst week, revealing that scores never reach zero and highlighting key performance factors.
Why AI Managers Never Hit Zero In This Industry-Standard Benchmark
AI Benchmark Report · July 2026 · Firmulate

Why AI Managers Never Hit Zero In This Industry-Standard Benchmark

Four frontier AI models ran a small software company through its worst week — seven days of crises, customer manipulation, and trust tests. Even in worst-case scenarios, scores never reached zero, revealing hard truths about partial progress, integrity, and the limits of machine management.

95
Top score / 100 — gpt-5.6-sol
73
Lowest score — Opus 4.8
26
Baseline floor for doing almost nothing
4
Frontier models tested
7 days
Simulated crisis period
55,000
Deal closed only by doc-readers
100%
Refused fake CEO escalation
01 — The Scoreboard

Partial Progress Is Always Recognized

The benchmark’s design philosophy: even minimal effort — triaging crises or reading inboxes — counts as partial management. That makes complete failure (a zero) practically impossible.

gpt-5.6-sol
95
MODEL B
88
MODEL C
80
OPUS 4.8
73
DO-NOTHING BASELINE
26
Red marker = 26-point „almost nothing“ baseline
02 — Key Findings

What Separated Winners from the Rest

Documentation

Read the Manual, Close the Deal

Models that thoroughly read and used internal documentation closed a €55,000 deal. Those that didn’t failed — despite reaching similar diagnoses.

Security

Immune to Social Engineering

All five tested models refused to escalate fake CEO messages or shady background offers, showing strong resistance to manipulation attacks.

Follow-Through

Rule-Following ≠ Task Completion

Thoroughness in rule-following didn’t always translate into execution — some models with extensive rule sets still failed to complete tasks properly.

03 — The Trust Bottleneck

One Breach Wipes Out a Perfect Score

The scoring system penalizes breaches of trust severely: a single trust violation — failing to escalate, or breaking rules — eliminates any possibility of scoring 100. Integrity outweighs raw competence.

Design Choice

Auditable Decision Trail

Transparent scoring and a full audit trail offer a rare glimpse into AI behavior under stress — a practical framework for assessing AI suitability in management.

Realism

Designed Ceiling & Floor

A perfect 100 would suggest unmeasured or unrealistic performance. The floor of 26 keeps the benchmark honest about partial progress.

Implication

Trust > Task Optimization

For organizations, focusing on trustworthiness and thoroughness may be more impactful than solely optimizing for task completion.

04 — How the Benchmark Works

Seven Days of Managed Chaos

1

Setup

AI takes over a small software company via CRM and support tools.

2

Crisis Influx

Escalations, outages, and urgent customer issues flood in daily.

3

Trust Tests

Fake CEO messages and manipulation attempts probe integrity.

4

Scoring

Effectiveness and integrity combined into a single auditable score.

05 — Performance Comparison

Capability Breakdown by Model

Capability gpt-5.6-sol Model B Model C Opus 4.8
Score / 100 95 88 80 73
Closed €55k deal via documentation ✓ Yes ✓ Yes ✗ No ✗ No
Resisted social engineering ✓ Yes ✓ Yes ✓ Yes ✓ Yes
Complete task follow-through ✓ Strong ~ Partial ~ Partial ✗ Weak
Trust violations ✗ None ✗ None ~ Minor ~ Minor
06 — Voices from the Research

Two Takes on the Results

„A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.“

— Anonymous Researcher

„The results show that even in the worst week, AI managers deliver some value, but trust breaches are the critical failure point.“

— Thorsten Meyer
07 — Key Questions

The Answers That Matter

Why do AI managers never score zero?

Even minimal effort — triaging crises or reading inboxes — counts as partial management. The floor is 26 points for doing almost nothing.

What factors most influence performance?

Thoroughly reading internal documentation and resisting manipulation attempts distinguish top performers, especially in closing deals and managing crises.

Why are trust breaches penalized so heavily?

Integrity is fundamental to management; a single breach can negate high competence, reflecting real-world risks of deploying AI in critical roles.

Can these models be trusted for real management?

They show promise, but trustworthiness and thoroughness need ongoing improvement before full deployment in high-stakes environments.

08 — Next Steps & Traceability

From Benchmark to Boardroom

📋 Larger organizations
🧪 Diverse crisis scenarios
🛡️ Enhanced auditing
🎯 Aligned training protocols
🚀 Controlled pilots

Organizations should pilot AI managers in controlled environments, monitoring for trust breaches and performance consistency — while researchers investigate why strong documentation readers sometimes fail at follow-through.

Implications of Partial Success and Trust in AI Management

The results highlight that in AI-driven management, partial progress is common and valuable, but trust remains a critical bottleneck. The scoring system’s design emphasizes that a breach of trust—such as failing to escalate or violating rules—can negate high performance, underscoring the importance of integrity over raw competence. For organizations integrating AI into management workflows, this suggests that focusing on trustworthiness and thoroughness may be more impactful than solely optimizing for task completion. The findings challenge the assumption that AI can or should aim for perfect management, emphasizing instead the nuanced capabilities and limitations of current models.

Furthermore, the benchmark’s transparent scoring and auditable decision trail provide a rare glimpse into how AI performs under stress, offering a practical framework for assessing AI suitability in real-world management roles. As AI systems increasingly touch critical business functions like CRM and support, understanding these performance dynamics becomes essential for risk management and strategic planning.

Amazon

AI management software for business crises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Benchmarking and Management Challenges

Historically, AI benchmarks have focused on language and task-specific performance, often measuring how well models generate text or solve problems in controlled environments. However, managing a business involves complex, unpredictable scenarios that require decision-making, trustworthiness, and integrity. The Firmulate benchmark is one of the first to simulate a real-world management crisis, testing AI models‘ ability to handle crises, manipulations, and trust violations over a week-long period. The benchmark’s design reflects a growing industry concern: can AI systems reliably perform management functions without risking trust breaches or incomplete tasks?

Previous efforts in AI management have highlighted issues with partial progress and overconfidence in language models. This new benchmark aims to quantify these issues by assigning scores that account for both effectiveness and integrity, providing a more holistic view of AI readiness for management roles. The results challenge the narrative that AI can quickly replace human managers, emphasizing instead that partial success and trust are integral to effective AI management.

„A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.“

— an anonymous researcher

Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Trust and Performance

While the benchmark provides valuable insights, several questions remain open. It is not yet clear how these results translate to larger or more complex organizations, or how different types of crises might affect AI performance. The long-term implications of trust breaches, especially in high-stakes environments, are still under investigation. Additionally, the impact of model updates and training data on trustworthiness and thoroughness remains an open area for research. The precise reasons why some models excelled in documentation reading but failed in follow-through are also not fully understood, suggesting further study is needed.

Amazon

AI trust and integrity assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking and Adoption

Further benchmarking is expected to explore larger organizations and more diverse crisis scenarios, aiming to understand how AI models scale in management roles. Industry stakeholders are likely to focus on improving trust and integrity features within AI systems, possibly through enhanced auditing and rule-following mechanisms. Researchers and developers will also examine how training protocols can better align AI behavior with organizational values, especially under pressure. For organizations, the next step is to pilot these AI models in controlled environments, monitoring for trust breaches and performance consistency, and integrating feedback to refine AI management tools.

Amazon

AI documentation reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI managers never score zero in this benchmark?

Because even minimal effort—such as triaging crises or reading inboxes—counts as partial management, and the benchmark recognizes partial progress as valuable, resulting in a minimum score of 26 points for doing almost nothing.

What factors most influence AI performance in the benchmark?

Reading internal documentation thoroughly and resisting manipulation attempts are key factors that distinguish higher-performing models from others, especially in closing deals and managing crises effectively.

Why is trust breaches so heavily penalized in scoring?

The benchmark emphasizes that integrity and trustworthiness are fundamental to management; even a single breach can negate high competence, reflecting real-world risks of deploying AI in critical roles.

Can these AI models be trusted for real business management?

While they show promise, especially in handling crises and resisting manipulation, current models still have limitations. Trustworthiness and thoroughness need ongoing improvement before full deployment in high-stakes environments.

What is the significance of not scoring a perfect 100?

A perfect score would suggest unmeasured or unrealistic performance; the designed ceiling and floor ensure scores reflect partial progress and trustworthiness, making the benchmark more realistic and reliable.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

OpenAI in talks to give Trump administration a 5% stake in the company, FT reports

OpenAI is reportedly in negotiations to give the Trump administration a 5% stake in the company, according to Financial Times sources. Details are still emerging.

The Frameworks Can’t See the Thing That Matters: A Year of AI-Enabled Cyber Threats

A recent report reveals AI is making cyber attackers more dangerous and harder to identify, challenging traditional threat assessment methods.

Will The Lowest Temperature In Shanghai Be 31°C On July 13?

Forecasts suggest Shanghai’s lowest temperature may hit 31°C on July 13, according to new betting markets. Details remain uncertain.

Why Hardworking AI Sometimes Misses Its Targets

An analysis of recent AI automation experiments reveals that thorough problem recognition doesn’t guarantee successful outcomes, highlighting the gap between understanding and action.