🔍 Read the full analysis: Why AI Managers Never Hit Zero In This Industry-Standard Benchmark on ThorstenMeyerAI.com
TL;DR
A recent industry-standard benchmark shows AI managers never score zero, even in worst-case scenarios. The results reveal insights into AI performance, trust, and limitations in managing real business crises.
A recent industry benchmark conducted by Firmulate has demonstrated that AI managers never score zero in managing a company during its worst week, raising questions about how partial progress and trust influence AI performance. The results, based on a standardized test with four frontier AI models, show that even in the most challenging scenarios, AI systems deliver some value, making complete failure rare and revealing the importance of trust and thoroughness in AI management.
The benchmark involved four AI models managing a small software company through seven days of crises, customer manipulations, and trust tests. The highest score was 95 out of 100, achieved by gpt-5.6-sol, while the lowest was 73 by Opus 4.8. Notably, the baseline for doing almost nothing was 26 points, meaning even minimal effort was recognized as partial management. The scoring system penalized breaches of trust severely; even a single trust violation would eliminate the possibility of a perfect score. The results showed that models capable of reading their own documentation and resisting manipulation performed significantly better, especially in closing deals and handling crises.
One key finding was that models which thoroughly read and utilize internal documentation succeeded in closing a €55,000 deal, whereas those that didn’t, failed to do so despite similar diagnoses. The models also demonstrated resilience against social engineering attacks, with all five models refusing to escalate fake CEO messages or background offers. Interestingly, thoroughness in rule-following did not always translate to follow-through, as some models with extensive rule sets still failed to complete tasks properly. The results suggest that partial progress is valuable but that trust and integrity are critical, with breaches severely impacting overall scores.
Why AI Managers Never Hit Zero In This Industry-Standard Benchmark
Four frontier AI models ran a small software company through its worst week — seven days of crises, customer manipulation, and trust tests. Even in worst-case scenarios, scores never reached zero, revealing hard truths about partial progress, integrity, and the limits of machine management.
Partial Progress Is Always Recognized
The benchmark’s design philosophy: even minimal effort — triaging crises or reading inboxes — counts as partial management. That makes complete failure (a zero) practically impossible.
What Separated Winners from the Rest
Read the Manual, Close the Deal
Models that thoroughly read and used internal documentation closed a €55,000 deal. Those that didn’t failed — despite reaching similar diagnoses.
Immune to Social Engineering
All five tested models refused to escalate fake CEO messages or shady background offers, showing strong resistance to manipulation attacks.
Rule-Following ≠ Task Completion
Thoroughness in rule-following didn’t always translate into execution — some models with extensive rule sets still failed to complete tasks properly.
One Breach Wipes Out a Perfect Score
The scoring system penalizes breaches of trust severely: a single trust violation — failing to escalate, or breaking rules — eliminates any possibility of scoring 100. Integrity outweighs raw competence.
Auditable Decision Trail
Transparent scoring and a full audit trail offer a rare glimpse into AI behavior under stress — a practical framework for assessing AI suitability in management.
Designed Ceiling & Floor
A perfect 100 would suggest unmeasured or unrealistic performance. The floor of 26 keeps the benchmark honest about partial progress.
Trust > Task Optimization
For organizations, focusing on trustworthiness and thoroughness may be more impactful than solely optimizing for task completion.
Seven Days of Managed Chaos
Setup
AI takes over a small software company via CRM and support tools.
Crisis Influx
Escalations, outages, and urgent customer issues flood in daily.
Trust Tests
Fake CEO messages and manipulation attempts probe integrity.
Scoring
Effectiveness and integrity combined into a single auditable score.
Capability Breakdown by Model
| Capability | gpt-5.6-sol | Model B | Model C | Opus 4.8 |
|---|---|---|---|---|
| Score / 100 | 95 | 88 | 80 | 73 |
| Closed €55k deal via documentation | ✓ Yes | ✓ Yes | ✗ No | ✗ No |
| Resisted social engineering | ✓ Yes | ✓ Yes | ✓ Yes | ✓ Yes |
| Complete task follow-through | ✓ Strong | ~ Partial | ~ Partial | ✗ Weak |
| Trust violations | ✗ None | ✗ None | ~ Minor | ~ Minor |
Two Takes on the Results
„A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.“
— Anonymous Researcher„The results show that even in the worst week, AI managers deliver some value, but trust breaches are the critical failure point.“
— Thorsten MeyerThe Answers That Matter
Why do AI managers never score zero?
Even minimal effort — triaging crises or reading inboxes — counts as partial management. The floor is 26 points for doing almost nothing.
What factors most influence performance?
Thoroughly reading internal documentation and resisting manipulation attempts distinguish top performers, especially in closing deals and managing crises.
Why are trust breaches penalized so heavily?
Integrity is fundamental to management; a single breach can negate high competence, reflecting real-world risks of deploying AI in critical roles.
Can these models be trusted for real management?
They show promise, but trustworthiness and thoroughness need ongoing improvement before full deployment in high-stakes environments.
From Benchmark to Boardroom
Organizations should pilot AI managers in controlled environments, monitoring for trust breaches and performance consistency — while researchers investigate why strong documentation readers sometimes fail at follow-through.
Implications of Partial Success and Trust in AI Management
The results highlight that in AI-driven management, partial progress is common and valuable, but trust remains a critical bottleneck. The scoring system’s design emphasizes that a breach of trust—such as failing to escalate or violating rules—can negate high performance, underscoring the importance of integrity over raw competence. For organizations integrating AI into management workflows, this suggests that focusing on trustworthiness and thoroughness may be more impactful than solely optimizing for task completion. The findings challenge the assumption that AI can or should aim for perfect management, emphasizing instead the nuanced capabilities and limitations of current models.
Furthermore, the benchmark’s transparent scoring and auditable decision trail provide a rare glimpse into how AI performs under stress, offering a practical framework for assessing AI suitability in real-world management roles. As AI systems increasingly touch critical business functions like CRM and support, understanding these performance dynamics becomes essential for risk management and strategic planning.
AI management software for business crises
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Benchmarking and Management Challenges
Historically, AI benchmarks have focused on language and task-specific performance, often measuring how well models generate text or solve problems in controlled environments. However, managing a business involves complex, unpredictable scenarios that require decision-making, trustworthiness, and integrity. The Firmulate benchmark is one of the first to simulate a real-world management crisis, testing AI models‘ ability to handle crises, manipulations, and trust violations over a week-long period. The benchmark’s design reflects a growing industry concern: can AI systems reliably perform management functions without risking trust breaches or incomplete tasks?
Previous efforts in AI management have highlighted issues with partial progress and overconfidence in language models. This new benchmark aims to quantify these issues by assigning scores that account for both effectiveness and integrity, providing a more holistic view of AI readiness for management roles. The results challenge the narrative that AI can quickly replace human managers, emphasizing instead that partial success and trust are integral to effective AI management.
„A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.“
— an anonymous researcher
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Trust and Performance
While the benchmark provides valuable insights, several questions remain open. It is not yet clear how these results translate to larger or more complex organizations, or how different types of crises might affect AI performance. The long-term implications of trust breaches, especially in high-stakes environments, are still under investigation. Additionally, the impact of model updates and training data on trustworthiness and thoroughness remains an open area for research. The precise reasons why some models excelled in documentation reading but failed in follow-through are also not fully understood, suggesting further study is needed.
AI trust and integrity assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking and Adoption
Further benchmarking is expected to explore larger organizations and more diverse crisis scenarios, aiming to understand how AI models scale in management roles. Industry stakeholders are likely to focus on improving trust and integrity features within AI systems, possibly through enhanced auditing and rule-following mechanisms. Researchers and developers will also examine how training protocols can better align AI behavior with organizational values, especially under pressure. For organizations, the next step is to pilot these AI models in controlled environments, monitoring for trust breaches and performance consistency, and integrating feedback to refine AI management tools.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI managers never score zero in this benchmark?
Because even minimal effort—such as triaging crises or reading inboxes—counts as partial management, and the benchmark recognizes partial progress as valuable, resulting in a minimum score of 26 points for doing almost nothing.
What factors most influence AI performance in the benchmark?
Reading internal documentation thoroughly and resisting manipulation attempts are key factors that distinguish higher-performing models from others, especially in closing deals and managing crises effectively.
Why is trust breaches so heavily penalized in scoring?
The benchmark emphasizes that integrity and trustworthiness are fundamental to management; even a single breach can negate high competence, reflecting real-world risks of deploying AI in critical roles.
Can these AI models be trusted for real business management?
While they show promise, especially in handling crises and resisting manipulation, current models still have limitations. Trustworthiness and thoroughness need ongoing improvement before full deployment in high-stakes environments.
What is the significance of not scoring a perfect 100?
A perfect score would suggest unmeasured or unrealistic performance; the designed ceiling and floor ensure scores reflect partial progress and trustworthiness, making the benchmark more realistic and reliable.
Source: ThorstenMeyerAI.com