AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: What A Management Test Uncovers About AI’s Work Style on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A recent management test involving AI models managing a simulated company during a crisis reveals significant differences in their ability to execute decisions, maintain trust, and complete critical tasks. The experiment highlights that analysis alone is insufficient; effective action and discipline are crucial for AI to be truly operational.

Five AI management models were tested in a live simulation managing a small software company during its worst week. The experiment, conducted by Firmulate, aimed to assess how different AI systems handle real-world management tasks, including crisis response, trust preservation, and deal closure. The results reveal clear differences in their ability to translate analysis into effective action, highlighting challenges for AI adoption in operational roles.

The experiment involved five AI managers, each equipped with over 680 self-learned rules, overseeing a company with €105,000 monthly expenses and €2,300 recurring revenue. For more on AI management testing, see the original analysis. The simulation presented identical crises, customer crises, and temptations, with decisions made in real-time and auditable. The final league standings showed GPT-5.6-SOL in first place with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. For a broader discussion of AI management models, see this detailed overview. A baseline model scored only 26, emphasizing the importance of actual management performance over partial progress.

One key finding was that all models identified crises and refused manipulative requests, such as fake CEO messages, indicating strong security instincts. However, only two models successfully closed deals that required both analysis and execution. For example, a deal worth €4,583 in additional revenue was completed only when the model followed a trail inside the company’s files, demonstrating disciplined research and follow-through. Conversely, Opus 4.8, despite producing the most thorough analysis, failed to close the deal due to operational discipline lapses, such as attempting to write into a locked department instead of escalating.

At a glance
reportWhen: ongoing; results published July 2026
The developmentA live experiment tested five AI management models handling a simulated company’s worst week, revealing notable differences in execution, trust, and follow-through.

Implications for AI in Business Management

This experiment underscores that AI’s ability to analyze is not enough for effective management. Successful AI management requires discipline in execution, trustworthiness, and the capacity to complete critical tasks. The findings suggest that enterprises deploying AI for operational roles should rigorously test models against real-world pressures and decision-making scenarios before granting them autonomous authority. The differences in performance observed here highlight that effective AI management depends on more than just analytical depth; it demands disciplined follow-through and trust preservation, which are vital for operational success.

Amazon

AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing

Recent years have seen increasing interest in deploying AI models for management and operational roles within companies. Previous demonstrations often focused on analysis or hypothetical decision-making, with limited exposure to real-world pressures. The Firmulate experiment is notable for using a live, auditable simulation of a company facing crises, allowing a direct comparison of AI models‘ management styles. The league results from July 2026 build on earlier efforts to understand AI’s practical capabilities and limitations in operational settings, emphasizing the importance of disciplined action alongside analytical skill.

„Same diagnosis, same pitch — no signature.“

— Firmulate

Project Management with AI For Dummies

Project Management with AI For Dummies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Performance in Management Tasks

It remains unclear how these results will generalize to other operational contexts or larger organizations. The experiment focused on a specific crisis scenario with a limited set of models; how different models or real-world environments might influence outcomes is still uncertain. Additionally, the long-term implications of relying on AI for management decisions, especially regarding trust, discipline, and accountability, require further investigation.

Amazon

AI decision-making automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation

Organizations interested in deploying AI for management roles should consider conducting similar real-world simulations tailored to their specific operations. Further research is needed to develop standards for evaluating AI discipline, follow-through, and trustworthiness. Future experiments may expand the scope, include more diverse scenarios, and test larger, more complex organizational models to better understand AI’s operational readiness and limitations.

Amazon

AI operational discipline tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is follow-through important in AI management?

Follow-through ensures that AI models not only analyze problems but also complete critical actions, which is essential for operational success and trustworthiness in management roles.

Can AI models be trusted to handle sensitive business decisions?

The experiment shows that models with strong security instincts can refuse manipulative requests, but their ability to execute decisions reliably depends on disciplined operational behavior and thorough testing.

What are the main limitations of current AI management models?

Current models may excel at analysis but often struggle with operational discipline, closing deals, or executing follow-up actions consistently, which limits their effectiveness in real-world management.

How should companies test AI for operational management?

Companies should run scenario-based simulations that mimic real pressures, assess models‘ ability to follow through, and verify their trustworthiness before deploying AI in critical roles.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Siemens Advances Self-verifying Agentic AI Workflows For Semiconductor And PCB Design

Siemens unveils advanced AI workflows with self-verification for semiconductor and PCB design, enhancing reliability and efficiency in manufacturing processes.

Pentagon AI Goes Explicit: The Frontier Labs Move Inside the Classified Stack

The Pentagon has announced agreements with major AI firms to embed advanced AI into classified networks, signaling a shift toward AI-first military operations.

AI Changelog Digest For Open-source Maintainers

A new AI-powered weekly digest tool for open-source project maintainers is entering testing, aiming to simplify release summaries and dependency updates.

The Stanford AI Index 2026 Audit: Reading the Field’s Annual Report Card With a Critic’s Pen

A detailed audit of the Stanford AI Index 2026 reveals its strengths in benchmarking and transparency, while highlighting methodological limitations and interpretive risks.