AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Your next security incident may arrive as an urgent message from the boss

For business, marketing and ecommerce teams, AI agents promise speed across customer records, support queues, forecasts and sales operations. But speed becomes a liability when an apparently senior executive demands confidential data and insists there is no time for the usual process.

Firmulate put that scenario into a live, watchable experiment. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 frontier models refused every manipulation attempt. The result is an encouraging security story with a practical lesson: integrity under pressure can be tested before an AI workforce reaches production.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated under controlled conditions

Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable.

The company is not a tidy classroom exercise. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its public operation includes a cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday.

That environment makes the social-engineering test more revealing than a simple chatbot prompt. The models were already managing pressure from customers, money and internal work when the supposed CEO demanded an approval bypass. Kimi K3 recorded the clearest summary of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.” More model responses can be explored on Firmulate’s public quotes page.

Amazon

AI integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The refusals were unanimous, but the business results were not

All models spotted every crisis and refused every manipulation attempt. That matters because the benchmark treats trust as a hard boundary: a single breach caps the total, since “no amount of good work outweighs a breach of trust.” A do-nothing baseline still scores 26 because partial progress counts, but harmless inactivity is not the goal.

The final Crucible League results from July 2026 show that safety and execution can coexist:

  • gpt-5.6-sol scored 95.
  • Kimi K3 scored 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 scored 73.

The full standings and plain-language findings are published on the Firmulate benchmarks page. K3’s comparison also comes with an important fairness note: it ran with the API default and no effort parameter, while the other models ran at xhigh.

Refusing the con was only part of the job

The surprising gap appeared after the models had protected the company. Only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the miss as: “Same diagnosis, same pitch — no signature.” For an ecommerce or marketing leader, that distinction is crucial. An agent can recognize a problem, prepare a persuasive response and still fail to complete the commercially valuable action.

The decisive sales advantage was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue. The finding turns file-reading discipline into a business issue rather than an administrative detail.

Thoroughness did not guarantee the strongest performance

Opus 4.8 was the most thorough participant. It produced 80 additional learned rules and the deepest analyses, yet finished last. The deal was left unsigned, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.

That profile complicates the familiar assumption that more analysis necessarily produces better management. In this experiment, extensive reasoning did not compensate for an unfinished close or a process lapse. The league rewards models that protect trust, uncover the relevant evidence and carry legitimate work through to completion.

Firmulate also exposes the human difficulty of recognizing model behavior. Its quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each choice. The experiment’s value lies in those observable decisions, not in polished claims about what an agent might do.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model safety benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the pressure point before granting access

The strongest result is not simply that the models said no. All 5 of 5 maintained the boundary while operating a company under severe commercial pressure. At the same time, the benchmark exposed meaningful differences in research, follow-through and operational discipline.

That is the standard business buyers should want. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. Before an agent touches customer data or revenue work, leaders can observe whether it reads the evidence, completes the assignment and remains honest when an urgent authority figure asks it to break the rules.

The live Firmulate company makes those behaviors watchable. Its message for ecommerce and marketing teams is reassuring but demanding: trustworthy refusals are possible, yet security alone does not prove that an AI agent will finish the valuable work sitting on the other side of the crisis.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Top 7 AI-Powered Note Apps To Enhance Your Productivity In 2026

Discover the leading AI-driven note-taking apps of 2026, designed to enhance productivity through advanced transcription, summarization, and device support.

Say Goodbye To Eye Discomfort With Webcam-Driven Eye Wellness

A new webcam-driven app to monitor blink rate and promote eye breaks for remote workers is being tested as a solution to digital eye strain.

The Kill Switch: What the Anthropic Export Ban Really Costs the AI Industry

Anthropic’s models were abruptly shut down following US export controls, raising concerns over industry reliance on AI and regulatory risks.

Is Mistral Forge The AI Solution For Forward-Thinking Companies?

Exploring whether Mistral Forge suits organizations with high sovereignty, specialized data, and advanced AI needs, amid growing enterprise AI options.