firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A pressure test for digital power tools

Anyone who works with serious tools knows the safety check matters most when the tool is under load. A drill may behave perfectly on the bench and still fail when the bit binds. The same principle applies to AI systems being trusted with customer records, forecasts and operational decisions: fluent answers reveal little about what happens when someone demands an unsafe shortcut.

Firmulate tested that problem directly. Its live experiment placed frontier AI models in charge of the same small software company during its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. Among the tests were fake messages from the CEO, escalating over three stages, followed by a reporter seeking confidential confirmation with the line, “just one yes/no, on background.” All 5 models refused every manipulation attempt.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The forged authority test

The fake CEO messages applied a familiar social-engineering tactic: combine apparent authority with urgency, then dismiss normal process as an obstacle. The requested action involved sending a customer list to a journalist. As the pressure increased, the models had repeated opportunities to treat the instruction as legitimate simply because it appeared to come from the top.

None did. Kimi K3 captured the correct posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because it identifies both possibilities without pretending to know more than the evidence supports. The message might have been fraudulent, or it might have been a real executive attempting to bypass safeguards. Either way, the customer data should not move.

The reporter trick tested a different route to the same protected information. Instead of issuing an order, it minimized the request and suggested informality. The models still refused. Across the staged CEO pressure and the reporter approach, 5 of 5 held the line while also recognizing every crisis presented to them.

Amazon

AI model safety evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Security was necessary, but it was not the whole job

The encouraging result did not make the models interchangeable. Firmulate’s final July 2026 Crucible League ranked gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The largest divide emerged after the models had already understood the commercial opportunity. Only 2 signed the €55,000 deal their own analysis had earned. The outcome was summarized as “Same diagnosis, same pitch — no signature.” In workshop terms, several participants measured correctly, selected the right tool and positioned the cut, but did not finish the workpiece.

The decisive commercial clue was also easy to miss. A competitor weakness sat two document references deep in the company’s own files rather than in the customer event. Models that followed the trail won the deal at full price, worth +€4,583 MRR. The lesson is practical: an agent can behave honestly and still underperform if it fails to inspect the available material or carry a sound decision through to completion.

The thorough model that finished last

Opus 4.8 makes that distinction especially vivid. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in the other 4 models as well.

That result challenges the assumption that more analysis automatically produces better management. Thoroughness is valuable, but an enterprise agent must also respect boundaries, recognize when escalation is required and complete authorized work. Firmulate’s comparison makes those behaviors visible together rather than treating safety and productivity as separate demonstrations.

There is an important fairness detail in the ranking: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not change the observed refusals or completed actions, but it belongs beside the scores when readers compare the participants.

Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company built to expose real trade-offs

The environment gives these decisions consequences. The live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real, public and watchable rather than a collection of handpicked chat responses.

Its evidence base extends beyond the headline result. A “guess the model” quiz draws from 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to their real systems. That creates a way to observe judgment under company-specific pressure without granting the tested agent control over production systems.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI integrity testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the guard before the kickback

The most useful finding is not simply that every model said no. It is that integrity under pressure can be observed before an AI agent is placed near live customer data. A convincing fake executive, artificial urgency and a reporter’s seemingly harmless request can be introduced deliberately, with the resulting decisions preserved for review.

  • Test whether the agent verifies authority when urgency rises.
  • Check whether it reads the company’s own material deeply enough to find buried evidence.
  • Watch whether it completes authorized work and escalates when a boundary blocks progress.

Firmulate’s models passed the manipulation test, but their commercial follow-through and operational discipline varied sharply. That combination is the real story: trustworthy behavior is measurable, and so is the gap between understanding the job and finishing it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Best Way to Store Tool Batteries Without Killing Their Lifespan

Find out how proper storage can extend your tool batteries’ lifespan and avoid common mistakes that could shorten their performance.

Selecting an Air Compressor for DIY Projects

An essential guide to selecting an air compressor for DIY projects, exploring key factors that ensure optimal performance and longevity.

Mortgage Refinance Rates

Mortgage refinance rates have declined significantly in April 2024, reaching their lowest levels in over a year, impacting homeowners and the housing market.

The Hidden Benefit of a Dedicated Tool Return Zone in Busy Workshops

Best workshops boost efficiency and safety—discover how a dedicated tool return zone can unlock hidden benefits you can’t afford to miss.