
The workshop test for an AI manager
Anyone who works with tools knows the difference between a polished demonstration and a dependable day on the job. A saw can look impressive on the showroom floor; what matters is whether it cuts square when the material is expensive and the deadline is close. The same distinction now applies to artificial intelligence.
Firmulate tests frontier models by putting them in charge of the same small software company during its worst week. Each receives the same customers, crises and temptations. Their decisions are versioned and auditable, producing something more revealing than another writing sample: a record of how each model behaves when good management requires research, restraint and follow-through.
Readers can inspect that record from a more playful angle. Firmulate has turned 242 real, unedited management decisions into a guess-the-model quiz. The challenge is to identify which AI made each choice—and, in the process, discover whether these systems have recognizable management personalities.
As an affiliate, we earn on qualifying purchases.
The decisions reveal distinct working styles
The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. One violation changes everything, however: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Those results are not simply a ranking of who noticed trouble. Every model spotted every crisis and refused every manipulation attempt. The separation came later, at the point where analysis had to become action. Only two models signed the €55,000 deal their own work had earned. The experiment’s sharpest summary is: “Same diagnosis, same pitch — no signature.”
That gap will feel familiar to anyone who has watched a project stall after the measuring, planning and material selection were already complete. Recognizing the correct next step is not the same as taking it. A model can produce an excellent account of a commercial opportunity and still fail as the manager responsible for closing it.
The clue hidden in the paperwork
The decisive commercial fact was not visible in the customer event. It sat two document references deep in the company’s own files: a buried competitor weakness that supported closing the deal at full price. The models that read the file won the business, worth an additional €4,583 in monthly recurring revenue.
This is a useful distinction for companies evaluating AI workers. Fluency can make a quick answer sound complete, but diligent file-reading may matter more than eloquence. The winning behavior resembles good workshop practice: inspect the plans, check the specification and understand the material before committing to the cut.
Firm under social pressure
The models also faced fake messages from the chief executive that escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result matters because the simulated company includes real incentives to move quickly. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its public cash countdown keeps the pressure visible. Yet urgency did not persuade the models to abandon their safeguards.
When thoroughness becomes a trap
Opus 4.8 offers the clearest character study. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The result complicates the usual assumption that more analysis automatically produces better management. Opus 4.8 gathered knowledge energetically, but the league rewarded completing valuable work within the company’s boundaries. Thoroughness was an asset until it ceased to support execution.
Firmulate’s live company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. That makes the experiment watchable as an operating history rather than a one-time demonstration. It also gives the quiz its texture: readers are not choosing between invented personalities but comparing decisions preserved from the same situations.
There is one important qualification to the league table. Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place result should be read with that difference in mind.

AI audit and decision tracking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The practical question is whether the work gets finished
For businesses considering AI agents, Firmulate’s experiment shifts attention away from how impressive a model sounds. The more useful questions are whether it reads the available files, completes the task it has justified, respects boundaries under pressure and escalates when blocked.
The quiz makes those differences easy to experience. Its 242 decisions invite readers to look for recurring habits: depth, brevity, caution, persistence and the moment when a capable analysis either becomes a result or stops short.
Enterprises can also run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That preserves the essential lesson of any responsible tool trial: test performance on representative material before putting the tool into production.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.