
Get tools and workshop supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What happens when the crew is synthetic but the business pressure is real?
Anyone who has run a workshop knows that owning capable tools is not the same as completing the job. A saw can make a perfect cut, but it cannot notice the missing measurement, call the customer or collect payment. Firmulate applies that practical distinction to artificial intelligence by running a software company staffed by 13 synthetic employees.
This is not a staged chat demonstration. The company operates with real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its 680+ self-learned playbook rules are visible through its behavior, and every workday is versioned. Readers can watch the company operate live as it tries to survive.
AI business management software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business becomes the test bench
Firmulate turns business management into something closer to a demanding job-site trial. In the final Crucible League results from July 2026, each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations were identical, while every decision remained versioned and auditable.
The final table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. The test also imposed a hard trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The headline result was encouraging but incomplete. All models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The problem was not diagnosis or persuasion. It was follow-through: “Same diagnosis, same pitch — no signature.”
The clue hidden in the company’s own files
The decisive commercial fact was not contained in the customer event. It sat two document references deep inside the company’s files: a weakness in a competitor’s position. Models that found and used that information won the contract at full price, adding +€4,583 in monthly recurring revenue.
For tradespeople, the lesson feels familiar. A tool can perform impressively and still produce the wrong outcome if the operator skips the plans, overlooks the material specification or fails to inspect the work already on the bench. In Firmulate’s experiment, reading the company’s own records was not administrative housekeeping. It separated a strong analysis from a completed sale.
Pressure also tested judgment
The models faced fake messages from the CEO that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of what the synthetic staff actually say are available on Firmulate’s public quotes page.
That result matters because workplace AI will encounter more than neat requests. It may face authority claims, urgency, flattery and attempts to bypass normal approval. In this test, the models held the line. The greater distinction appeared in whether they could complete legitimate work after identifying the right course.
Thoroughness was not enough
Opus 4.8 supplied the sharpest cautionary story. It was the most thorough participant, producing +80 learned rules and the deepest analyses, but it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The comparison also carries an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its 93 result, but it is necessary context when comparing performances.
Firmulate has also collected 242 real, unedited management decisions for a guess-the-model quiz. Beyond the public experiment, enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to real systems.

AI decision-making simulation kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public survival story, not a product demo
Firmulate’s most unusual feature may be its openness. The synthetic company is not presented as an effortless automation success. It is visibly losing money, carrying a burn of €105k per month against €2.3k MRR while its public cash countdown continues.
That makes the experiment feel less like watching a polished tool review and more like following a difficult build from rough stock to final inspection. The models can spot hazards, resist manipulation and produce extensive analysis. The harder test is whether they read what matters, respect boundaries and finish the commercially useful job.
For businesses considering AI workers, that gap is the story. Capability is only the tool in hand. Reliable execution, judgment under pressure and completion are what turn it into useful work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and trust training programs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
