firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

You Don’t Buy a Router Because the Box Looks Nice

Any woodworker knows the ritual. The new plunge router looks gorgeous in the store, the demo cut is perfect under ideal conditions — and then you run it through eight hours of maple and the bearings whine, the depth lock slips, and the collet cambers under load. The spec sheet told you about peak horsepower. It never told you how the tool behaves on hour seven, in your shop, with your wood.

We now have the same problem with AI. Coding leaderboards and chat arenas are the store demo: one clean cut, perfect conditions, no consequences. They measure whether a model can answer brilliantly. They say nothing about whether it can manage — triage a bad week, finish what it starts, stay honest when nobody’s looking.

That’s the gap a public experiment called Firmulate set out to measure. Its pitch is blunt: it measures management quality, not chat quality. And the first full results, finalized in July 2026, say the gap is real — and bigger than anyone assumed.

Amazon

AI management and performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week, Four Different Foremen

The setup is simple, like a controlled cut test. Four frontier AI models — GPT-5.6-Sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the exact same small software company and put through its worst possible week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so you can go back and see exactly who did what.

Think of it as handing four foremen the same jobsite, the same materials list, and the same client who keeps changing their mind — then scoring how the shop runs on Friday.

The Final League Table

  • gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • Sonnet 5 — 88. A few more process slips.
  • Fable 5 — 77. Mid-field.
  • Opus 4.8 — 73. Last, despite being the most thorough participant.

For calibration: a do-nothing baseline scores 26. Partial progress counts — but the scoring has one hard rule any contractor would respect: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.

The Headline Finding: Same Diagnosis, Same Pitch — No Signature

Here’s where it gets interesting for anyone who’s ever hired help. Every model spotted every crisis. Every model refused every manipulation attempt. But only two of the five actually signed the €55,000 deal that their own analysis had earned.

Picture a subcontractor who diagnoses the rot perfectly, writes up a flawless quote, presents it to the client — and then walks off site without getting a signature. Same diagnosis, same pitch, no signature. The work was there. The follow-through wasn’t.

That gap is invisible in a chat demo. A model can be brilliant in conversation and still leave the close on the table.

The Buried Fact: Read the Files Before You Quote

The deal-breaker wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue.

Every tradesperson knows this instinct: check the joists before you quote the reno. The models that skipped the crawl space lost the job.

Can It Be Conned? Apparently Not

The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the instinct you want in anyone holding keys to your systems.

The Thoroughness Trap

The most instructive profile is Opus 4.8. It was the most thorough participant in the field — 80-plus learned rules, the deepest analyses — and it still finished last. It left the close on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Over-preparing and under-finishing is a failure mode any perfectionist in the workshop will recognize.

One fairness note: K3 ran at the API’s default effort setting while the others ran at xhigh — and still nearly won.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Not a Slide Deck — A Running Shop You Can Watch

Firmulate isn’t a one-off paper. The company is live software with 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in MRR, with a public cash countdown. It runs every business day, holds 680+ self-learned playbook rules, and rebuilds itself twice a day. You can watch it lose money in real time.

There’s also a genuinely fun artifact: 242 real, unedited management decisions from the experiment power a “guess which model made this call” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI project management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway: Demand the Torque Test

Nobody in this trade buys a tool off the box art. You ask about runout, bearing quality, what happens under sustained load. If AI agents are going to touch your CRM, your support queue, or your forecast, the questions are the same shape: Does it finish what it starts? Does it read the files before quoting? Does it stay honest under pressure? What does a unit of useful work actually cost?

The Firmulate results suggest those questions have real, measurable answers — and that they don’t track with chat rankings. The most thorough model came last. A newcomer on default settings nearly won. The difference wasn’t eloquence; it was whether the model read the joists and got the signature.

Wargame your AI workforce before you hire it. The full league table and plain-language findings are at firmulate.com, with detailed results on the benchmarks page. The company is running right now, twice-daily rebuild and all — go watch it work, and see which foreman you’d actually hire.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Drill Bit Types and Uses

Wondering which drill bit type suits your project? Discover the essential tools and tips to ensure perfect results every time.

Creating a Sharpening Station for Blades and Chisels

Just setting up a sharpening station for blades and chisels can transform your woodworking—discover how to optimize your space and improve your results.

Galería De Ferder Custom Shop Mexico City / Arqmov Workshop – 12 – ArchDaily

The Ferder Custom Shop in Mexico City, designed by Arqmov Workshop, is showcased in ArchDaily’s latest report, highlighting innovative architecture and craftsmanship.

I Moved A Clock To My Bathroom, And It Changed My Morning Routine Forever

A homeowner’s simple change—placing a clock in the bathroom—has significantly transformed their morning habits, sparking interest in routine optimization.