
Every shop owner knows the type: a technician who nails the diagnosis every single time — misfire on cylinder three, parasitic drain, the intermittent fault nobody else could find — but never picks up the pen to close the repair order with the customer. The diagnosis was perfect. The revenue didn’t happen. In the garage, we call that leaving money on the lift.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A public experiment called Firmulate just proved that the exact same failure mode exists in frontier AI models — and it built a scoring system honest enough to catch it. The headline finding from its final July 2026 “Crucible League”: four elite AI models each ran the same small software company through its worst week. All four spotted every crisis. All four refused every manipulation attempt. Only two actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The Do-Nothing Floor at 26
Here’s the detail that tells you this benchmark is serious: a baseline run that does nothing still scores 26 points. Not zero. That’s deliberate. In the Firmulate methodology, partial progress counts — if a model correctly triages a crisis, documents it, and gets halfway to a resolution before stalling out, that work has value, just like a correct partial diagnosis in the bay has value. A zero would imply that nothing short of a perfect close matters, and no honest manager runs a business that way.
But the floor comes with a ceiling. The methodology caps the total score if a model commits a single breach of trust — the principle being, in the experiment’s own words, “no amount of good work outweighs a breach of trust.” Think of it like a service advisor who upsells beautifully for eleven months and gets caught billing for work never performed once. That’s not a deduction. That’s a firing.
AI diagnostic and troubleshooting software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the League Actually Showed
The final July 2026 standings: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. The spread matters less than the pattern behind it.
The €55,000 deal — the biggest revenue event of the simulated week — turned on what Firmulate calls the buried fact. The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Models that didn’t, didn’t. It’s the AI equivalent of the tech who skips the service history in the glovebox and misses the pattern failure.
Then there’s the Opus 4.8 profile — the most instructive result of the run. It was the most thorough participant in the field: over 80 learned rules added and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped — including write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Brilliance without follow-through loses to competence with it.
Pressure and Impersonation
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” If you’ve ever had a “vendor” email your parts counter asking to change payment details, you know why this test exists.
One fairness note Firmulate discloses openly: K3 ran without an effort parameter while the others ran at xhigh — and still took second at 93.
It’s Live, and It’s Watchable
Firmulate isn’t a paper — it’s a running company you can watch at firmulate.com. Thirteen synthetic employees, real money mechanics: €105k monthly burn against €2.3k in monthly recurring revenue, a public cash countdown, 680+ self-learned playbook rules, every workday versioned and auditable. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for anyone whose livelihood depends on operations — a garage, a fleet shop, a dealership service lane — is simple. When you evaluate an AI tool, don’t ask how well it chats. Ask what happens at the closing counter. Does it finish what it starts? Does it read your files before it answers? Does it stay honest when pressure escalates? Firmulate’s scoring system bakes those questions in: partial progress earns points, unfinished deals cost them, and one breach of trust ends the conversation entirely. That’s not just a benchmark. That’s how a good shop owner has always graded people.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
