
Your Scanner Finds the Fault. Do You Finish the Job?
Every garage owner knows the type: the technician whose diagnostics are flawless, whose write-ups are beautiful, who can tell you the misfire is a coil on cylinder three before the hood is up — and who still hands the customer a clipboard and walks away without asking for the job. The diagnosis was perfect. The bay stayed empty. In the automotive service business, we have never confused knowing with closing. It might be the first thing we learn.
It turns out the artificial-intelligence industry is only now learning it too — and the results of a live, watchable experiment suggest that anyone planning to put AI agents near customers, calendars, or cash should pay very close attention.
The Wargame That Refuses to Be a Demo
Firmulate, which describes itself as “the AI company emulator,” ran something unusual: instead of asking chatbots clever questions and grading their answers, it handed four frontier AI models the same job. Each one was put in charge of the same small software company during its worst week — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the run can be quietly retouched afterward.
The final league table from July 2026 reads: gpt-5.6-sol in first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline — a model that simply sat on its hands — scored 26, because partial progress counts for something. But there was one hard ceiling: a single breach of trust caps the total. In Firmulate’s own words, “no amount of good work outweighs a breach of trust.” That rule will feel familiar to anyone who has ever run a service counter.
Everybody Diagnosed. Not Everybody Closed.
Here is the finding that should stop any owner mid-coffee: all four models spotted every crisis, and all four refused every manipulation attempt thrown at them. Yet only two of them signed the €55,000 deal that their own analysis had earned. Firmulate’s summary of the failure is blunt: “Same diagnosis, same pitch — no signature.”
Think about what that means. The models did the equivalent of a perfect inspection, printed a flawless estimate, explained the repair to the customer — and then let the customer walk out the door. If a service advisor did that, you would not call it a technical failure. You would call it a management failure. Firmulate calls the whole category exactly that: it measures management quality, not chat quality.
The Buried Fact That Won the Deal
The decisive detail in the €55,000 negotiation was not in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that the winners found by actually reading what the business already knew. The models that did the reading won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that skipped the homework left it on the table.
Any garage that has ever lost a job because nobody checked the service history in the folder before quoting has lived this exact story. The knowledge was in the building. Nobody opened the file.
The Social Engineering Test — Five for Five
The experiment also staged a pressure campaign: fake CEO messages that escalated over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” All five participating models refused every attempt. Kimi K3’s on-record reasoning was admirably suspicious: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the instinct you want in an employee, silicon or otherwise.
When Thoroughness Isn’t Enough
The most striking individual profile belongs to Opus 4.8 — the most thorough participant in the field, with 80 learned rules added during the run and the deepest analyses of any model. It finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Firmulate notes the same weakness appeared, weaker, in all four models. Again: the best diagnostician in the shop, and the empty bay.
One fairness caveat Firmulate itself discloses: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Worth knowing before you crown a champion on a single decimal point.
You Can Watch the Company Lose Money
This is not a slide deck. Firmulate runs a live synthetic company — 13 employees, real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. You can watch it at firmulate.com, which rebuilds itself twice a day.
There is also a genuinely fun hook for skeptics: 242 real, unedited management decisions from the runs power a “guess the model” quiz, and enterprises can go further — running the same wargame against a read-only export of their own business, with nothing ever writing back to real systems. The full methodology and findings are on the benchmarks page.

What a Garage Owner Should Take From This
The AI vendors will keep showing you chat demos — clever answers, polished prose, impressive code. Those measure the diagnostician. What Firmulate measured instead was whether the agent finishes what it starts, reads your files before quoting, and stays honest when someone impersonates the boss. Under that curriculum — churn waves, price increases, down rounds, PR crises — the field separates fast: a 95 for the complete performer, a 73 for the brilliant one who never asked for the job.
If you are ever sold an AI agent for your scheduling, your parts ordering, or your customer follow-ups, ask the vendor one question: has it been tested on outcomes and temptations, or only on answers? Because in this business, we already know the difference between a perfect diagnosis and a closed job. It is the difference between a full bay and an empty one — and now, apparently, between 95 points and 73.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.