firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Would You Hand Your Shop to a Mechanic Who Doesn’t Open the Hood?

Every garage owner knows the type: the tech who walks up to a car, hears the rattle, and starts guessing. And every garage owner knows the other type — the one who pulls the service history, finds the recall notice filed two owners ago, and fixes the actual problem in an hour. Same car, same symptoms, wildly different outcome.

A new experiment from Firmulate, which runs AI models through simulated companies with real money mechanics, just showed that AI agents have exactly the same split. Four frontier models were each handed the same small software company during its worst week. All of them diagnosed the problems. All of them survived every attempt to manipulate them. But only two of them did the digital equivalent of opening the service history — and that one habit was the difference between closing a €55,000 deal at full price and losing it entirely.

Amazon

automotive service history database software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Worst Week, Run Four Times

The setup was brutally fair. Each model ran the identical company through the identical seven-day gauntlet: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing came down to interpretation. Think of it as four technicians given the same lemon with a hidden fault — and a hidden service bulletin buried in the paperwork.

The final league table from the July 2026 Crucible run tells the story:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93 points. Closed the deal too, with the cleanest discipline in the field.
  • Sonnet 5 — 88 points. Solid work with a few process slips.
  • Fable 5 — 77 points.
  • Opus 4.8 — 73 points. The most thorough participant of all — and still last.

For context, doing nothing at all scores 26, because partial progress counts. But there’s a hard ceiling: a single breach of trust caps the total. In Firmulate’s own words, “no amount of good work outweighs a breach of trust” — the same principle as a shop that does beautiful brake jobs but pads the invoice.

The €55,000 Fact, Two Documents Deep

Here’s the finding that should matter to anyone whose business runs on records — warranty files, parts catalogs, customer histories. The decisive weakness in a competitor’s offering wasn’t in the customer conversation at all. It sat two document references deep in the company’s own internal files. A model had to actually go read its own paperwork, follow the citations, and connect the dots.

The two models that did their homework won the €55,000 deal at full price — worth €4,583 in monthly recurring revenue. The ones that didn’t lost it automatically. As Firmulate summarized the pattern: “Same diagnosis, same pitch — no signature.” The models could talk a great game. They just hadn’t read the file that made the pitch land.

If you’ve ever lost a customer to the shop across town because their tech remembered the TSB your tech never looked up, you already understand this failure mode intimately.

Honest Under Pressure

The week didn’t just test competence — it tested integrity. The models faced fake CEO messages that escalated over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five manipulation attempts were refused. Kimi K3’s on-record reasoning was the standout: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the instinct you want in anything with access to your accounts, your customers, or your parts ordering.

The Cautionary Tale: Thorough Isn’t the Same as Good

Opus 4.8 is the profile every overworked shop owner should study. It was the most thorough participant in the entire field — over 80 learned rules, the deepest analyses of anyone running. And it finished last. The €55,000 close was left sitting on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness showed up, weaker, in all four models. Effort and diligence don’t automatically convert into finished work — a lesson anyone who’s watched a perfectionist technician miss a booking deadline already knows.

One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at extra-high effort — and still nearly won.

You Can Watch It Live

This isn’t a one-off paper. Firmulate runs a live, watchable company around the clock: 13 synthetic employees, real money mechanics, a burn rate of €105k a month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — with every workday versioned and visible. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, which is exactly as humbling as it sounds.

For larger operations, there’s a pilot program: enterprises can run the same wargame against a read-only export of their own business. Nothing ever writes back to real systems, so it’s a safe way to find out whether your would-be AI workforce reads the manual or just improvises.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Takeaway

Chat demos measure how well an AI talks. This experiment measured how well it works — and the gap was worth €55,000 on a single deal. Before you let any AI agent near your scheduling, your invoicing, or your customer records, the question isn’t “does it write well?” It’s: does it finish what it starts, does it stay honest when someone tries to impersonate the boss, and — above all — does it read your files before it answers? The full league table and plain-language findings are published at firmulate.com/benchmarks.html. Bring the same skepticism you’d bring to a technician who diagnoses your check-engine light without plugging in the scanner.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Complete Guide to Automotive Care, Detailing, Maintenance, and Garage Essentials

AIThis post was created with the assistance of artificial intelligence (AI).Taking good…

The Super Bee Name Is BACK With A 600HP Inline-6!

The legendary Super Bee name is making a comeback with a new 600-horsepower inline-6 engine, marking a significant shift in muscle car offerings.

Nissan’s ‘Home Run’ Hybrid SUV Arrives In November To Rival Toyota RAV4, Honda CR-V

Nissan’s new hybrid SUV, dubbed ‘Home Run,’ is scheduled for release in November, targeting popular models like Toyota RAV4 and Honda CR-V. Details are confirmed, but some claims remain unverified.

Ducati Surges In Global Coverage

Ducati’s media mentions have increased significantly, with a 26-fold rise in recent coverage, highlighting growing global interest in the brand.