firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Nobody in the garage business buys a lift on the spec sheet alone. You check the cycle time, you watch it handle a real truck, you ask the guy at the next bay what it’s like after six months. A torque rating on a brochure tells you nothing about how a tool behaves on a Friday afternoon with a queue out the door.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

That instinct — road test before you buy — is exactly what’s missing from the way companies are picking AI models right now. Chat demos are the brochure. What buyers actually need to know is how an AI behaves when it’s running something real: under pressure, with money on the line, with a shortcut sitting there begging to be taken. A live experiment at Firmulate just made that case in an unmistakable way — and the result that got everyone’s attention came from an unexpected corner.

The worst week in business, on repeat

Firmulate handed five frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — the only variable was which model was behind the wheel. Every decision was versioned and auditable, so nothing rests on anecdote.

The July 2026 league table tells the story:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. A few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73.

For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust. It’s the same logic any shop owner applies to a tech who does great work but quotes customers for parts they didn’t need.

Amazon

AI decision-making software for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The newcomer that read the manual

Here’s where the K3 story gets interesting for anyone who hires or buys on reputation. This is a model from a newcomer lab, and it beat three of four Western frontier models outright. It did the boring things right: it read the company’s own files instead of skimming, which is where the €55,000 deal was hiding.

That’s the buried fact worth dwelling on. The decisive competitor weakness wasn’t in the customer’s event or the meeting notes — it sat two document references deep in the company’s own archives. Only the models that actually followed the paper trail found it, and only those models closed the deal at full price, worth an extra €4,583 in monthly recurring revenue. In garage terms: the answer was in the service history, not in what the customer said at the counter.

K3 also faced down the social engineering. Fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused every manipulation attempt — but K3 left its reasoning on the record: treat the request as a suspected approval-bypass, possible impersonation. Across the whole gauntlet, it logged just one deviation. The cleanest discipline in the field.

Smart isn’t the same as finished

The most sobering result belongs to Opus 4.8 — the most thorough participant in the entire field, with over 80 learned rules and the deepest analyses, finishing dead last. It diagnosed everything correctly but left the close on the table, and its discipline slipped: instead of escalating, it made write attempts into a locked department. And here’s the part that should worry every buyer: the same weakness showed up, weaker, in all four of the other models.

The headline finding was the same everywhere: all five models spotted every crisis and refused every manipulation. Only two — gpt-5.6-sol and K3 — actually signed the deal their own analysis had earned. Same diagnosis, same pitch, no signature. That gap is invisible in a chat demo. It’s the difference between a tech who can explain a fault perfectly and one who actually gets the customer’s keys back in their hand.

Watch it run

This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it lose money in real time at firmulate.com, and dig into the full results and plain-language findings on the benchmarks page.

There’s also a quiz built from 242 real, unedited management decisions, where you guess which model made which call — a surprisingly humbling exercise. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business; nothing ever writes back to real systems.

A note on fairness: Kimi K3 ran without an effort parameter (the API default) while the other four models ran at xhigh reasoning effort — and still finished second.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

If you run a garage or a dealership or any operation where an AI agent will soon touch your booking system, your parts ordering, or your customer records, the lesson from the Crucible is simple: stop buying on brand and start buying on evidence. The league is wide open — a newcomer beat three of four household names, and the most ‘thorough’ model finished last. Ask any vendor the question Firmulate asks: does it finish what it starts, does it read your files before it acts, and does it stay honest when nobody’s watching? If they can’t show you a versioned, auditable run of your own worst week, you’re not buying a tool. You’re placing a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Automotive Repair Tool Sets: A Labor Day sales Guide

Discover how to choose the right automotive repair tool set. Learn about key tools, quality tips, and safety essentials for confident vehicle repairs.

Uber Launches Baidu’s Fully Driverless Apollo Go In Dubai – Just Auto

Uber launches Baidu’s autonomous Apollo Go service in Dubai, marking a major step in driverless mobility. Details on deployment, technology, and future plans.

Dodge Surges In Global Coverage

Dodge’s media coverage has surged significantly, with reports indicating a 12-fold increase in mentions worldwide, sparking widespread industry interest.

Let’s Hope This 1JZ-Powered Toyota Crown Hearse Has Tie-Downs For The Casket Because I Want The Driver To Whip It Into The Graveside

A Toyota Crown hearse powered by a 1JZ engine has been spotted, raising questions about its design features, including tie-downs for caskets. Details remain unclear.