
A garage’s worst week can start with a parts delay, a surge of repair requests or a message that appears to come from the boss. If AI is going to help manage bookings, customer questions or forecasts, a polished demo won’t show how it handles those pressures. Firmulate’s experiment offers a more demanding test: give models the same company and crises, then watch what they decide.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Same crises, different finishes
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Decisions were versioned and auditable, making it possible to examine what happened rather than judge a model by its conversation alone.
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee completing it.
The final standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”
The clue was buried in the company’s own files
The deal turned on a competitor weakness hidden two document references deep in the company’s files, not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a garage, the parallel is practical: useful facts may sit in service records, customer history or internal procedures, while the urgent request in front of an AI tells only part of the story.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusal is a meaningful safeguard, though the deal results show that caution and follow-through both matter.
Thoroughness did not ensure a strong finish
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are a record of this experiment, with that difference in conditions made explicit.
The company being watched is synthetic, but its business pressures are designed to be legible: 13 synthetic employees, burn of €105k a month against €2.3k MRR, and a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The live experiment can be watched at Firmulate. A separate quiz draws on 242 real, unedited management decisions and asks readers to guess which model made each one.
From watching to a company-specific trial
For an automotive business, the point is not whether an AI can sound confident about a service booking. It is whether it can handle a difficult week across customer commitments, changing conditions and company rules—and whether people can inspect its decisions. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s business, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Testing AI against your own business can reveal where it misses context, fails to complete a sound decision or needs a clearer escalation path. To discuss a Firmulate pilot using a read-only export, visit the pilot page and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
