firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A garage’s worst week can start with a parts delay, a surge of repair requests or a message that appears to come from the boss. If AI is going to help manage bookings, customer questions or forecasts, a polished demo won’t show how it handles those pressures. Firmulate’s experiment offers a more demanding test: give models the same company and crises, then watch what they decide.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Same crises, different finishes

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Decisions were versioned and auditable, making it possible to examine what happened rather than judge a model by its conversation alone.

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee completing it.

The final standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”

The clue was buried in the company’s own files

The deal turned on a competitor weakness hidden two document references deep in the company’s files, not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a garage, the parallel is practical: useful facts may sit in service records, customer history or internal procedures, while the urgent request in front of an AI tells only part of the story.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusal is a meaningful safeguard, though the deal results show that caution and follow-through both matter.

Thoroughness did not ensure a strong finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are a record of this experiment, with that difference in conditions made explicit.

The company being watched is synthetic, but its business pressures are designed to be legible: 13 synthetic employees, burn of €105k a month against €2.3k MRR, and a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The live experiment can be watched at Firmulate. A separate quiz draws on 242 real, unedited management decisions and asks readers to guess which model made each one.

From watching to a company-specific trial

For an automotive business, the point is not whether an AI can sound confident about a service booking. It is whether it can handle a difficult week across customer commitments, changing conditions and company rules—and whether people can inspect its decisions. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s business, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Testing AI against your own business can reveal where it misses context, fails to complete a sound decision or needs a clearer escalation path. To discuss a Firmulate pilot using a read-only export, visit the pilot page and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Teach Your Kids How V8s And Stick Shifts Work With This $40 Model Car

A $40 model car is gaining attention for helping children learn how V8 engines and manual transmissions work, sparking increased interest in educational toys.

Diesel Prices Are About To Break Records. Here’s Why They Keep Rising

Diesel prices are approaching record levels due to supply constraints and global market pressures. Here’s what is driving the increase and what it means for consumers.

Inside The $2 Billion Battery Factory Making Cells For Tesla And Toyota

A new $2 billion battery factory is under construction to supply Tesla and Toyota, marking a major step in EV battery manufacturing expansion.

Honda Surges In Global Coverage

Search interest in Honda has surged, with media mentions rising dramatically, signaling increased global attention amid ongoing industry developments.