firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A garage’s worst week can start with a parts delay, a surge of repair requests or a message that appears to come from the boss. If AI is going to help manage bookings, customer questions or forecasts, a polished demo won’t show how it handles those pressures. Firmulate’s experiment offers a more demanding test: give models the same company and crises, then watch what they decide.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garage and car supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same crises, different finishes

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Decisions were versioned and auditable, making it possible to examine what happened rather than judge a model by its conversation alone.

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was strikingly simple: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee completing it.

The final standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”

The clue was buried in the company’s own files

The deal turned on a competitor weakness hidden two document references deep in the company’s files, not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a garage, the parallel is practical: useful facts may sit in service records, customer history or internal procedures, while the urgent request in front of an AI tells only part of the story.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusal is a meaningful safeguard, though the deal results show that caution and follow-through both matter.

Thoroughness did not ensure a strong finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The standings are a record of this experiment, with that difference in conditions made explicit.

The company being watched is synthetic, but its business pressures are designed to be legible: 13 synthetic employees, burn of €105k a month against €2.3k MRR, and a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The live experiment can be watched at Firmulate. A separate quiz draws on 242 real, unedited management decisions and asks readers to guess which model made each one.

From watching to a company-specific trial

For an automotive business, the point is not whether an AI can sound confident about a service booking. It is whether it can handle a difficult week across customer commitments, changing conditions and company rules—and whether people can inspect its decisions. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s business, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Testing AI against your own business can reveal where it misses context, fails to complete a sound decision or needs a clearer escalation path. To discuss a Firmulate pilot using a read-only export, visit the pilot page and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

BMW That Plunged Off Mumbai’s Coastal Road Bridge Had Earlier Hit Couple, Video Viral – The Times Of India

A BMW vehicle plunged off Mumbai’s Coastal Road bridge after hitting a couple, with a video of the incident going viral. Details are still emerging.

Nissan Surges In Global Coverage

Media coverage of Nissan has increased significantly, with reports indicating a 9.8-fold spike in mentions. The reasons behind this surge remain unconfirmed.

2027 Taipei AMPA、E-Mobility Taiwan全球徵展啟動 以360° MOBILITY平台整合售後零件與智慧移動 – 經貿透視雙周刊

Taipei’s 2027 AMPA and E-Mobility Taiwan events initiate a 360° mobility platform integrating parts and smart mobility solutions, signaling a shift in the industry.

Jetour Wmotors Surges In Global Coverage

Jetour Wmotors experiences a significant surge in international coverage, reflecting growing global interest in the brand amid expanding markets.