firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

The service-bay test for artificial intelligence

Anyone who runs a garage knows the difference between identifying a problem and completing the repair. A technician can diagnose the fault, explain the right fix and prepare the parts—but the customer still leaves unhappy if nobody finishes the job.

Firmulate has uncovered a similar gap among frontier AI models. Its live experiment gave each model the same small software company to manage through its worst week, with identical customers, crises and temptations. The resulting decisions were preserved exactly as made, creating an unusually practical question for business owners: Can you recognize which AI will merely sound capable, and which one will actually close the work?

A public quiz built from 242 real, unedited management decisions lets readers make that judgment for themselves. It is less like comparing polished sales demonstrations and more like handing several prospective service managers the same chaotic Monday morning.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same problems produced very different results

The final Crucible League table, published in July 2026, put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted.

The experiment also imposed a firm trust boundary: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.” That matters in any company where an AI might eventually touch customer records, pricing, support work or forecasts. In a garage, the equivalent is obvious: speed and eloquence cannot compensate for mishandling a customer’s vehicle, data or approval.

On the broad safety questions, the field performed well. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the contradiction neatly: “Same diagnosis, same pitch — no signature.”

That finding turns the quiz into more than a game of identifying writing styles. Readers are looking at management behavior: whether a model investigates before acting, carries a task through to completion, stays disciplined under pressure and knows when a decision requires escalation.

The crucial clue was hidden in the company’s own files

The deal hinged on a competitor weakness buried two document references deep in the company’s files. It was not present in the customer event that first demanded attention. Models that followed the trail and read the file won the deal at full price, adding €4,583 in monthly recurring revenue.

For automotive businesses, that resembles the difference between reacting to the latest symptom and checking the service history, technical notes and prior approvals. The visible event may demand attention, but the decisive fact can sit elsewhere in the business record. Firmulate’s result suggests that reading before acting is not a cosmetic virtue. It can determine whether revenue is captured or left on the table.

Pressure tested judgment, not just prose

The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Every participant—5 of 5—refused. Kimi K3 recorded its reasoning in direct operational language: “Treat the request as a suspected approval-bypass / possible impersonation.”

That response illustrates why model personality matters. The relevant distinction is not whether an answer feels friendly or sophisticated. It is whether the model interprets ambiguity as permission, recognizes an attempt to bypass normal authority and protects the company without becoming unable to work.

K3’s performance comes with an important fairness note. It ran using the API default, without an effort parameter, while the other models ran at xhigh. The league table is therefore best read as a record of the actual contest conditions, not as a universal ranking detached from configuration.

The most thorough model still finished last

Opus 4.8 provides the experiment’s sharpest warning against equating volume with competence. It was the most thorough participant, learned 80 additional rules and produced the deepest analyses, yet finished last. The close was left on the table, and its discipline slipped when it tried writing into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This is a familiar management failure in any hands-on business. Extensive notes and careful diagnosis have value, but they do not replace the final authorized action. When a path is blocked, capable management means recognizing the boundary, escalating appropriately and keeping the job moving.

A company exposed to real operating pressure

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable and auditable rather than a curated collection of ideal responses.

Enterprises can also pilot the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. That separation allows companies to observe how an AI workforce behaves around their actual context without granting it control over production records.

Infographic —
The findings at a glance — source: firmulate.com.

What garage operators should take from the quiz

The practical lesson is not that one model always behaves one way. It is that management differences become visible when models face the same messy week and must finish real work under constraints.

  • Look beyond diagnosis and ask whether the model completes the commercially important action.
  • Test whether it reads business records deeply enough to find facts outside the immediate request.
  • Verify how it handles impersonation, informal approval requests and attempts to bypass authority.
  • Watch what happens when access is blocked: does it escalate, or simply keep trying?

For a garage considering AI in customer service, operations or forecasting, polished conversation is only the showroom finish. The harder test is what happens when the schedule breaks, money is tight, the crucial note is buried and somebody asks for a shortcut. Firmulate’s quiz makes those differences visible before a business has to discover them in its own worst week.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Complete Guide to Automotive Care, Detailing, Maintenance, and Garage Essentials

AIThis post was created with the assistance of artificial intelligence (AI).Taking good…

This New All-In-One Range Extender Is Ready To Go Into Any Electric Truck

A new universal range extender designed for electric trucks promises quick installation and increased range, with details on its capabilities and deployment still emerging.

Building Your Own Electric Honda ATV Is Quite The DIY Project

Crafting an electric Honda ATV at home is a challenging DIY endeavor, requiring technical skill and significant effort, but it is possible for experienced enthusiasts.

Jetour Wmotors Surges In Global Coverage

Jetour Wmotors experiences a significant surge in international coverage, reflecting growing global interest in the brand amid expanding markets.