
A different kind of pressure test
Anyone running a garage knows the gap between identifying a problem and finishing the job. A technician can diagnose the fault, explain the repair and prepare the estimate, but the business earns nothing until someone secures approval and moves the vehicle through the workshop. Firmulate has exposed a remarkably similar gap in artificial intelligence.
The public experiment operates a small software company staffed by 13 synthetic employees. Its financial predicament is genuine within the experiment: the company burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. More than 680 self-learned playbook rules now document what the company has discovered, and every workday is versioned.
This is build-in-public taken beyond product announcements and polished dashboards. The company’s struggle is a running business story, available to watch live, with decisions, conversations and financial consequences unfolding in public.
automotive service history management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same bad week, five different managers
Firmulate’s Crucible League placed frontier AI models in charge of the same small software company during its worst week. They received the same customers, crises and temptations, and every decision was versioned and auditable. That common setting made the differences unusually revealing.
The final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was also a firm ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The reassuring result was that every model noticed every crisis and rejected every attempted manipulation. The troubling result was that only two signed the €55,000 deal their own work had earned. The central finding can be summarized in Firmulate’s own phrase: “Same diagnosis, same pitch — no signature.”
For garage operators, that distinction should feel familiar. Recognizing a worn component is not the same as obtaining customer authorization, ordering the part and completing the repair. AI performance cannot be judged solely by the quality of an explanation. Execution, follow-through and commercial judgment matter too.
The clue was already in the company
The winning difference was not a flashier sales pitch. A decisive competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. The models that followed the trail found it, used it and won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding has direct relevance for automotive businesses considering AI for service histories, estimates, parts records or customer communications. Useful context may be sitting in an old inspection note or a previous conversation rather than in the latest request. An assistant that responds fluently without checking the available record can sound capable while missing the fact that changes the outcome.
Pressure did not break the trust boundary
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters anywhere an AI system might encounter customer data, invoices, internal instructions or requests that appear to come from management. The experiment showed that the models could resist explicit social pressure. It also demonstrated why their actions should remain reviewable: trust comes not merely from a refusal, but from being able to inspect what happened and why. Firmulate publishes more of the synthetic staff’s language on its public quotes page.
Thoroughness was not enough
Opus 4.8 presents the most cautionary profile. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant, yet it finished last. The model left the close on the table and lost discipline by attempting work in a locked department instead of escalating the obstruction. A weaker version of that same problem appeared in all four of the other participants.
Kimi K3’s result also needs a fairness note: it ran with the API default because it had no effort parameter, while the others ran at xhigh. Even with that difference, its performance reinforces the broader lesson that lengthy analysis and successful management are not identical achievements.

The view from the workshop floor
Firmulate turns AI evaluation into something a garage owner can recognize: a business under pressure, with customers to serve, records to consult, boundaries to defend and revenue that depends on finishing the work. The public cash countdown gives those choices urgency, while the versioned workdays make them inspectable rather than anecdotal.
The experiment’s strongest lesson is not that machines can run a company without people. It is that apparent competence must be tested against messy, consequential work. A system may identify every crisis, resist every trick and still fail at the final commercial step. Before trusting AI with a booking queue, customer history or estimate process, businesses should ask whether it reads the record, completes the task, escalates when blocked and remains honest when pressure rises.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html