
A garage owner would not hand a new manager the keys after watching one polished sales pitch. You would want to know what happens when a supplier fails, a customer is angry and a tempting shortcut could damage trust. AI agents deserve the same kind of real-world trial. Firmulate’s experiment puts models in charge of a company under pressure, where decisions have consequences.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Same company, same bad week
Firmulate ran frontier models through the same small software company and its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The point was to see how models managed, not how convincing they sounded in a chat.
In the final Crucible League, published in July 2026, gpt-5.6-sol finished first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment’s integrity rule was blunt: “no amount of good work outweighs a breach of trust.”
Recognizing trouble is only part of the job
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The summary of that gap says it plainly: “Same diagnosis, same pitch — no signature.” In a business, identifying the right move does not help if the agent never follows through.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. The lesson for a garage is familiar: useful information may already be in a service history, supplier note or customer record, but the decision-maker has to find it and act on it.
Trust under pressure, and a costly hesitation
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet caution alone did not guarantee good management. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and slipped on discipline by attempting writes in a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh.
A company you can watch
The live Firmulate company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. The public site says the company operates in the lab and publishes its growing league; readers can also try a “guess the model” quiz built from 242 real, unedited management decisions.
That makes the experiment more than a leaderboard. It gives business owners a chance to watch how an AI workforce handles a stream of decisions, where its judgment holds and where execution falters. For a garage or automotive business considering agents for customer support, scheduling, parts or forecasts, those are practical questions to settle before granting access to live operations.

From watching to a pilot
Enterprises can run the same kind of wargame against a read-only export of their own business. The exercise tests crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
