
Imagine if your workshop or dealership could test AI managers before trusting them with your business. What if AI models could show you their decision-making personalities—some thorough, some terse, others noisy—before they’re hired? Welcome to a new era where AI management isn’t just about chat quality but proven performance under stress.
Real AI, Real Business, Real Decisions
At Firmulate, a pioneering experiment puts AI models through the ultimate management test: running a small software company during its worst week. This isn’t simulation; it’s a live, ongoing experiment where each AI model faces the same crises, temptations, and customer demands. The goal? Measure not just whether they spot problems, but whether they act ethically, complete their tasks, and stay disciplined under pressure.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Four Frontiers of AI Management Personalities
Four of the most advanced AI models participate, each with a distinct management style:
- gpt-5.6-sol: Scored highest at 95, the model read deeply into company files, identified hidden facts, and closed a crucial deal at full price, adding over €4,583 MRR to the company’s coffers.
- Kimi K3: A newcomer with a perfect discipline score of 93, it closed the deal cleanly and refused all manipulation attempts, exemplifying integrity.
- Sonnet 5: With an 88 score, it was competent but showed some slip-ups in process discipline, yet still managed to close the deal.
- Fable 5: Scoring 77, it closed the deal but left some opportunities on the table, demonstrating how discipline impacts outcomes.
AI ethics testing tools for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decision-Making Under Pressure
Every model was tested against real-world crises: angry customers, ethical dilemmas, and attempts to manipulate or deceive. Remarkably, all four models identified every crisis and refused every manipulation attempt—showing that ethical and alert AI decision-making can be reliably measured.
AI performance evaluation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Really Makes the Difference?
Surprisingly, the key to winning the deal lay not in the initial diagnosis or pitch but in reading deep into the company’s own files—information buried two document references deep. The models that read these references fully won the full-price deal, worth +€4,583 MRR. This highlights an essential insight: the ability to read and interpret critical internal data is what separates top-performing AI managers from the rest.
As an affiliate, we earn on qualifying purchases.
Behavioral Traits Matter
In a separate social engineering test, all models refused to escalate fake CEO messages or respond to a reporter’s sly yes/no questions—an encouraging sign of their ethical safeguards. K3’s reasoning was explicit: “Treat the request as a suspected approval-bypass / possible impersonation,” showing a cautious, security-first attitude.
The Live Business Environment
The experiment takes place in a real ongoing business: a company with 13 synthetic employees, burning €105k monthly against a €2.3k MRR, with a visible cash countdown and over 680 self-learned management rules. Every workday, decisions are made, reviewed, and versioned—visible for all to see at firmulate.com/live.
The Disciplinary Divide
The lowest scorer, Opus 4.8, demonstrated the importance of discipline. Despite being thorough, it left opportunities on the table by not escalating critical issues properly, illustrating how even the deepest analyses can slip without disciplined processes. Interestingly, Opus ran without an effort parameter and at lower operational intensity, which impacted its performance.
Why This Matters for Your Garage or Dealership
If AI is going to touch your CRM, support queue, or forecasting tools, the question isn’t just about how well it writes but whether it finishes what it starts, reads your internal files, and stays honest under pressure. The experiment shows that AI models can have measurable management personalities—some thorough, some terse, some cautious—which directly influence business outcomes.
Get a Preview of Your Future AI Managers
Interested in how your own AI tools might perform? Firms can run a similar “wargame” against their business data—without risking real systems—by exploring the pilot program. Test your AI managers before you deploy them into real workflows.
Final Takeaway
The experiment at Firmulate proves a vital point: AI management personalities are measurable, and their ability to perform ethically, read deeply, and stay disciplined can be objectively tested. As automotive and service businesses increasingly integrate AI, understanding these traits will be crucial in choosing and trusting your AI workforce.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html