firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine if your garage could be run by an AI that not only handles customer crises but also spots buried information and stays honest under pressure. That’s no longer science fiction—it’s the real-world test happening now, with AI models competing in a fierce league of performance. The latest results reveal surprising leaders and highlight crucial gaps that could shape the future of AI-driven business management.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Benchmarking AI in a Live Business Environment

In a groundbreaking experiment, four leading AI models were tasked with managing a real, small software company facing its worst week. The company, with synthetic employees and real money mechanics, experiences daily crises—customer issues, security threats, and fraudulent attempts. The goal? See which AI could best diagnose problems, resist manipulation, and close profitable deals.

This isn’t a simulation with canned responses. Every decision was real, auditable, and consistent across models. The models faced the same crises, same customer files, and same temptation to cheat or cut corners. The experiment, run by Firmulate, offers a rare window into how AI might perform in actual business settings—far beyond chatbots or simple automation.

Key Findings: Performance Scores and Deal Closures

  • The models were scored on a 100-point scale based on their ability to identify issues, resist manipulation, and close deals.
  • The top performer, gpt-5.6-sol, scored 95 and successfully closed a €55,000 deal, matching analytical accuracy with practical results.
  • Close behind was the newcomer, Kimi K3, with a score of 93. K3 not only closed the deal but also demonstrated the cleanest discipline, resisting all three social engineering bait attempts.
  • Sonnet 5 scored 88 and closed the deal but with some process slips, while Fable 5 scored 77, and Opus 4.8 lagged at 73.

All models identified every crisis and refused manipulation attempts, but only two completed the deal based on their own analysis. Notably, the key to winning the deal was uncovering a buried fact—hidden two document references deep in the company’s files—not in the customer event itself. Models that reviewed the internal documents succeeded in closing the full-price deal, worth over €4,580 in additional monthly revenue.

Understanding the Weaknesses and Discipline

The experiment also revealed a crucial weakness shared among the models. The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last because it left the close on the table and slipped in discipline—writing attempts into a locked department instead of escalating. These slips, though subtle, suggest that even the most capable models can falter under pressure or in discipline.

Resisting Social Engineering and Manipulation

All models were tested against social engineering tactics, including fake CEO messages and a reporter trick. Remarkably, each one refused these manipulative requests, with Kimi K3 explicitly reasoning that such requests could be impersonation attempts. This demonstrates that models can be trained not just to diagnose but also to recognize and resist attempts to deceive—an essential trait for trustworthy AI in management roles.

Fairness and Experiment Conditions

It’s important to note that Kimi K3 ran without an effort parameter (the API default), while other models operated at xhigh. This fairness note helps contextualize the performance differences and underscores that even with different settings, K3’s performance remains impressive.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond Demos: Real Business at Stake

The live company managed by these models operates with 13 synthetic employees, spending €105,000 monthly against €2,300 in monthly recurring revenue. Its daily operations, decision-making, and crisis management are publicly observable at firmulate.com/live. This transparency underscores a vital point: these models aren’t just performing in controlled tests—they’re managing a real business in real time.

Why does this matter for automotive or garage operations? Because whether managing customer relationships, inventory, or security, the core questions are the same: does the AI finish what it starts, does it read all relevant information before acting, and does it stay honest under pressure? The performance of Kimi K3 suggests that a newcomer can outperform established models in critical management tasks—making choosing AI tools a matter of testing, not assumptions.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Broader Implication for Business Decisions

This experiment isn’t about chat quality or superficial demos. It’s about real work—diagnosing problems, resisting manipulation, closing deals, and maintaining discipline. As the league table shows, the performance gap is clear, and the field is open for new entrants with the right approach.

For your garage or automotive business, the takeaway is clear: don’t rely solely on reputation or demo scripts. Test your AI solutions in real-world scenarios—like Firmulate’s live wargame—and see which model truly delivers. The future of AI-driven management hinges on trustworthiness, discipline, and actual results, not just clever conversation.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In a live business test, an AI model from Moonshot outperformed three Western frontier models, with a score of 93, closing the deal and resisting manipulation—proof that choosing the right AI requires real testing, not assumptions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and fraud detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Real Meaning of “Shear Stability” in Gear Oils

Her understanding of shear stability in gear oils is essential to prevent gear wear; discover what this crucial property truly means.

Liqui Moly

Liqui Moly announces new distribution centers in Ukraine to meet rising demand for automotive products, signaling strategic growth in Eastern Europe.

The Feds Diluted Gas To Cut Prices. Diesel Has No Such Fix

Federal authorities have reportedly diluted gasoline supplies to reduce prices, but diesel prices remain unchanged, raising questions about fuel market strategies.

AI’s Hidden Strength: Why Focusing on Closing Matters More Than Chat Quality for Garage Business Automation

Live AI experiments show that reading internal files and resisting manipulation are key to closing deals, not just generating convincing chat. Automate smarter for your garage.