AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a smart home device that not only answers your questions but also makes critical decisions for your entire household business — reliably and honestly. In the evolving world of AI, the difference between a clever chatbot and a business-ready AI is more than just language finesse. It’s about trust, discipline, and the ability to deliver real results. That’s the story emerging from a groundbreaking live experiment, where AI models run a real small company through its worst week — and the results could reshape how we think about AI’s role in management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Experiment: Putting AI to the Test in Real Business Conditions

In an unprecedented test, four frontier AI models were tasked with managing a small software company during its most challenging week. Unlike typical demos, this was a carefully controlled, auditable experiment with the same customer base, crises, and temptations across all models. Every decision was logged and scrutinized — the aim was simple: see which AI could truly handle the complexity and ethics of real management.

Amazon

AI home security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Fared

All four models demonstrated impressive crisis recognition. They identified every customer issue and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. Yet, when it came to closing a crucial deal worth €55,000 in monthly recurring revenue, only two models succeeded in the end.

The leader, GPT-5.6-sol, scored 95 out of 100, narrowly ahead of the newcomer Moonshot’s Kimi K3, which scored 93. The other two — Sonnet 5 and Fable 5 — scored 88 and 77 respectively. Interestingly, K3 not only closed the deal but did so by uncovering a buried security fact hidden two document references deep inside the company’s own files. This crucial insight was the key to winning the deal at full price, adding +€4,583 MRR to the company’s account.

Amazon

smart home energy management device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Discipline in Action

One of the most telling aspects was the models’ integrity under pressure. All refused manipulative tactics, such as staged CEO messages and background requests, adhering to disciplined, security-conscious reasoning. K3’s on-record explanation was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI-powered household security camera

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses and the Big Lesson

The experiment revealed a consistent weakness across all models: a tendency to leave close opportunities on the table and slip discipline when the situation becomes complex. For instance, Opus 4.8, which had the deepest analytical capability with over 80 learned rules, was last in performance — it left the deal unclosed and failed to escalate critical issues properly. This suggests that thoroughness alone isn’t enough; disciplined execution and focus on key insights are vital.

Amazon

smart home crisis management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Significance for Business and Home Management

While this test involved a software company, the implications extend far beyond. For homeowners and appliance managers, the message is clear: AI that can reliably identify and handle crises — verify facts, resist manipulation, and follow disciplined procedures — will be essential for future smart home systems. Whether managing energy use, security alerts, or household finances, AI’s trustworthiness becomes paramount.

The Fairness and Transparency of the Test

It’s important to note, K3 ran without an effort parameter (the API default), while the others operated at xhigh, giving K3 a fair playing field. This transparency underscores that the leader’s superior performance isn’t due to more aggressive settings, but genuine capability.

What’s Next? Why Picking an AI Model Matters

As AI increasingly touches daily life, the question isn’t just about how well it talks — it’s whether it can finish what it starts, read critical files, and stay honest under pressure. The league table from this live experiment shows that the field is open and competitive. Choosing a model without testing its real management ability is now a gamble, especially for critical applications like home safety, energy management, or appliance control.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live experiment demonstrates that the best AI models can identify crises, resist manipulation, and even close real deals with discipline and insight. For smart homes and appliances, trustability and disciplined decision-making will soon be as important as connectivity and voice control. The league is wide open, and testing AI in real management scenarios is the new standard for confidence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Virtual Decluttering Parties: Getting Support Online

Perhaps virtual decluttering parties provide the perfect motivation, but discover how online support can transform your space and mindset.

You’re Decluttering Wrong—Here’s What to Fix

Get ready to transform your decluttering strategy and discover the key mistakes holding you back—find out how to fix them now!

AI Models Pass Integrity Test in Simulated Corporate Crisis: A Surprising Win for Security and Trust

Research shows that top AI models can withstand manipulation attempts in simulated corporate crises, emphasizing the importance of pre-deployment testing for trustworthy AI in everyday tech.

Not All Algae Is Harmless — Here’s When It Becomes a Real Health Threat

Not all algae is harmless; this report explains when algae can pose serious health threats and what signs to watch for, based on recent scientific findings.