
Imagine your smart home devices not only managing your appliances but also making critical decisions about your business operations. How reliable would they be under pressure? Recently, a groundbreaking live experiment tested four advanced AI models by running a real small software company through its worst week — with eye-opening results that could reshape how we view AI trustworthiness in everyday life and business.
The Live AI Management Challenge
At the heart of this experiment was a unique setup: four frontier AI models were tasked with running the same small software company facing a series of crises, temptations, and ethical dilemmas. This company, which relies on real money mechanics and daily operations, is openly available for observation at firmulate.com/live. Every decision made by these AI ‘managers’ was recorded, versioned, and transparent, providing an unprecedented look into how AI behaves as a business leader under pressure.
AI management software for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings from the AI Showdown
Despite all four models identifying every major crisis and resisting every attempt at manipulation — like social engineering attempts and fake CEO messages — only two of them managed to close a crucial €55,000 deal based on their own analysis and diagnosis. The other two, despite similar decision-making, left the deal on the table, showing a notable lapse in follow-through or discipline.
Interestingly, the decisive advantage lay in the models’ ability to read deeply into company files. The winning models uncovered a hidden document reference that led to the full deal, adding +€4,583 MRR — a performance invisible in typical chat demos where models are often praised for surface-level interaction.
As an affiliate, we earn on qualifying purchases.
Personality Profiles of the AI Models
Each model displayed a distinct management personality:
- gpt-5.6-sol: The top performer, scored 95, demonstrated thoroughness and meticulous analysis, ensuring all critical information was considered before closing deals.
- Kimi K3: The newcomer, scoring 93, showed the cleanest discipline of the field, refusing manipulative tactics and sticking to principles, yet slightly more cautious in closing deals.
- Sonnet 5: Scoring 88, and another with a similar score, both closed the deals but with some slips in process discipline, leaving opportunities on the table.
- Opus 4.8: With a score of 73, displayed the most depth in analysis (+80 learned rules) but faltered in execution, particularly in escalating issues appropriately, leaving discipline slip-ups.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Ethical and Practical Implications
All models refused to engage in social engineering or fake CEO message escalations, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This suggests that these models are not only capable of avoiding unethical shortcuts but also do so in a manner that is consistent and explainable.
In the real-world live company’s environment — which burns €105k monthly against a €2.3k MRR — these AI systems are tested daily in a high-stakes setting. The experiment proves that AI models can be honest, attentive, and disciplined, but the level of thoroughness varies and impacts performance significantly.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Everyday Smart Homes
While the experiment focuses on a software company’s management, the implications extend beyond. If AI can reliably read deeply into documents, make honest decisions, and resist manipulation under pressure, the same principles apply to smart home management — from energy use to security and maintenance. As these AI systems become integrated into our daily lives, knowing their decision-making personalities helps us choose models that best align with our values and trust levels.
The Takeaway
As AI models increasingly handle critical tasks, their management personalities matter. The live experiment shows that even in high-pressure business situations, some models demonstrate thoroughness, honesty, and discipline, leading to better outcomes. For consumers and enterprises alike, understanding these traits is key to selecting AI that not only performs well on surface but also finishes what it starts, reads important information deeply, and stays honest under pressure.
Curious to see how your AI workforce stacks up? Try the interactive quiz at firmulate.com/quiz.html and discover which model might be managing your future smart home or business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html