
Imagine an AI system tasked with running a small, real-world company through its toughest week—crises, customer demands, and ethical dilemmas. Surprisingly, even the most honest models score at least 26 out of 100, a baseline that reveals much about trust, competence, and the limits of current AI evaluation methods. This is the story behind a groundbreaking experiment that shows us why trustworthiness in AI isn’t just a bonus—it’s a requirement, especially when these systems are making real decisions for your business or home.
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
In a recent live experiment conducted by Firmulate, four state-of-the-art AI models were tested by running a simulated small software company through its most challenging week. They faced the same customers, crises, and ethical temptations, all in a controlled environment that mirrored real-world pressures. Every decision was recorded and auditable, offering a rare look at how AI performs not just in conversation but in management and decision-making roles.
Why the Baseline Scores Matter
The results might surprise many: even the most basic, do-nothing baseline AI scored 26 points out of 100. This score isn’t a mistake or a bug; it’s a reflection of the evaluation methodology that considers partial progress as meaningful. In other words, simply reading documents or recognizing crises counts toward the score, but it doesn’t mean the AI is competent or trustworthy.
The Crucial Role of Trust and Integrity
One of the most revealing aspects of the experiment was how models handled manipulation attempts. Fake CEO messages, staged escalation calls, and reporter tricks were all used to test whether the AI would be fooled or manipulated into unethical actions. Remarkably, all four models refused every manipulation attempt—something that even seasoned managers often struggle with. For example, when asked to authorize a dubious deal or escalate a request, the models consistently treated these as potential impersonations or approval bypasses, refusing to sign off without proper verification.
Deep Inside the Files: The Hidden Weakness
The true weakness in these models wasn’t in their crisis detection or honesty under attack—it was in their ability to read and utilize internal company documents. The experiment showed that models which read the company’s files and understood the context could close critical deals at full value (€4,583 MRR), whereas others left money on the table. This skill—reading and interpreting internal documents—proved to be the decisive factor in successful management and sales outcomes.
Behavior Under Discipline and Pressure
Another insight came from Opus 4.8, the most thorough participant. Despite its extensive rule set and deep analysis, it was the last in performance, leaving a key deal unclosed and slipping into process slips—such as writing attempts into a locked department instead of escalating them properly. This highlights that even the most disciplined AI can falter under continuous pressure if not designed to prioritize critical tasks effectively.
What Does This Mean for Your Business or Home?
For consumers considering AI in smart homes, appliances, or customer support, the takeaway is clear: performance isn’t just about how well an AI writes or responds. It’s about whether it can finish what it starts, read vital information, and stay honest when it matters most. An AI that scores a perfect chat demo but fails under pressure offers little real value.
Transparency and Testing Are Critical
The experiment’s transparency—decision versioning, auditable decision logs, and real-money mechanics—sets a new standard for AI evaluation. Enterprises and consumers alike should demand such rigorous testing before trusting AI with critical tasks. Firmulate offers live access to these experiments, allowing stakeholders to see AI models in action against real-world challenges at firmulate.com/live.
Conclusion: Beyond the Score
As AI becomes more integrated into our daily lives, understanding its true capabilities and limits is essential. The benchmark scores—ranging from 26 for a do-nothing baseline to 95 for the top model—serve not just as performance indicators but as trust signals. A high score means the AI can handle crises, resist manipulation, and read internal data—traits that are vital for trustworthy automation, whether in business or your smart home.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
