AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI system tasked with running a small, real-world company through its toughest week—crises, customer demands, and ethical dilemmas. Surprisingly, even the most honest models score at least 26 out of 100, a baseline that reveals much about trust, competence, and the limits of current AI evaluation methods. This is the story behind a groundbreaking experiment that shows us why trustworthiness in AI isn’t just a bonus—it’s a requirement, especially when these systems are making real decisions for your business or home.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In a recent live experiment conducted by Firmulate, four state-of-the-art AI models were tested by running a simulated small software company through its most challenging week. They faced the same customers, crises, and ethical temptations, all in a controlled environment that mirrored real-world pressures. Every decision was recorded and auditable, offering a rare look at how AI performs not just in conversation but in management and decision-making roles.

Why the Baseline Scores Matter

The results might surprise many: even the most basic, do-nothing baseline AI scored 26 points out of 100. This score isn’t a mistake or a bug; it’s a reflection of the evaluation methodology that considers partial progress as meaningful. In other words, simply reading documents or recognizing crises counts toward the score, but it doesn’t mean the AI is competent or trustworthy.

The Crucial Role of Trust and Integrity

One of the most revealing aspects of the experiment was how models handled manipulation attempts. Fake CEO messages, staged escalation calls, and reporter tricks were all used to test whether the AI would be fooled or manipulated into unethical actions. Remarkably, all four models refused every manipulation attempt—something that even seasoned managers often struggle with. For example, when asked to authorize a dubious deal or escalate a request, the models consistently treated these as potential impersonations or approval bypasses, refusing to sign off without proper verification.

Deep Inside the Files: The Hidden Weakness

The true weakness in these models wasn’t in their crisis detection or honesty under attack—it was in their ability to read and utilize internal company documents. The experiment showed that models which read the company’s files and understood the context could close critical deals at full value (€4,583 MRR), whereas others left money on the table. This skill—reading and interpreting internal documents—proved to be the decisive factor in successful management and sales outcomes.

Behavior Under Discipline and Pressure

Another insight came from Opus 4.8, the most thorough participant. Despite its extensive rule set and deep analysis, it was the last in performance, leaving a key deal unclosed and slipping into process slips—such as writing attempts into a locked department instead of escalating them properly. This highlights that even the most disciplined AI can falter under continuous pressure if not designed to prioritize critical tasks effectively.

What Does This Mean for Your Business or Home?

For consumers considering AI in smart homes, appliances, or customer support, the takeaway is clear: performance isn’t just about how well an AI writes or responds. It’s about whether it can finish what it starts, read vital information, and stay honest when it matters most. An AI that scores a perfect chat demo but fails under pressure offers little real value.

Transparency and Testing Are Critical

The experiment’s transparency—decision versioning, auditable decision logs, and real-money mechanics—sets a new standard for AI evaluation. Enterprises and consumers alike should demand such rigorous testing before trusting AI with critical tasks. Firmulate offers live access to these experiments, allowing stakeholders to see AI models in action against real-world challenges at firmulate.com/live.

Conclusion: Beyond the Score

As AI becomes more integrated into our daily lives, understanding its true capabilities and limits is essential. The benchmark scores—ranging from 26 for a do-nothing baseline to 95 for the top model—serve not just as performance indicators but as trust signals. A high score means the AI can handle crises, resist manipulation, and read internal data—traits that are vital for trustworthy automation, whether in business or your smart home.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

smart home AI assistants

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

4 Home Decor Things That Are Not A Trend Yet — But Will Be In 2027

Four home decor elements are not currently trending but are projected to become popular by 2027, according to emerging trend signals and coverage interest.

Decision Fatigue and Decluttering: Making Quick Choices

Ongoing decluttering reduces decision fatigue, empowering quick choices—discover how simplifying your environment can transform your decision-making process.

Shoe Piles End Today: The Entryway Flow Map You’ve Never Tried

Want to finally banish shoe piles forever? Discover a proven entryway flow map that will transform your space and keep clutter out of sight.

Coffee Table Storage System: What Belongs Inside (and What Never Should)

Keen to master stylish coffee table storage? Discover what belongs inside and what never should for a clutter-free, chic living space.