firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a test for AI honesty—one so straightforward that even a do-nothing baseline scores 26 out of 100. In a world where trust and accountability are critical, understanding how AI models are measured can change everything.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Unpacking the AI Benchmark: Simplicity, Standards, and Trust

At first glance, one might assume that an AI model that doesn’t do anything at all would score zero. But in the rigorous world of AI benchmarking, the reality is far more nuanced. The recent Firmulate experiment shows that even the most passive baseline scores a surprising 26 points, setting a clear floor for what honesty and minimal competence look like in AI management.

Why Does the Baseline Score 26?

This score isn’t a fluke or a flaw; it’s a reflection of the benchmark’s design. The methodology rewards partial progress—recognizing that even doing the bare minimum to identify crises, read important documents, or refuse manipulative requests counts as some level of competence. Conversely, it also emphasizes that a single breach of trust, like signing a manipulative deal, caps the total score at a certain point. In this case, that cap is 26.

The Value of Partial Progress

In real-world business, an AI isn’t expected to be perfect out of the gate. Instead, the benchmark measures how well it handles complex, pressure-filled scenarios, and whether it maintains honesty and diligence. Even models that refused manipulative tricks and identified critical information scored well—up to a point—but their failure to complete key tasks or breach trust kept scores from soaring higher.

Amazon

AI ethics and trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Experiment Works

The experiment placed four top AI models in the same scenario: running a small software company through a simulated, stressful week. The setup involved real crises—customer issues, ethical dilemmas, and manipulative attempts—all designed to test the model’s integrity and decision-making.

Consistent Challenges, Varied Outcomes

All four models successfully identified crises and refused manipulative requests, such as fake CEO messages or reporter tricks. For example, five out of five models refused to sign off on unethical requests, with the Kimi K3 model explicitly reasoning that the request could be impersonation or approval-bypass.

However, when it came to closing deals, only two models signed the €55,000 contract their own analysis had earned, while others left money on the table. The key difference was access to deeper documents—those that sat two references deep in the company’s files. Models that read the full files were able to win the deal at full price—a potential €4,583 monthly recurring revenue (MRR)—demonstrating that thoroughness and reading comprehension directly impact business outcomes.

Honest Models, Honest Outcomes

What’s notable is that models performed equally well in crisis detection and refusal, but their ability to read deeply and act decisively made the difference in closing the deal. The entire process kept transparency intact: decisions were versioned, auditable, and comparable—giving a clear view of how each model behaved under pressure.

Amazon

AI model transparency and benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Limitations and Lessons

One striking insight is that the most thorough model, Opus 4.8, despite analyzing more rules and performing the deepest assessments, ranked last in deal closure. Its discipline slipped, and it failed to escalate issues appropriately, leaving opportunities unexploited. This shows that even in perfection, discipline and process matter.

Additionally, the setup avoided bias by running models without effort parameters—meaning they operated at default settings. This ensured the results reflected genuine model behavior, not tuned incentives or shortcuts.

Amazon

AI decision-making audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and Trust

If AI is to become a trusted partner in customer management, support, or decision-making, the question isn’t just whether it can generate convincing language. It’s whether it can finish what it starts, verify facts from internal files, and resist manipulation or deceit. The benchmark’s design demonstrates that honesty and diligence are measurable, and that even a do-nothing baseline can set a meaningful floor for trustworthiness.

Amazon

AI deep reading comprehension tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Larger Takeaway

What do scores of 26 and above tell us? They show that AI models are capable of partial progress—identifying crises, refusing manipulation—and that the foundation of trust is built on simple, transparent behaviors. Importantly, a single breach of trust caps the score, reinforcing that honesty can’t be compromised without consequences.

Looking Ahead

Businesses considering AI integration need to evaluate how models perform under pressure—not just in chat demos, but in real decision-making scenarios where trust is tested. The Firmulate benchmark offers a transparent, watchable lab for this purpose, making the path toward trustworthy AI clearer and more accountable.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Signs You Are Experiencing Spiritual Awakening

I believe these signs indicate a profound spiritual awakening, and understanding them can help you navigate your transformative journey effectively.

How to Search for Angel Numbers

Yearning for spiritual guidance? Discover how to search for angel numbers and unlock personalized messages for your life path.

Angel Number 717: Trusting Divine Timing

Discover why Angel Number 717 urges you to trust divine timing and how embracing patience can transform your spiritual journey.

Angel Number 1212: Twin Flame Connection

Fascinated by angel number 1212? Discover how this powerful sign can illuminate your twin flame journey and what it truly means for your love life.