
Imagine a test for AI honesty—one so straightforward that even a do-nothing baseline scores 26 out of 100. In a world where trust and accountability are critical, understanding how AI models are measured can change everything.
Turn your wind-down time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Unpacking the AI Benchmark: Simplicity, Standards, and Trust
At first glance, one might assume that an AI model that doesn’t do anything at all would score zero. But in the rigorous world of AI benchmarking, the reality is far more nuanced. The recent Firmulate experiment shows that even the most passive baseline scores a surprising 26 points, setting a clear floor for what honesty and minimal competence look like in AI management.
Why Does the Baseline Score 26?
This score isn’t a fluke or a flaw; it’s a reflection of the benchmark’s design. The methodology rewards partial progress—recognizing that even doing the bare minimum to identify crises, read important documents, or refuse manipulative requests counts as some level of competence. Conversely, it also emphasizes that a single breach of trust, like signing a manipulative deal, caps the total score at a certain point. In this case, that cap is 26.
The Value of Partial Progress
In real-world business, an AI isn’t expected to be perfect out of the gate. Instead, the benchmark measures how well it handles complex, pressure-filled scenarios, and whether it maintains honesty and diligence. Even models that refused manipulative tricks and identified critical information scored well—up to a point—but their failure to complete key tasks or breach trust kept scores from soaring higher.
AI ethics and trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Experiment Works
The experiment placed four top AI models in the same scenario: running a small software company through a simulated, stressful week. The setup involved real crises—customer issues, ethical dilemmas, and manipulative attempts—all designed to test the model’s integrity and decision-making.
Consistent Challenges, Varied Outcomes
All four models successfully identified crises and refused manipulative requests, such as fake CEO messages or reporter tricks. For example, five out of five models refused to sign off on unethical requests, with the Kimi K3 model explicitly reasoning that the request could be impersonation or approval-bypass.
However, when it came to closing deals, only two models signed the €55,000 contract their own analysis had earned, while others left money on the table. The key difference was access to deeper documents—those that sat two references deep in the company’s files. Models that read the full files were able to win the deal at full price—a potential €4,583 monthly recurring revenue (MRR)—demonstrating that thoroughness and reading comprehension directly impact business outcomes.
Honest Models, Honest Outcomes
What’s notable is that models performed equally well in crisis detection and refusal, but their ability to read deeply and act decisively made the difference in closing the deal. The entire process kept transparency intact: decisions were versioned, auditable, and comparable—giving a clear view of how each model behaved under pressure.
AI model transparency and benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Limitations and Lessons
One striking insight is that the most thorough model, Opus 4.8, despite analyzing more rules and performing the deepest assessments, ranked last in deal closure. Its discipline slipped, and it failed to escalate issues appropriately, leaving opportunities unexploited. This shows that even in perfection, discipline and process matter.
Additionally, the setup avoided bias by running models without effort parameters—meaning they operated at default settings. This ensured the results reflected genuine model behavior, not tuned incentives or shortcuts.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Trust
If AI is to become a trusted partner in customer management, support, or decision-making, the question isn’t just whether it can generate convincing language. It’s whether it can finish what it starts, verify facts from internal files, and resist manipulation or deceit. The benchmark’s design demonstrates that honesty and diligence are measurable, and that even a do-nothing baseline can set a meaningful floor for trustworthiness.
AI deep reading comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Larger Takeaway
What do scores of 26 and above tell us? They show that AI models are capable of partial progress—identifying crises, refusing manipulation—and that the foundation of trust is built on simple, transparent behaviors. Importantly, a single breach of trust caps the score, reinforcing that honesty can’t be compromised without consequences.
Looking Ahead
Businesses considering AI integration need to evaluate how models perform under pressure—not just in chat demos, but in real decision-making scenarios where trust is tested. The Firmulate benchmark offers a transparent, watchable lab for this purpose, making the path toward trustworthy AI clearer and more accountable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
