
Imagine testing your DIY woodworking project, only to find that even the most basic measures still score some points — not zero. That’s a bit like how AI benchmarks work today, where even doing nothing at all earns a baseline score of 26. For business leaders, understanding this hidden floor is crucial, because it reveals what honest AI performance really looks like, especially when stakes are high.
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Understanding the Baseline: More Than Just Zero
In a recent public experiment run by Firmulate, four advanced AI models were tasked with managing a small software company’s worst week — handling customer crises, reading critical files, and resisting manipulation attempts. Interestingly, even the most passive baseline model scored 26 out of 100. This isn’t a mistake or a flaw; it’s a reflection of how the benchmarking methodology is designed.
The Role of Partial Progress
While the models had to confront real-world challenges, the system acknowledges partial successes, giving points for progress rather than all-or-nothing results. For example, if an AI reads a crucial document deep in a file and acts accordingly, it earns credit for that insight — even if it fails to close a deal or fully resist manipulation. This approach ensures that incremental improvements are recognized, aligning with how actual business decisions often work.
Trust Matters More Than Good Work
Perhaps most revealing is that a single breach of trust caps the overall score. No matter how many crises the AI handles flawlessly, one successful manipulation attempt — like accepting a fake CEO message — prevents it from reaching top marks. This highlights a fundamental truth: in business, trustworthiness is paramount. A model that slips once can undermine all its achievements, which is why honest benchmarks account for this vulnerability explicitly.
As an affiliate, we earn on qualifying purchases.
What the Experiment Reveals About AI in Business
All four tested models successfully identified crises and refused manipulative offers. They recognized fake CEO messages, escalated appropriately, and stayed honest. Yet, only two completed the transaction that their analysis suggested was correct, signing a €55,000 deal at full price. The other two, despite good diagnoses, faltered in closing — demonstrating that discipline and process integrity are just as vital as insight.
The Hidden Weakness: Reading Deep in Files
The key vulnerability was not apparent at first glance. While models performed well in handling visible crises, their weakness was hidden two document references deep inside the company’s files. Those that read these files correctly secured the full deal, adding €4,583 in monthly recurring revenue. This emphasizes that thorough attention to detail and deep information gathering are essential for real-world performance.
For Business Leaders: What Matters Beyond the Score
The experiment underscores that evaluating AI isn’t just about how well it chats or answers questions. It’s about whether it can finish tasks, stay honest under pressure, and read deeply into complex information. For decision-makers, this means focusing on how AI models perform in realistic, high-stakes environments, not just in scripted demos.
Implications for Your AI Strategy
For companies exploring AI tools, the takeaway is clear: don’t be fooled by superficial performance scores. A good benchmark reflects whether the AI can be trusted to follow through on commitments, resist manipulation, and act in your best interest — even when tested under pressure.
Firmulate’s Live Experiment and Ongoing Benchmarks
Through Firmulate’s live platform, businesses can simulate their own operations against these models — running real crises, assessing decision quality, and identifying weaknesses before deploying AI in critical roles. Every decision is versioned and auditable, providing transparency and confidence in the AI’s capabilities.
What the Future Holds
As AI continues to integrate into decision-making workflows, understanding these nuanced performance metrics becomes vital. A model’s ability to stay honest, read deeply, and finish what it starts will determine whether it adds value or introduces risk. The benchmark’s honest design, acknowledging a baseline of 26 points for doing nothing, sets a truthful standard that business leaders should heed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
