
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Why Even a Do-Nothing AI Gets a Score of 26
For crypto and Bitcoin investors, trusting an AI system to handle complex tasks is a leap of faith. But how do we measure whether these systems are honest and effective? The recent Firmulate experiment offers insight: even the baseline AI, which does nothing but passively observe, scores 26 points out of 100. This isn’t an error—it’s a carefully designed floor that reflects fundamental issues of trust and partial progress in AI performance benchmarks.
As an affiliate, we earn on qualifying purchases.
Understanding the Benchmark Methodology: More Than Just Scores
The Firmulate live experiment pits four frontier AI models against a simulated weekly crisis at a small software company. Each model is tasked with managing customer relationships, navigating crises, and making decisions that could lead to profit or loss. All decisions are fully versioned and auditable, making transparency a core feature of the test. The key takeaway? The scoring isn’t just about who gets the most crises right. It incorporates partial progress and, crucially, recognizes that even a simple, do-nothing approach scores 26 points.
The Significance of Partial Progress
In this benchmark, every decision matters—whether it’s spotting a hidden document reference or refusing a manipulative request. Even if a model refuses to act, it still earns partial credit for recognizing the crisis. This means that the baseline, which does nothing but observe, still earns 26 points because it at least acknowledges the existence of an issue. This lower limit sets a realistic bar: an AI must do more than just ignore problems to excel.
Trust Breaches Cap the Score at 26
Another vital aspect: a single breach of trust—such as signing a fraudulent deal—immediately caps the overall score. For example, even if an AI performs flawlessly on all other tasks, one lapse in honesty drops its score to the baseline level. This emphasizes that reliability and trustworthiness are non-negotiable qualities for AI systems in real-world business settings, including crypto and Bitcoin sectors where integrity is paramount.
As an affiliate, we earn on qualifying purchases.
Surprising Findings from the Live Experiment
Despite the complexity of the scenario, all models successfully identified crises and refused manipulative requests, including staged social engineering attacks. In particular, all models refused a staged CEO request that escalated through multiple stages, a fake reporter query, and other manipulative tactics. Kimi K3, for instance, explicitly treated such requests as potential impersonation, reflecting a cautious and trustworthy stance.
However, the real competitive edge came from reading internal company files. The models that examined documents that were two references deep in the company’s own files managed to close the deal at full price—+€4,583 MRR—while those that failed to do so left revenue on the table. This underscores that comprehensive, diligent data review is crucial for trustworthy and effective AI performance.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Crypto and Bitcoin Stakeholders
In the fast-moving world of crypto and Bitcoin, where trusting AI-driven solutions for support, compliance, or trading decisions is becoming commonplace, knowing what a benchmark measures is critical. It’s not enough for AI to generate convincing chat responses. The question is: does it follow through on commitments, read the right documents before acting, and stay honest under pressure?
Firmulate’s live experiment lays bare these considerations. It shows that even the least active, do-nothing baseline scores 26 points—highlighting the importance of genuine progress. More importantly, it illustrates that trustworthiness in AI isn’t just about avoiding mistakes but actively recognizing and resisting manipulation and deception.
h2>Key Takeaways for Investors and Developers
- The benchmark assesses management quality, not just chat quality.
- Partial progress—like recognizing crises or reading internal documents—contributes to the score.
- A single breach of trust caps the overall score, underscoring the importance of honesty.
- Models that read deeper into company files close bigger deals, showing that diligence pays off.
For crypto and Bitcoin enterprises deploying AI solutions, these findings stress that trust, diligence, and integrity are non-negotiable. The real test of an AI system isn’t just how well it responds but whether it can finish what it starts—reading files thoroughly, refusing manipulation, and maintaining honesty under pressure.

AI compliance monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Trust and Integrity Matter More Than Scores
The Firmulate experiment confirms that in high-stakes environments—like crypto and Bitcoin—AI performance is measured not just by output but by trustworthiness. A score floor of 26 for doing nothing reminds us that partial progress and honesty are foundational. As AI becomes more embedded in financial systems, what truly matters is whether these systems can be relied upon to act ethically and thoroughly, not just perform well in demos.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
