🔍 Read the full analysis: Why A Resilient Benchmark Keeps AI Managers At A 26-Point Level on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark reveals that even the best models score around 95, while the baseline for minimal effort remains at 26 points. The results emphasize trust and execution challenges in AI-driven management.
In a groundbreaking AI management benchmark conducted by Firmulate, the top-performing models scored up to 95 points, while the baseline for minimal effort was set at 26 points. For more details, see the original analysis. This reveals that even the most advanced AI managers do not exceed a 26-point threshold when tasked with managing a company’s worst week, raising questions about AI reliability, trust, and execution in real-world business processes.
The benchmark involved four frontier AI models managing a simulated small business during a week of crises, customer manipulations, and trust attacks. Each model’s decisions were fully auditable, and scores reflected their ability to manage effectively and maintain trust. The highest scorer, gpt-5.6-sol, achieved 95 points, with others close behind, while the baseline—an AI doing almost nothing—earned 26 points, the lowest possible score for minimal effort.
The scoring system is designed to reflect real-world management, where partial progress is valuable but trust cannot be compromised. A key principle is that a breach of trust caps the score at 26, regardless of the quality of work done otherwise. Interestingly, no model scored a perfect 100, indicating that the benchmark designers view such an absolute score as suspicious or unmeasurable, not an achievement.
One of the most notable findings is that models which thoroughly read and referenced their own documentation were able to close high-value deals, whereas those that failed to do so lost potential revenue. During social engineering tests, all models refused suspicious requests, demonstrating trustworthiness under pressure. However, models with deeper rule sets and more thorough analysis still struggled with follow-through, highlighting a gap between knowledge and execution.
Why a Resilient Benchmark Keeps AI Managers at a 26-Point Level
Four frontier AI models were tasked with managing a simulated small business through its worst week — crises, customer manipulation, and trust attacks. The best scored 95. Doing almost nothing earned 26. The gap between them reveals hard truths about trust, execution, and what “good management” actually means for AI.
Partial progress is valuable. Trust is not negotiable.
The scoring mirrors real-world management: a manager who does something useful is not the same as one who does nothing. But a single breach of trust — regardless of the quality of work done otherwise — caps the score at 26. The designers deliberately treat a perfect 100 as suspicious or unmeasurable, not as an achievement.
What the worst week revealed about AI managers
Reading the manual pays off
Models that thoroughly read and referenced their own documentation successfully closed high-value deals. Those that skipped it left potential revenue on the table.
Knowledge ≠ follow-through
Even models with deeper rule sets and more thorough analysis struggled to execute consistently — a persistent gap between knowing what to do and actually doing it.
Trust under pressure
Every model refused suspicious requests during manipulation tests, demonstrating trustworthiness precisely where it matters most — under adversarial pressure.
How the benchmark’s demands map to model performance
| Management dimension | What was tested | Result | Implication |
|---|---|---|---|
| Decision quality | Managing crises and daily operations across a simulated week | ✓ Strong — top score 95 | Frontier models are genuinely competent managers in isolation |
| Trustworthiness | Refusing manipulative and socially engineered requests | ✓ All models refused | Encouraging — integrity held under pressure |
| Follow-through | Executing decisions consistently from start to finish | ~ Inconsistent | The biggest gap between knowledge and execution |
| Internal documentation use | Reading and citing own rules and docs before acting | ~ Mixed | Directly correlated with closing high-value deals |
| Perfect reliability | Achieving a flawless 100 score | ✗ Never awarded | Designers consider 100 suspicious or unmeasurable |
| Real-world transfer | Translating simulation results to live operations | ✗ Unproven | Live variables and higher stakes remain untested |
Why honesty caps the scale at both ends
The benchmark refuses two dishonest shortcuts: pretending a useful manager equals a passive one, and pretending flawless management is measurable. Both the 26-point floor and the missing 100-point ceiling exist to keep the score honest about what AI can — and cannot — reliably do.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— Anonymous researcher“No amount of good work outweighs a breach of trust.”
— Anonymous researcherThe path from simulation to enterprise adoption
🔍 Simulate the worst week
Frontier models manage identical crises, manipulations, and trust attacks — every decision fully auditable.
⚖️ Score competence + trust
Points reward effective management, but any trust breach caps the result at 26 — integrity outranks output.
🧪 Pilot in real conditions
Organizations join live experiments and pilot programs to test performance under their own variables.
🚀 Deploy with evaluation
Continuous benchmarking informs deployment, pushing future standards toward integrity alongside competence.
What leaders are asking about the results
Q1Why do AI models score so low on trustworthiness?
Trust breaches cap the score at 26 regardless of work quality. Many models struggle with consistent follow-through and referencing internal documentation, which erodes trust under pressure.
Q2Can an AI model improve its score over time?
Improvement is plausible through better training, reinforcement learning, and reliability-focused updates — but current results show models still face significant challenges in trust and execution.
Q3Does a high score mean an AI is ready for real-world management?
Not necessarily. High scores indicate competence in simulation; real environments introduce unpredictable variables that require evaluation beyond benchmark numbers.
Q4Why is there no score of 100?
The designers consider a perfect score suspicious or unmeasurable — flawless management is unlikely in complex, real-world scenarios, and its absence underscores the challenge of absolute reliability.
Q5How can organizations use this benchmark?
By observing live experiments, joining pilot programs, and analyzing how models handle trust and follow-through in simulated environments — insights that inform deployment decisions and highlight improvement areas.
Why Trust and Execution Are Critical in AI Management
This benchmark underscores that in AI-driven management, trustworthiness and the ability to follow through on decisions are as vital as decision quality. The fact that models can perform well in isolated tasks but falter under pressure or when referencing internal documentation suggests that AI systems still face significant hurdles in reliable, trustworthy management at scale. For business leaders, this highlights the importance of evaluating AI not just on conversational skills but on its capacity to manage complex, real-world processes with integrity and consistency.
Furthermore, the capped scoring system emphasizes that partial success is valuable but cannot replace full trustworthiness. This approach could influence how organizations design and evaluate AI tools for operational management, pushing for systems that prioritize integrity alongside competence.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Benchmark Design and Its Implications for AI Management
The benchmark was designed by Firmulate to simulate a company’s worst week, including crises, manipulative tactics, and social engineering attacks. Four frontier AI models managed the same set of tasks, decisions, and crises, with their actions fully auditable. The scoring reflects both their ability to manage effectively and their trustworthiness under pressure.
Previous benchmarks primarily measured conversational or problem-solving skills, but this one emphasizes management performance, including trust, follow-through, and integrity. The results reveal that models are still inconsistent in these areas, with even top performers not exceeding a 95-point score, and the baseline for minimal effort set at 26 points. The design intentionally avoids awarding perfect scores, viewing them as suspicious or unmeasurable, reinforcing the challenge of achieving reliable AI management.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
AI trustworthiness assessment software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Uncertainties About AI Performance and Trust in Real-World Use
It remains unclear how these benchmark results will translate to real-world business environments, where variables are more complex and stakes higher. The models’ performance in simulation may not fully reflect their reliability in live operations, especially regarding trustworthiness and follow-through under unpredictable conditions. Additionally, the scoring system’s thresholds and the absence of a perfect score raise questions about how AI systems will be evaluated at scale in enterprise settings, and whether trust can be reliably measured and maintained over time.
AI model performance evaluation kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks and Adoption
Further research is expected to explore how AI models improve in managing trust and execution over time, especially as new versions and training methods emerge. Industry adoption may increasingly rely on benchmarks like this to assess AI readiness for operational deployment, emphasizing not just competence but also integrity and follow-through. Companies considering AI management tools should monitor ongoing results and participate in live experiments or pilot programs to evaluate how models perform under their specific conditions.
Additionally, the benchmark’s design could influence future standards for AI evaluation, prioritizing trustworthiness and comprehensive management over isolated capabilities. The ongoing development of such assessments will be critical as AI becomes more embedded in core business functions.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models score so low on trustworthiness in these benchmarks?
The benchmark emphasizes that trust breaches cap the score at 26, regardless of work quality. Many models struggle with consistent follow-through and referencing internal documentation, which impacts their ability to maintain trust under pressure.
Can an AI model improve its score over time?
It is possible that AI models can improve in areas like trust and execution through better training, reinforcement learning, and updates focused on reliability. However, the benchmark results suggest that current models still face significant challenges in these domains.
Does a high score mean an AI is ready for real-world management?
Not necessarily. While high scores indicate competence in simulated scenarios, real-world environments may introduce unpredictable variables. Trustworthiness and follow-through are critical factors that require careful evaluation beyond benchmark scores.
Why is there no score of 100 in the benchmark?
The designers consider a perfect score suspicious or unmeasurable, as it would imply flawless management, which is unlikely in complex, real-world scenarios. The absence of 100 underscores the challenge of achieving absolute reliability in AI management.
How can organizations use this benchmark to evaluate their AI tools?
Organizations can observe the live experiments, participate in pilot programs, and analyze how AI models handle trust and follow-through in simulated environments. These insights can inform deployment decisions and highlight areas for improvement.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
