TL;DR
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A live experiment tested AI models managing a small company under crisis conditions. The results show that management skills, not just chat quality, determine success. The top model scored 95, but trust breaches and decision accuracy proved crucial.
During a live, ongoing experiment, five AI models managed a small software company through its worst week, revealing a new performance hierarchy based on management quality rather than chat responses. The top model, GPT-5.6-SOL, scored 95 points, but the experiment also exposed critical gaps in decision execution and trust management, which are not captured by traditional benchmarks.
The experiment, conducted by Firmulate, simulated a crisis week for a small company with real money mechanics and 13 synthetic employees. For more details on the methodology, see the original analysis. Each AI model was tasked with diagnosing issues, communicating with stakeholders, negotiating deals, and escalating problems, all under strict trust conditions. The models’ performance was scored based on their ability to identify crises, reject manipulation, and complete key business tasks, including signing deals worth over €55,000.
While all models detected crises and refused social engineering attempts, only two managed to secure the commercial deal. This highlights the importance of management skills, as discussed in the original analysis. The highest scorer, GPT-5.6-SOL, achieved a 95-point rating, while others lagged behind, with Fable 5 scoring 77 and Opus 4.8 at 73. Notably, Opus 4.8 provided the most detailed analysis but failed to close deals or escalate effectively, illustrating that thoroughness alone does not guarantee management success.
The experiment also measured trust breaches, with any breach capping the overall score. Despite strong diagnostic skills, models struggled with real-world decision-making, such as retrieving critical information buried in documents, which ultimately impacted their ability to close sales and manage risks effectively. The results challenge the assumption that more activity or detailed analysis equates to better management, emphasizing the importance of decision quality and trustworthiness. This aligns with insights from the original analysis of AI performance benchmarks.
Why Management Skills Trump Chat Quality in AI Evaluations
This experiment underscores that AI’s ability to manage complex, real-world business scenarios depends on decision-making, trustworthiness, and operational discipline, not just generating convincing responses. It suggests that future AI benchmarks should incorporate management and execution metrics to better reflect practical utility. For organizations exploring AI for decision support, understanding these qualities is essential to avoid overestimating an AI’s capabilities based solely on superficial performance or chat fluency.
As an affiliate, we earn on qualifying purchases.
The Shift Toward Management-Centric AI Evaluation
Traditional AI benchmarks have focused on technical proficiency—such as code accuracy or conversational fluency—without assessing how models perform under real-world pressures. The Firmulate experiment is a pioneering effort to evaluate models in a simulated business environment, exposing the gap between apparent competence and effective management. Previous assessments rarely captured trust, escalation, or decision quality, which are critical in operational settings. The July 2026 Crucible League results mark a significant step toward more holistic AI evaluation frameworks that prioritize management skills.
Prior to this, most benchmarks relied on static tests or simulated tasks that do not replicate the complexities of ongoing organizational crises. The live, versioned decision-making process used here provides a more accurate picture of how models might perform in actual enterprise contexts, especially when trust and accountability are at stake.
“This experiment reveals that management quality—trust, decision-making, escalation—must become its own category in AI evaluation, beyond chat and coding benchmarks.”
— Thorsten Meyer, Lead Researcher at Firmulate
As an affiliate, we earn on qualifying purchases.
Unclear Impact of Long-Term Trust and Decision Quality
It remains uncertain how these findings will translate to larger organizations or different industries. The experiment focused on a small, simulated company, and real-world complexities could introduce additional variables. The long-term impact of integrating management-focused AI evaluation metrics into standard benchmarks is still under discussion, and whether models can improve in these areas with further training is not yet known.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Adoption
Future research will likely explore refining management-centric benchmarks, including real-time decision tracking and trust metrics. Organizations considering AI for management tasks should pilot models with live scenarios, emphasizing trust, escalation, and decision execution. The industry may also develop new standards for evaluating AI in operational roles, moving beyond chat and coding scores to include management effectiveness.
Additionally, ongoing experiments like Firmulate’s are expected to expand, testing models’ ability to handle more complex, long-term management challenges across diverse industries, ultimately shaping how AI is integrated into enterprise decision-making processes.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the new leaderboard reveal about AI management skills?
The leaderboard shows that management skills—trustworthiness, decision quality, escalation—are critical for success, often more so than chat fluency or diagnostic accuracy. The top model scored 95, but many others failed in execution or trust management, highlighting the importance of operational discipline.
Why is trust important in AI management models?
Trust determines whether an AI model can be relied upon to make ethical, accurate, and appropriate decisions, especially under pressure. Breaches of trust cap performance and can lead to failures in critical tasks like signing deals or escalating issues.
How does this experiment affect future AI evaluation standards?
It suggests that benchmarks need to include real-world management metrics—such as decision accuracy, escalation, and trustworthiness—rather than focusing solely on response quality or technical tasks.
Can models improve in management skills over time?
Potentially, yes. Ongoing training and better evaluation frameworks could help models develop stronger decision-making and trust management capabilities, but this remains an area for further research.
What should organizations consider before deploying AI in management roles?
Organizations should assess whether AI models can read organizational context, escalate issues properly, and maintain trust. Relying solely on superficial performance metrics could lead to ineffective or risky management decisions.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
