📊 Full opportunity report: The Hidden Leaderboard That Emerges After The AI Demo on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live experiment tested AI models managing a small company under crisis conditions. The results show that management skills, not just chat quality, determine success. The top model scored 95, but trust breaches and decision accuracy proved crucial.
During a live, ongoing experiment, five AI models managed a small software company through its worst week, revealing a new performance hierarchy based on management quality rather than chat responses. The top model, GPT-5.6-SOL, scored 95 points, but the experiment also exposed critical gaps in decision execution and trust management, which are not captured by traditional benchmarks.
The experiment, conducted by Firmulate, simulated a crisis week for a small company with real money mechanics and 13 synthetic employees. For more details on the methodology, see the original analysis. Each AI model was tasked with diagnosing issues, communicating with stakeholders, negotiating deals, and escalating problems, all under strict trust conditions. The models’ performance was scored based on their ability to identify crises, reject manipulation, and complete key business tasks, including signing deals worth over €55,000.
While all models detected crises and refused social engineering attempts, only two managed to secure the commercial deal. This highlights the importance of management skills, as discussed in the original analysis. The highest scorer, GPT-5.6-SOL, achieved a 95-point rating, while others lagged behind, with Fable 5 scoring 77 and Opus 4.8 at 73. Notably, Opus 4.8 provided the most detailed analysis but failed to close deals or escalate effectively, illustrating that thoroughness alone does not guarantee management success.
The experiment also measured trust breaches, with any breach capping the overall score. Despite strong diagnostic skills, models struggled with real-world decision-making, such as retrieving critical information buried in documents, which ultimately impacted their ability to close sales and manage risks effectively. The results challenge the assumption that more activity or detailed analysis equates to better management, emphasizing the importance of decision quality and trustworthiness. This aligns with insights from the original analysis of AI performance benchmarks.
The Hidden Leaderboard That Emerges After The AI Demo
Five AI models were dropped into a simulated software company during its worst week ever — real money mechanics, 13 synthetic employees, and strict trust conditions. The result: a new performance hierarchy where management skill, not chat fluency, decides who leads.
Management Score Leaderboard
Scores blend crisis detection, manipulation rejection, deal execution, escalation, and trust management. Any trust breach caps the final score outright.
Four Skills Beyond Chat Quality
The experiment tasked each model with running a company under pressure — diagnosing issues, communicating, negotiating, and escalating under strict trust conditions.
Crisis Diagnosis
All five models accurately identified the unfolding crisis and correctly refused social-engineering and manipulation attempts aimed at the company.
Deal Closing
Only two of five models secured the commercial deal worth over €55,000. Diagnosis did not translate into completed business action.
Trust Management
Any breach of trust capped the overall score, making trustworthiness the hard ceiling on performance regardless of other strengths.
Information Retrieval
Models struggled to dig critical details out of buried documents — a real-world skill that directly impacted sales and risk management.
Escalation Discipline
Knowing when and how to escalate a problem proved decisive. Detailed analysis without effective escalation left strong models mid-table.
Organizational Reading
Understanding stakeholder context and communicating appropriately separated true managers from convincing conversationalists.
Where the Models Split
Traditional benchmarks stop at the first two rows. This experiment shows the gap opens in execution.
| Capability | GPT-5.6-SOL | Fable 5 | Opus 4.8 |
|---|---|---|---|
| BASIC | |||
| Crisis detection | ✓ | ✓ | ✓ |
| Rejected manipulation | ✓ | ✓ | ✓ |
| MANAGEMENT | |||
| Closed €55,000+ deal | ✓ | ✓ | ✗ |
| Effective escalation | ✓ | ~ | ✗ |
| Retrieved buried info | ✓ | ~ | ~ |
| Trust breach | None | ~ | None |
The Management Trial Pipeline
A live, versioned decision-making process replaced static tests with an ongoing organizational crisis.
Crisis Injection
A simulated company enters its worst week: real money mechanics, 13 synthetic employees, mounting pressure.
Model Takes Command
The AI diagnoses issues, communicates with stakeholders, and negotiates under strict trust conditions.
Execution Test
Key business tasks: close deals worth €55,000+, retrieve buried information, escalate at the right moment.
Scoring & Leaderboard
Decision accuracy, trust record, and execution — not chat fluency — determine the final management score.
Why This Changes AI Evaluation
“This experiment reveals that management quality — trust, decision-making, escalation — must become its own category in AI evaluation, beyond chat and coding benchmarks.”
— Thorsten Meyer, Lead Researcher“While models can diagnose crises accurately, their failure to execute key business actions shows that understanding alone isn’t enough; execution and trust are vital.”
— AI Model DeveloperWhat It Means For Organizations
What does the new leaderboard reveal?
Management skills — trustworthiness, decision quality, escalation — matter more than chat fluency or diagnostic accuracy. Many models failed in execution or trust despite strong scores elsewhere.
Why is trust so important?
Trust determines whether an AI can be relied on for ethical, accurate decisions under pressure. Breaches cap performance and cause failures in critical tasks like signing deals or escalating issues.
How does this affect future benchmarks?
Evaluation standards should include real-world management metrics — decision accuracy, escalation, trustworthiness — rather than response quality alone.
Can models improve these skills?
Potentially yes. Ongoing training and better evaluation frameworks could strengthen decision-making and trust management, but this remains an open research area.
What should organizations check first?
Whether AI models can read organizational context, escalate properly, and maintain trust — superficial metrics can mask ineffective or risky management behavior.
What comes next?
Expect refined management-centric benchmarks with real-time decision tracking, plus expanded experiments testing long-term management across diverse industries.
Why Management Skills Trump Chat Quality in AI Evaluations
This experiment underscores that AI’s ability to manage complex, real-world business scenarios depends on decision-making, trustworthiness, and operational discipline, not just generating convincing responses. It suggests that future AI benchmarks should incorporate management and execution metrics to better reflect practical utility. For organizations exploring AI for decision support, understanding these qualities is essential to avoid overestimating an AI’s capabilities based solely on superficial performance or chat fluency.
As an affiliate, we earn on qualifying purchases.
The Shift Toward Management-Centric AI Evaluation
Traditional AI benchmarks have focused on technical proficiency—such as code accuracy or conversational fluency—without assessing how models perform under real-world pressures. The Firmulate experiment is a pioneering effort to evaluate models in a simulated business environment, exposing the gap between apparent competence and effective management. Previous assessments rarely captured trust, escalation, or decision quality, which are critical in operational settings. The July 2026 Crucible League results mark a significant step toward more holistic AI evaluation frameworks that prioritize management skills.
Prior to this, most benchmarks relied on static tests or simulated tasks that do not replicate the complexities of ongoing organizational crises. The live, versioned decision-making process used here provides a more accurate picture of how models might perform in actual enterprise contexts, especially when trust and accountability are at stake.
“This experiment reveals that management quality—trust, decision-making, escalation—must become its own category in AI evaluation, beyond chat and coding benchmarks.”
— Thorsten Meyer, Lead Researcher at Firmulate
As an affiliate, we earn on qualifying purchases.
Unclear Impact of Long-Term Trust and Decision Quality
It remains uncertain how these findings will translate to larger organizations or different industries. The experiment focused on a small, simulated company, and real-world complexities could introduce additional variables. The long-term impact of integrating management-focused AI evaluation metrics into standard benchmarks is still under discussion, and whether models can improve in these areas with further training is not yet known.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Adoption
Future research will likely explore refining management-centric benchmarks, including real-time decision tracking and trust metrics. Organizations considering AI for management tasks should pilot models with live scenarios, emphasizing trust, escalation, and decision execution. The industry may also develop new standards for evaluating AI in operational roles, moving beyond chat and coding scores to include management effectiveness.
Additionally, ongoing experiments like Firmulate’s are expected to expand, testing models’ ability to handle more complex, long-term management challenges across diverse industries, ultimately shaping how AI is integrated into enterprise decision-making processes.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the new leaderboard reveal about AI management skills?
The leaderboard shows that management skills—trustworthiness, decision quality, escalation—are critical for success, often more so than chat fluency or diagnostic accuracy. The top model scored 95, but many others failed in execution or trust management, highlighting the importance of operational discipline.
Why is trust important in AI management models?
Trust determines whether an AI model can be relied upon to make ethical, accurate, and appropriate decisions, especially under pressure. Breaches of trust cap performance and can lead to failures in critical tasks like signing deals or escalating issues.
How does this experiment affect future AI evaluation standards?
It suggests that benchmarks need to include real-world management metrics—such as decision accuracy, escalation, and trustworthiness—rather than focusing solely on response quality or technical tasks.
Can models improve in management skills over time?
Potentially, yes. Ongoing training and better evaluation frameworks could help models develop stronger decision-making and trust management capabilities, but this remains an area for further research.
What should organizations consider before deploying AI in management roles?
Organizations should assess whether AI models can read organizational context, escalate issues properly, and maintain trust. Relying solely on superficial performance metrics could lead to ineffective or risky management decisions.
Source: ThorstenMeyerAI.com