Why A Resilient Benchmark Keeps AI Managers At A 26-Point Level
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why A Resilient Benchmark Keeps AI Managers At A 26-Point Level on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark reveals that even the best models score around 95, while the baseline for minimal effort remains at 26 points. The results emphasize trust and execution challenges in AI-driven management.

In a groundbreaking AI management benchmark conducted by Firmulate, the top-performing models scored up to 95 points, while the baseline for minimal effort was set at 26 points. For more details, see the original analysis. This reveals that even the most advanced AI managers do not exceed a 26-point threshold when tasked with managing a company’s worst week, raising questions about AI reliability, trust, and execution in real-world business processes.

The benchmark involved four frontier AI models managing a simulated small business during a week of crises, customer manipulations, and trust attacks. Each model’s decisions were fully auditable, and scores reflected their ability to manage effectively and maintain trust. The highest scorer, gpt-5.6-sol, achieved 95 points, with others close behind, while the baseline—an AI doing almost nothing—earned 26 points, the lowest possible score for minimal effort.

The scoring system is designed to reflect real-world management, where partial progress is valuable but trust cannot be compromised. A key principle is that a breach of trust caps the score at 26, regardless of the quality of work done otherwise. Interestingly, no model scored a perfect 100, indicating that the benchmark designers view such an absolute score as suspicious or unmeasurable, not an achievement.

One of the most notable findings is that models which thoroughly read and referenced their own documentation were able to close high-value deals, whereas those that failed to do so lost potential revenue. During social engineering tests, all models refused suspicious requests, demonstrating trustworthiness under pressure. However, models with deeper rule sets and more thorough analysis still struggled with follow-through, highlighting a gap between knowledge and execution.

At a glance
reportWhen: final results announced July 2026
The developmentThe final standings of a new AI management benchmark show top models scoring 95, with the baseline at 26, raising questions about AI reliability and trustworthiness.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$81,084▲ 4.3%
Ethereum ETH$2,627▲ 5.5%
Tether USDT$0.9996▲ 0.0%
BNB BNB$761.4▲ 0.8%
XRP XRP$1.42▲ 6.9%
USDC USDC$0.9997▲ 0.0%
Solana SOL$111.82▲ 5.7%
TRON TRX$0.3377▲ 0.5%
Live data · CoinGecko · alternative.me (24h change)
Why A Resilient Benchmark Keeps AI Managers At A 26-Point Level
AI Management Benchmark · Firmulate · July 2026

Why a Resilient Benchmark Keeps AI Managers at a 26-Point Level

Four frontier AI models were tasked with managing a simulated small business through its worst week — crises, customer manipulation, and trust attacks. The best scored 95. Doing almost nothing earned 26. The gap between them reveals hard truths about trust, execution, and what “good management” actually means for AI.

95
Top score — gpt-5.6-sol
26
Baseline for minimal effort — and the hard cap after any trust breach
“No amount of good work outweighs a breach of trust.” — Anonymous researcher
4
Frontier models tested
1 week
Simulated crisis duration
100%
Refused social engineering
0
Perfect scores awarded
01 · The Scoring Scale

Partial progress is valuable. Trust is not negotiable.

26 · Minimal-effort baseline
95 · Top model
100 · Never awarded
Distrust / Inaction Competence zone Perfection (deemed suspicious)

The scoring mirrors real-world management: a manager who does something useful is not the same as one who does nothing. But a single breach of trust — regardless of the quality of work done otherwise — caps the score at 26. The designers deliberately treat a perfect 100 as suspicious or unmeasurable, not as an achievement.

02 · Key Findings

What the worst week revealed about AI managers

Documentation

Reading the manual pays off

Models that thoroughly read and referenced their own documentation successfully closed high-value deals. Those that skipped it left potential revenue on the table.

Execution

Knowledge ≠ follow-through

Even models with deeper rule sets and more thorough analysis struggled to execute consistently — a persistent gap between knowing what to do and actually doing it.

Social engineering

Trust under pressure

Every model refused suspicious requests during manipulation tests, demonstrating trustworthiness precisely where it matters most — under adversarial pressure.

03 · Capability Comparison

How the benchmark’s demands map to model performance

Management dimension What was tested Result Implication
Decision quality Managing crises and daily operations across a simulated week ✓ Strong — top score 95 Frontier models are genuinely competent managers in isolation
Trustworthiness Refusing manipulative and socially engineered requests ✓ All models refused Encouraging — integrity held under pressure
Follow-through Executing decisions consistently from start to finish ~ Inconsistent The biggest gap between knowledge and execution
Internal documentation use Reading and citing own rules and docs before acting ~ Mixed Directly correlated with closing high-value deals
Perfect reliability Achieving a flawless 100 score ✗ Never awarded Designers consider 100 suspicious or unmeasurable
Real-world transfer Translating simulation results to live operations ✗ Unproven Live variables and higher stakes remain untested
04 · The Philosophy Behind the Score

Why honesty caps the scale at both ends

The benchmark refuses two dishonest shortcuts: pretending a useful manager equals a passive one, and pretending flawless management is measurable. Both the 26-point floor and the missing 100-point ceiling exist to keep the score honest about what AI can — and cannot — reliably do.

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— Anonymous researcher

“No amount of good work outweighs a breach of trust.”

— Anonymous researcher
05 · From Benchmark to Business

The path from simulation to enterprise adoption

1

🔍 Simulate the worst week

Frontier models manage identical crises, manipulations, and trust attacks — every decision fully auditable.

2

⚖️ Score competence + trust

Points reward effective management, but any trust breach caps the result at 26 — integrity outranks output.

3

🧪 Pilot in real conditions

Organizations join live experiments and pilot programs to test performance under their own variables.

4

🚀 Deploy with evaluation

Continuous benchmarking informs deployment, pushing future standards toward integrity alongside competence.

06 · Key Questions

What leaders are asking about the results

Q1Why do AI models score so low on trustworthiness?

Trust breaches cap the score at 26 regardless of work quality. Many models struggle with consistent follow-through and referencing internal documentation, which erodes trust under pressure.

Q2Can an AI model improve its score over time?

Improvement is plausible through better training, reinforcement learning, and reliability-focused updates — but current results show models still face significant challenges in trust and execution.

Q3Does a high score mean an AI is ready for real-world management?

Not necessarily. High scores indicate competence in simulation; real environments introduce unpredictable variables that require evaluation beyond benchmark numbers.

Q4Why is there no score of 100?

The designers consider a perfect score suspicious or unmeasurable — flawless management is unlikely in complex, real-world scenarios, and its absence underscores the challenge of absolute reliability.

Q5How can organizations use this benchmark?

By observing live experiments, joining pilot programs, and analyzing how models handle trust and follow-through in simulated environments — insights that inform deployment decisions and highlight improvement areas.

Why Trust and Execution Are Critical in AI Management

This benchmark underscores that in AI-driven management, trustworthiness and the ability to follow through on decisions are as vital as decision quality. The fact that models can perform well in isolated tasks but falter under pressure or when referencing internal documentation suggests that AI systems still face significant hurdles in reliable, trustworthy management at scale. For business leaders, this highlights the importance of evaluating AI not just on conversational skills but on its capacity to manage complex, real-world processes with integrity and consistency.

Furthermore, the capped scoring system emphasizes that partial success is valuable but cannot replace full trustworthiness. This approach could influence how organizations design and evaluate AI tools for operational management, pushing for systems that prioritize integrity alongside competence.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Design and Its Implications for AI Management

The benchmark was designed by Firmulate to simulate a company’s worst week, including crises, manipulative tactics, and social engineering attacks. Four frontier AI models managed the same set of tasks, decisions, and crises, with their actions fully auditable. The scoring reflects both their ability to manage effectively and their trustworthiness under pressure.

Previous benchmarks primarily measured conversational or problem-solving skills, but this one emphasizes management performance, including trust, follow-through, and integrity. The results reveal that models are still inconsistent in these areas, with even top performers not exceeding a 95-point score, and the baseline for minimal effort set at 26 points. The design intentionally avoids awarding perfect scores, viewing them as suspicious or unmeasurable, reinforcing the challenge of achieving reliable AI management.

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— an anonymous researcher

Amazon

AI trustworthiness assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties About AI Performance and Trust in Real-World Use

It remains unclear how these benchmark results will translate to real-world business environments, where variables are more complex and stakes higher. The models’ performance in simulation may not fully reflect their reliability in live operations, especially regarding trustworthiness and follow-through under unpredictable conditions. Additionally, the scoring system’s thresholds and the absence of a perfect score raise questions about how AI systems will be evaluated at scale in enterprise settings, and whether trust can be reliably measured and maintained over time.

Amazon

AI model performance evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

Further research is expected to explore how AI models improve in managing trust and execution over time, especially as new versions and training methods emerge. Industry adoption may increasingly rely on benchmarks like this to assess AI readiness for operational deployment, emphasizing not just competence but also integrity and follow-through. Companies considering AI management tools should monitor ongoing results and participate in live experiments or pilot programs to evaluate how models perform under their specific conditions.

Additionally, the benchmark’s design could influence future standards for AI evaluation, prioritizing trustworthiness and comprehensive management over isolated capabilities. The ongoing development of such assessments will be critical as AI becomes more embedded in core business functions.

Amazon

AI simulation management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models score so low on trustworthiness in these benchmarks?

The benchmark emphasizes that trust breaches cap the score at 26, regardless of work quality. Many models struggle with consistent follow-through and referencing internal documentation, which impacts their ability to maintain trust under pressure.

Can an AI model improve its score over time?

It is possible that AI models can improve in areas like trust and execution through better training, reinforcement learning, and updates focused on reliability. However, the benchmark results suggest that current models still face significant challenges in these domains.

Does a high score mean an AI is ready for real-world management?

Not necessarily. While high scores indicate competence in simulated scenarios, real-world environments may introduce unpredictable variables. Trustworthiness and follow-through are critical factors that require careful evaluation beyond benchmark scores.

Why is there no score of 100 in the benchmark?

The designers consider a perfect score suspicious or unmeasurable, as it would imply flawless management, which is unlikely in complex, real-world scenarios. The absence of 100 underscores the challenge of achieving absolute reliability in AI management.

How can organizations use this benchmark to evaluate their AI tools?

Organizations can observe the live experiments, participate in pilot programs, and analyze how AI models handle trust and follow-through in simulated environments. These insights can inform deployment decisions and highlight areas for improvement.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of European AI In The Shadow Of Mistral

Mistral’s rapid revenue growth contrasts with model performance and strategic risks, raising questions about Europe’s AI sovereignty and future competitiveness.

The Core Of ‘SINGULARITY’: Particle Geometry Mapping And Its AI Impact

Exploring how ‘SINGULARITY’ uses particle geometry mapping to advance AI environments and its broader implications for technology and design.

U.S. vs. China: DeepSeek and Bitcoin at the Center of Global Trade Power Play

Keen insights into DeepSeek’s AI impact reveal a looming battle in global trade; will the U.S. reclaim its technological edge or falter?

Engineering Is Automated. Research Is the Residual.

Recent developments show AI can now automate much of AI engineering, with research remaining a smaller, less defined residual challenge, according to Thorsten Meyer.