Mistral Large 4 For AI: Strengths Beyond The US And China, Limits For Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 For AI: Strengths Beyond The US And China, Limits For Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4, released as a research preview, scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a major improvement on its predecessor. The score puts it below leading US and Chinese models, while the source analysis raises concerns about output volume, cost per task and observed hallucinations; model weights and licensing details are still pending.

Mistral AI has released Mistral Large 4 as a research preview, and the model scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The result is a steep rise from earlier Mistral models and places it among the leading models from outside the United States and China in the source comparison, but it remains below the top US and Chinese systems—an important distinction for buyers considering it for demanding AI work.

Artificial Analysis’s current index lists Large 4 at 38.4 points. In the same version, Mistral Large 3 scored 9 and Medium 3.5 scored 14, making Large 4’s improvement substantial. The source’s ranking puts it below several US models, including Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7, as well as Chinese models such as GLM-5.3 at 44.8 and Kimi K3 at 43.6. These are benchmark results, not a guarantee of performance on every workload.

Mistral describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, able to process text and images and produce text, with a 512,000-token context window. It is available through Mistral’s API as a research public preview. The source reports standard prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens and a 50% discount for the first two weeks.

The source analysis says Artificial Analysis measured Large 4 at $1.13 per Intelligence Index task and that it generated 200 million output tokens across the benchmark, compared with a median of 81 million for comparable models. The author also reports seeing confident false statements in hands-on testing. That observation is separate from the Artificial Analysis scores and should not be treated as a published benchmark result.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, with benchmark data showing a sharp improvement over its predecessor but a continued gap from leading US and Chinese models.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$85,081▼ 0.6%
Ethereum ETH$2,681▼ 1.0%
Tether USDT$1▲ 0.0%
BNB BNB$775.21▼ 0.9%
XRP XRP$1.49▼ 0.8%
USDC USDC$1▲ 0.0%
Solana SOL$119.68▼ 0.8%
TRON TRX$0.3349▼ 0.4%
Live data · CoinGecko · alternative.me (24h change)
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Benchmark Gains, Agent-Use Questions

The release matters because it shows a large reported improvement for Mistral, a European AI company seeking to compete with the largest model developers. For organisations weighing dependence on US or Chinese providers, a stronger European option may broaden the field. But the benchmark does not establish parity with the leading systems: Large 4’s score is below the highest-scoring models listed, and the source says it trails some Chinese models that are expected to have open weights.

The gap and the reported output volume may matter most for multi-step agent tasks. A model that makes errors during one step can pass them into later actions, while longer outputs can add both latency and usage charges. Artificial Analysis’s index includes agentic tasks, according to the source, but a single overall score cannot show how Large 4 will perform in every company’s workflow. Buyers would need to test relevant tasks, costs and safeguards directly.

Cost comparisons in the source also complicate the case for using Large 4 at scale. It reports that GLM-5.3-Flash costs about $0.25 per index task and scores 41.8, while DeepSeek V4.1 Flash costs about $0.27 and scores 39.5. These are comparisons using the source’s stated benchmark and pricing assumptions; actual expenses will vary with workload, prompt length, caching and provider terms. Performance, reliability and total cost all matter in a procurement decision.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Large 4 Compares

The Artificial Analysis Intelligence Index version cited by the source is v4.3.2, allowing the listed models to be compared on the same version of the index. Its scores combine evaluations that the source describes as including agentic knowledge work, real-world work tasks, software workflows and coding. This makes the result relevant to more than general question answering, though it remains a benchmark rather than a direct test of a particular deployment.

The source frames the result as the strongest model from outside the US and China, while also cautioning that this comparison says little about whether Large 4 can match those countries’ leading labs. It reports that Large 4 exceeds the earlier Mistral models and some older Chinese models, but falls behind newer Chinese entries in its open-model ranking. The comparison set and benchmark version matter: rankings can change as models and evaluations are updated.

Large 4 is currently a proprietary API preview, according to the supplied material. Mistral has said weights are expected at the end of October, but the source says the licence has not been published. Until those details are available, users cannot judge the terms for running or adapting the weights from the information provided.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— ThorstenMeyerAI.com source author

Amazon

large language model for developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Reliability

Several points remain unresolved in the supplied material. Mistral has promised model weights for the end of October, but their release date, licence and usage restrictions are not confirmed here. The preview is currently accessible through Mistral’s API, and the source says the weights have not yet been released.

The benchmark may also change: Mistral says reinforcement learning is continuing, and the source does not provide a later score or a final evaluation. Nor does it establish how often the author’s reported hallucinations occur across a representative set of prompts. The reported cost and token totals reflect the benchmark setup; the source does not provide enough detail to calculate costs for every production use case.

Amazon

AI text and image processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Expected Weights and Further Testing

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. The licence and the actual timing will determine how organisations can run or adapt the model outside Mistral’s API. Mistral’s continuing reinforcement learning may also lead to updated performance figures, though no schedule for revised benchmark results is given.

For now, teams considering the preview can compare it with alternatives on their own tasks, tracking accuracy, hallucination rates, latency and total token costs rather than relying only on the index score. The source material does not establish that Large 4 is unsuitable for agents across the board; it does give reasons to test carefully before assigning it long-running or consequential workflows.

Amazon

AI token usage calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral announce?

Mistral released Mistral Large 4 as a research public preview through its API. The model is described as multimodal for text and images, with a 512,000-token context window.

How did Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the supplied source. The source reports scores of 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version.

Is Mistral Large 4 the leading model overall?

No. The rankings cited place it below several US and Chinese models. The source describes it as leading among models from outside those two countries, a narrower comparison.

Are the model weights available?

Not yet, according to the supplied material. Mistral has promised weights for the end of October; the licence has not been published in the source.

Does the report prove Large 4 is unreliable for agents?

No. The source raises concerns based on benchmark performance, output volume and the author’s reported hands-on observations, but it does not establish performance across all agent tasks. Organisations should test the model on their own workflows and account for the risk of errors and cost variation.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What’s New In MiniMax H3? Sound Features And The Meaning Of ‘Open’ In AI

MiniMax launched H3 on July 31, 2026, featuring joint audio-visual generation and an ‘open’ base model, with some limitations on openness and licensing.

Meta’s Reality Labs Losses Hit $17.7B, but Zuckerberg Calls 2024 a Pivotal Year for the Metaverse

In light of Meta’s staggering $17.7 billion losses, can Zuckerberg’s vision for the metaverse truly turn the tide in 2024? Find out more.

10 Best NAS Options For Home And Small Business: 2027 Guide

A 2027 NAS roundup compares named Synology, UGREEN and Buffalo options, with buying guidance on bays, drives, setup and backup limits.

The New Personal Agent Layer

A new personal agent layer is emerging, enabling persistent, action-oriented AI agents that operate across private and professional environments, raising questions of ownership and safety.