🔍 Read the full analysis: Mistral Large 4: A Strong Model Beyond The US And China, But Agent Use Is Tricky on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, up from 9 for Mistral Large 3, making it a substantial step forward for the French AI company. The score still trails current US and Chinese leaders, while its price, output volume and reported hallucinations raise questions about using it for long-running agent tasks.
Mistral AI has released Large 4, a research-preview model that scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, marking a steep rise from the company’s previous flagship but leaving it behind leading US and Chinese systems. The result makes Mistral a stronger European contender, while the benchmark and reported hands-on testing raise concerns about its cost and reliability on multi-step agent work.
Artificial Analysis’ current index places Large 4 below major US models, including Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7. It also trails several Chinese models: GLM-5.3 scored 44.8, Kimi K3 scored 43.6, GLM-5.3-Flash 41.8 and DeepSeek V4.1 Flash 39.5. Large 4’s score is close to OpenAI’s smaller GPT-6 Luna, listed at about 38. The figures are from the same index version, according to the source, allowing a like-for-like comparison.
The same dataset shows a sharp improvement within Mistral’s own lineup: Large 3 scored 9 and Medium 3.5 scored 14 on that index version. Large 4 is a one-trillion-parameter model with 49 billion active parameters, accepts text and images, and has a 512,000-token context window. Mistral says reinforcement learning is continuing, so the score could change. The source reports API pricing of $1.36 per million input tokens and $4.18 per million output tokens, with discounted cached input and an introductory offer.
Large 4 is currently a proprietary research preview available through Mistral’s API. The source says Mistral has promised to release model weights at the end of October, but that timing is a company commitment rather than a completed release. It also reports that the model’s licence has not been published. Until the weights and licence details are available, buyers cannot evaluate it as an open-weight option on those terms.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The trade-offs for agent deployments
The benchmark matters because Artificial Analysis’ Intelligence Index includes work designed to assess agentic tasks, such as knowledge work, software workflows and coding. It is not a measure of every use case, but it offers evidence relevant to systems asked to carry out multiple steps. In such workflows, an early error can affect later actions; a lower benchmark result is a reason to test carefully, not proof that a particular deployment will fail.
Cost and output volume add practical questions for organizations comparing models. The source reports that Large 4 used 200 million output tokens across the index tasks, versus a median of 81 million for comparable models, and calculates a cost of $1.13 per index task. It lists GLM-5.3-Flash at $0.25 per task with a score of 41.8, and DeepSeek V4.1 Flash at $0.27 with a score of 39.5. These are source-reported benchmark costs, not universal prices for every workload; actual bills depend on usage and pricing conditions. Still, the figures suggest that Large 4’s European origin alone may not outweigh performance and operating-cost differences for buyers.
The source author also reports seeing confident false answers in hands-on use. That is an individual observation, not an independently established benchmark finding, and the report provides no test protocol or sample size. It nevertheless points to a risk buyers should assess directly: whether the model can reliably ground answers and recover from mistakes in the tasks they plan to automate.
As an affiliate, we earn on qualifying purchases.
A European model gains ground
The launch lands amid competition among large US and Chinese AI labs, where benchmark scores, access terms and inference costs all shape adoption. Mistral’s result is notable within that market because its previous models scored considerably lower on the same index version. The move from 9 for Large 3 to 38.4 for Large 4 suggests a substantial improvement by the benchmark’s measure, though scores from one index do not capture every quality or deployment requirement.
The launch framing emphasized Large 4 as the most intelligent model from outside the United States and China, as presented in Artificial Analysis coverage cited by the source. That description depends on the comparison group. Within the index table provided, multiple US and Chinese models score higher; the claim is about geographic origin, not a top position across the full field. The source also notes that Cohere’s enterprise model is built around retrieval and tool use rather than competing directly in this frontier-model ranking, limiting its usefulness as a direct comparison.
There is an important distinction between a model being available through an API and its weights being available for others to run or inspect. Large 4’s preview status means customers can test it through Mistral now, while the promised weights release and unpublished licence leave open what independent deployment will permit.
Open questions on reliability and access
The benchmark is a snapshot, and Mistral says training work is continuing. It is not yet clear whether subsequent reinforcement learning will change Large 4’s ranking, output volume or reliability. Nor does the supplied material include enough detail about the author’s hands-on tests to establish how often the reported hallucinations occur or which tasks are affected.
The source does not specify the model’s final licence, the precise scope of the promised weights release, or whether the end-of-October date will hold. The introductory 50% discount is reported as lasting two weeks, but the material does not establish the date from which that period runs. Performance and cost on a given customer’s workload may also differ from the benchmark estimates.
Watch for weights and retesting
The next milestones are Mistral’s planned weights release at the end of October and publication of the licence terms. Those would clarify whether organizations can run the model outside Mistral’s API and under what conditions. Buyers considering the preview can compare it with alternatives on their own tasks, measuring successful completion, token use, latency and the frequency of unsupported answers rather than relying on an index score alone.
Further benchmark results may follow as Mistral continues reinforcement learning. Until updated evaluations and fuller release terms are available, Large 4’s clearest confirmed story is a major improvement over Mistral’s previous benchmark showing—not evidence that it has caught the leading US or Chinese models.
Key Questions
How did Mistral Large 4 score?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. The same source lists Mistral Large 3 at 9 on that index version.
Does Large 4 lead the US and Chinese models?
No. The rankings supplied put several US and Chinese models above it, including Claude Opus 5.5 at 57.6 and GLM-5.3 at 44.8. The launch description refers to its position among models from outside those two countries, not the overall leaderboard.
Can organizations download and run its weights now?
Not according to the source. Large 4 is available as a research preview through Mistral’s API. Mistral has promised weights for the end of October, while the licence remains unpublished in the supplied material.
Is Large 4 suitable for AI agents?
The supplied benchmark and source author’s observations raise issues to test, including cost, high output-token use and reported confident errors. They do not establish that the model will fail in every agent workflow. Organizations should evaluate it on representative tasks before relying on it for multi-step work.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
