🔍 Read the full analysis: Mistral Large 4 Still Faces A Gap To The AI Frontier on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral released Mistral Large 4 as an API preview on October 6, but its Artificial Analysis Intelligence Index score of 38 trails leading US models and several Chinese competitors. The model’s weights are scheduled for release later in October; benchmark results and one reviewer’s hallucination experience do not establish how it will perform on every real-world task.
Mistral AI launched Mistral Large 4 in public API preview on October 6, but an early score from Artificial Analysis places it behind leading US models and several Chinese alternatives. The result matters to developers weighing the model for complex work: the preview scored 38 on the Intelligence Index, while the company’s scheduled release of model weights has not yet happened.
Mistral describes Large 4 as its largest model to date, a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. The preview accepts text and images. Mistral said it trained the model on its own European infrastructure and is continuing to improve it. Its weights are scheduled to be released later in October, so the model was not yet publicly downloadable at the time of the source report.
In Artificial Analysis’s benchmark snapshot dated October 7, Mistral Large 4 Preview scored 38 on the Intelligence Index. That matched OpenAI’s GPT-6 Luna at maximum reasoning effort and was one point below DeepSeek V4.1 Flash at maximum effort. Higher scores on the index indicate stronger aggregate performance across its benchmark suite, but the comparison includes different reasoning settings and is not an evaluation under identical compute budgets.
The same snapshot listed Anthropic’s Claude Opus 5.5 at 58, Google’s Gemini 4 Argon at 53 and OpenAI’s GPT-6.1 Sol at 52. China’s Z.ai GLM-5.3 scored 45 and Moonshot AI’s Kimi K3 scored 44. Cohere Command A+, developed in Canada, scored 13. These are benchmark index points, not percentages, and developer locations do not indicate where an individual API request is processed.
Model briefing · October 7, 2026
Mistral Large 4 Still Faces a Gap to the AI Frontier
Mistral’s new API preview scored 38 on Artificial Analysis’s Intelligence Index. That puts it behind several leading US models and named Chinese competitors in this dated snapshot, while leaving real-world performance open to testing.
01 · Benchmark snapshot
A clear gap in this index
In the October 7 comparison, Large 4 Preview scored below several leading models. The list combines different reasoning settings, so it is not a like-for-like test under identical compute budgets.
Selected model scores
Artificial Analysis Intelligence Index · points, not percentages
*At maximum reasoning effort. DeepSeek V4.1 Flash scored 39 at maximum effort. Scale shown from 0–58.
How to read the result
The index helps compare broad benchmark performance. It cannot settle how a model will perform on a particular workflow.
What the snapshot supports: the preview trails the listed leading US models and the named Chinese alternatives GLM-5.3 and Kimi K3. It scores above Cohere Command A+ in the same comparison.
What it does not establish: success or failure on a customer’s coding, research, or agentic task. Different reasoning configurations also limit direct comparison.
02 · What launched
A capable preview, with more to come
Mistral describes Large 4 as its largest model to date: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images.
API preview
Mistral announced public API access on October 6, 2026. The source report says the model weights were not yet downloadable.
MoE at scale
The stated configuration has 1T total parameters and 49B active parameters, with text and image input support.
Still improving
Mistral says training used its own European infrastructure and that it continues to improve the model.
03 · Decision path
Turn the benchmark into a workload test
For teams considering Large 4, compare the API preview on their own tasks. Track outcomes that matter in production, then reassess when new benchmark data and weight-release details arrive.
Use representative coding, research, or business tasks.
Compare the preview against alternatives on the same workload.
Track accuracy, tool use, completion, supervision time, and cost.
Update the decision as the model and evaluation results change.
04 · Evidence limits
What remains unknown
The available material offers an early index score and one uncontrolled reviewer account. Neither establishes broad, comparative reliability across real customer workflows.
Workload fit
Index points do not predict results for every coding, research, or multi-step workflow. Specialized strengths need workload-specific evaluation.
Hallucination rate
One reviewer reported hallucinations during personal use, while noting the experience was uncontrolled. No controlled frequency study is provided.
Future details
The source gives no detailed cost table or methodology. Weight-release timing and contents, and future performance, still need confirmation.
Context capacity is not reliability: the reported capacity of about 512,000 tokens describes how much material a request may contain. It does not show accurate reasoning across that material or dependable completion of a long workflow.
05 · Quick answers
At a glance
The launch status and benchmark snapshot point to a preview-stage evaluation. The next useful evidence is repeatable performance on the tasks teams actually need done.
A public API preview on October 6, 2026.
38 on the October 7 Intelligence Index snapshot.
No. A specific workflow needs its own evaluation.
Not in the cited report; release was scheduled for later in October.
Test accuracy, tools, completion, supervision, and cost on their tasks.
Benchmark Results Shape Model Choices
For teams selecting a model for software development, research or other multi-step work, benchmark results can help narrow the field, but they do not settle performance on a specific task. A lower aggregate score may be a reason to test alternatives; it is not proof that a model will fail at a particular job. Mistral’s score puts the preview behind several competitors in this snapshot, a relevant consideration for organizations comparing providers.
The source report argues against selecting the preview for demanding agentic work, where a model plans, uses tools and carries decisions across multiple steps. That is the author’s judgment, not a finding established by the index. Artificial Analysis reports a context capacity of about 512,000 tokens, but capacity measures how much material a request can contain; it does not show that a model will reason accurately across that material or reliably finish a long workflow.
Mistral has promoted strengths in agentic coding and specialized professional tasks. Those claims call for evaluation on the workloads in question. One reviewer also reported encountering hallucinations during personal use, while explicitly describing that experience as uncontrolled and not a comparative study. For potential customers, the practical question is whether the model’s accuracy and supervision needs fit their tasks, not whether a single score or anecdote predicts every result.
As an affiliate, we earn on qualifying purchases.
A Preview, Not the Weight Release
The timing and status of the release limit what can yet be concluded. Mistral announced an API preview on October 6; the weights were scheduled for later in the month. The release therefore represents access to a preview version, rather than the completed public-weight release some readers might expect from coverage of an open-weight model.
The benchmark comparison is a dated snapshot, not a permanent ranking. Scores may change as models and evaluations are updated. The report also says the named reasoning configurations differ, so the results should not be treated as a controlled comparison with matching compute. Artificial Analysis’s index is one measure of aggregate benchmark performance; it does not directly measure reliability on every coding, research or business workflow.
The comparison supports a narrower conclusion than saying all competitors are ahead. Mistral’s score exceeds Cohere Command A+ in the cited snapshot, while remaining below the listed leading US models and the named Chinese alternatives GLM-5.3 and Kimi K3. The source report separately says DeepSeek V4.1 Flash has approximately comparable benchmark intelligence at a lower measured cost per task, but the provided material does not include the underlying cost figures or measurement details.
“The model was trained on its own infrastructure in Europe and continues to improve.”
— Mistral, as described in its announcement
Performance Beyond the Index
The available material does not establish how Large 4 performs across a broad set of real customer workflows, or whether it is more or less reliable than competitors in controlled, like-for-like testing. The Intelligence Index score is not a direct measure of long-task reliability, and the source provides no controlled hallucination study. The reviewer’s account of hallucinations is personal experience, not evidence of their frequency across users or tasks.
It is also unclear how the model’s performance, pricing and measured cost per task may change as Mistral continues development and releases the weights. The supplied material gives no detailed cost table, evaluation methodology for that cost comparison, or results from the future weight release. The benchmark scores may also shift over time, and the reported reasoning settings do not share an identical compute budget.
Weight Release and Further Testing
Mistral scheduled the model weights for release later in October 2026. That release would give developers a further opportunity to evaluate the model, but its timing and contents should be confirmed when announced. Mistral has also said it is continuing to improve the preview.
For now, teams considering Large 4 can compare its API preview against alternatives on their own tasks, tracking accuracy, tool use, completion rates, supervision time and cost. New benchmark results and details about the weights could change the picture; until then, the October 7 index is a dated snapshot, and claims about specialized strengths remain matters for workload-specific testing.
Key Questions
What did Mistral announce?
Mistral announced a public API preview of Mistral Large 4 on October 6, 2026. The model accepts text and images; its weights were scheduled for release later in October.
How did Mistral Large 4 score?
Artificial Analysis gave Mistral Large 4 Preview an Intelligence Index score of 38 in a snapshot dated October 7, 2026. The index is an aggregate benchmark measure, not a prediction of success on every task.
Does the score prove the model cannot handle agentic work?
No. The source report’s recommendation against using it for demanding agentic work is the author’s judgment. The benchmark result does not establish whether the model will succeed or fail on a specific workflow.
Are Mistral Large 4’s weights available?
Not according to the source material, which says the weights were scheduled for later in October. The report describes the current release as an API preview.
What remains unknown about the model?
Controlled comparisons of reliability, performance across customer workloads, and detailed cost measurements are not provided. Results may also change as Mistral updates the model and releases its weights.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
