AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Is Mistral Large 4 A Top AI Model Outside The US And China? on ThorstenMeyerAI.com

TL;DR

Mistral released Large 4 as a research preview, and Artificial Analysis’ Intelligence Index v4.3.2 gives it a score of 38.4. The score places it below current leading US and Chinese models, though it is a substantial improvement over Mistral’s previous scores. Its pricing, agent-task performance and unpublished license remain relevant questions for potential users.

Mistral AI has released Mistral Large 4 as a Research Public Preview, with Artificial Analysis’ Intelligence Index v4.3.2 scoring the model at 38.4. That score trails the leading US and Chinese models in the cited comparison, although it marks a sharp rise from the company’s previous-generation scores and puts Mistral among the stronger labs outside those two countries.

Artificial Analysis’ table places Large 4 below six listed US frontier models, whose scores range from 51.8 to 57.6. The leading model in that table, Claude Opus 5.5, scores 57.6. The highest listed Chinese model, GLM-5.3, scores 44.8; other Chinese models in the comparison also exceed Large 4, including Kimi K3 at 43.6, GLM-5.3-Flash at 41.8 and DeepSeek V4.1 Flash at 39.5. These figures describe performance on the stated index, not every possible use of the models.

The release is nevertheless a marked improvement for Mistral on the same index version: Large 3 scored 9, while Medium 3.5 scored 14. Artificial Analysis’ result makes Large 4 a notable European entry in the field, but does not place it at the level of the current US or Chinese leaders. The source report describes it as the strongest model from outside the US and China; that framing depends on the geographic comparison set and should not be confused with a top global ranking.

Mistral describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, text-and-image input, text output and a 512,000-token context window. It is available through Mistral’s API as a research preview. The source gives standard prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million; it also reports a 50% discount for the first two weeks. The source says Mistral’s reinforcement learning is continuing, so benchmark results could change.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral has released Large 4 as a research preview, with independent benchmark data placing it behind current leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has „essentially closed the gap.“ It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. „Most intelligent outside the US and China“ is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Benchmark Results and Buyer Trade-Offs

For companies choosing models, the score is one data point in a decision that also involves price, reliability, latency and task fit. The source report says Artificial Analysis’ index includes agent-oriented evaluations such as knowledge work, software workflows and coding tasks. That makes the result relevant to buyers considering multi-step automated work, but it does not establish how Large 4 will perform in a particular company’s own systems or on every task.

The report also says Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. It reports an estimated cost of $1.13 per index task for Large 4, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; both Chinese models scored higher on the index. These are source-reported benchmark costs, not a universal estimate for a customer’s workload. Actual bills depend on prompts, outputs, caching, task mix and provider pricing.

That trade-off gives the release significance beyond its national origin. A European model can offer buyers an additional provider and a potential alternative for particular deployment needs. But the benchmark figures alone do not show that Large 4 is the best choice for privacy, compliance, cost or performance in any specific setting. The report’s claim that the model can hallucinate confidently is based on the author’s hands-on experience, not a published Artificial Analysis measurement cited here, and should be treated accordingly.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Large 4 Compares

The source report bases its ranking on Artificial Analysis Intelligence Index v4.3.2, presenting scores for models from US, Chinese and European labs. In that comparison, Large 4’s 38.4 is near OpenAI’s small GPT-6 Luna model, listed at about 38, and above DeepSeek V4 Pro at 36.0 and GLM-5.2 at 33.7. Rankings can vary by benchmark, version and evaluation conditions, so the figures should be read as a snapshot of this index rather than a final judgment on model quality.

Large 4 is not yet an open-weights release. The source says Mistral has promised to release its weights at the end of October, but gives no year for that date. Until that release, users access the preview through the company’s API, and the model’s license is unpublished, according to the source. The distinction matters to developers weighing API access against self-hosting, model modification or license terms.

The report also compares Large 4 with Cohere, but says the comparison is limited because Cohere’s enterprise model is aimed at retrieval and tool use rather than frontier reasoning. That is a useful reminder that different systems may serve different purposes; a single benchmark ordering cannot substitute for evaluating a model against a real use case.

Amazon

text and image AI input devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Release Terms and Real-World Performance

Several points remain unsettled. Mistral has not yet released the weights, and the source says the license is unpublished. The promised end-of-October release is not accompanied by a year in the supplied material, so its exact timing cannot be confirmed here. Mistral’s continuing reinforcement learning also means the preview may not represent the model’s final performance.

The benchmark score does not settle how well Large 4 will perform on an individual buyer’s work. The source report’s concerns about verbosity and confident hallucinations include the author’s own observations; it supplies no detailed testing protocol for those observations. Nor does the supplied material include a complete comparison of reliability, latency, safety, privacy protections or performance across languages. Buyers would need task-specific tests and current provider terms before drawing conclusions.

The source’s cost comparisons are based on Artificial Analysis index tasks and listed prices. They do not show the total cost of a production workflow, which can depend on input and output volume, cached tokens and other implementation choices. The reported introductory discount is limited to two weeks, but the source does not state the dates of that offer.

Amazon

large language model tokens calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Updated Scores

The next milestones are Mistral’s promised weight release and publication of the license terms. Those details will help developers determine whether they can run the model themselves, how they may adapt it, and what restrictions apply. Mistral’s API preview remains the stated access route in the meantime.

Artificial Analysis scores may also change as Mistral continues reinforcement learning and the evaluator updates its comparisons. For now, the reported index offers a snapshot: Large 4 has improved substantially over Mistral’s earlier scores but remains below the leading US and Chinese systems listed. Prospective users will need to test it on their own tasks and compare the full cost and terms with alternatives before making deployment decisions.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

Artificial Analysis Intelligence Index v4.3.2 scored it at 38.4 in the source report. That places it below the leading US and Chinese models listed, but well above Mistral Large 3’s score of 9 on the same index version.

Is Mistral Large 4 the world’s leading AI model?

No. The supplied benchmark comparison places several US and Chinese models above it. The description of Large 4 as the strongest model outside those countries is a geographic comparison, not a claim that it leads globally.

Can developers download Large 4’s weights now?

Not according to the source report. It says Large 4 is available through Mistral’s API as a research preview and that Mistral has promised a weight release for the end of October. The supplied material does not specify the year or publish the license.

How much does Mistral Large 4 cost to use?

The source lists standard API prices of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks; the dates of that offer are not specified in the supplied material.

Should companies use Large 4 for AI agents?

The benchmark and cost figures do not answer that for every company. The source report raises concerns about agent-task performance, token use and hallucinations, but some of those concerns are the author’s observations rather than benchmark measurements. Organizations should test the model on their own workflows and review current pricing, reliability and license terms.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Should You Spend Money On Fable, Opus 5.5, Astra, Sol, Or Luna AI Models?

Analysis of AI models Fable, Opus 5.5, Astra, Sol, and Luna, focusing on performance, cost, and suitability for various tasks amid new benchmark data.

Why The Hidden AI Ban Matters In China’s Optical-Transceiver Innovation

Analysis of the proposed FCC measure to restrict Chinese optical transceivers highlights strategic, industrial, and security implications for AI infrastructure.

Outcome-First Decisions: The Friction Is the Feature

A new decision-making approach prioritizes testing and evidence, reducing wasted effort and building calibrated judgment over time.

The Bottleneck Moved: Inside Anthropic’s Expansion of Project Glasswing

Anthropic is extending its cybersecurity initiative, Project Glasswing, from 50 to 150 partners, shifting focus from vulnerability detection to patching and fixing threats.