The Short Answer
Mistral Large 4 is a 1.05T-parameter mixture-of-experts model with 52B active, released on 6 October 2026 as an API, with weights promised by month's end. Self-hosting it needs a full 8-GPU node; Aleph Alpha's 78B Kolibri fits one H200-class GPU. On ibl.ai you own all the code and the data, model-agnostic, so moving a workload between them is configuration.
Two European open-weight releases three days apart make a tempting headline about the model layer commoditizing. The more useful reading is narrower: they sit at opposite ends of the hardware range, and one of them is not downloadable yet.
What is Mistral Large 4 (Le Chonk), and is it open-weight yet?
Mistral Large 4 is Mistral AI's largest model, launched as a public preview on 6 October 2026. Mistral's announcement calls it "Unofficially ML4, very officially: le Chonk."
Mistral's model page lists 52B active parameters, 1.05T total, a 1.6B vision encoder and a 1M-token context. It is a natively multimodal mixture-of-experts.
It is open-weight by commitment, not yet in fact. Mistral says "We will release the weights by the end of the month," and is red-teaming the model meanwhile with cybersecurity leaders, vetted partners and state authorities.
The announcement does not state the licence the weights will ship under. That is the first thing to read when they arrive, because it decides whether "open-weight" means commercial self-hosting without terms.
Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters and serves the preview there. It claims the model significantly outperforms "any open-weight model developed in the US or Europe."
That is Mistral's own benchmark claim, and it is scoped: US and European open models, not every open model. Treat it as a vendor claim until an evaluation on your own workloads agrees.
How much GPU memory does it take to self-host Mistral Large 4?
Self-hosting Mistral Large 4 takes roughly a full 8-GPU node, because every one of its 1.05T parameters has to sit in memory even though only 52B are active per token.
The arithmetic is simple. At one byte per parameter (FP8), 1.05T parameters is about 1,050 GB of weights. At two bytes (BF16) it is about 2,100 GB. Whatever memory is left goes to the context cache and batching.
NVIDIA lists the DGX B200 at 1,440 GB of GPU memory across 8 GPUs, and the H200 at 141 GB per GPU. Here is how both models fit:
| Model | Total / active | Weights at FP8 | Weights at BF16 | Smallest fit at FP8 |
|---|---|---|---|---|
| Mistral Large 4 | 1.05T / 52B | ~1,050 GB | ~2,100 GB | One DGX B200 node (1,440 GB), ~390 GB headroom |
| Aleph Alpha Kolibri | 78.1B / 3.46B | ~78 GB | ~156 GB | One H200 (141 GB), ~63 GB headroom |
An 8-GPU H200 node holds 1,128 GB, which fits Mistral Large 4 at FP8 with under 80 GB to spare. That is enough to load it and not much else, so long-context serving pushes toward B200-class memory or quantization below 8 bits.
The figures are weights only, before runtime overhead, and assume the published parameter counts. They are a sizing floor, not a deployment spec. The same exercise for 2-trillion-parameter models is in The Open-Weight Tipping Point.
Why do mixture-of-experts models change the cost of self-hosting?
Mixture-of-experts models split the cost of self-hosting in two: total parameters set how much memory you buy, and active parameters set how much compute each token uses.
Both European releases are sparse in nearly the same proportion. Mistral Large 4 activates about 5% of its parameters per token (52B of 1.05T). Kolibri, per its model card, activates about 4.4% (3.46B of 78.1B).
So per-token compute for Mistral Large 4 is roughly 15 times Kolibri's (52B against 3.46B active), while its memory bill is roughly 13 times larger (1.05T against 78.1B total). Neither model is cheap or expensive in the abstract.
What follows for infrastructure is that a mixture-of-experts model is priced by its memory footprint first. A 52B-active model does not run on 52B worth of hardware; it runs on hardware that can hold 1.05T.
That is why the two models land in different tiers. One is a dedicated node you size, power and schedule. The other is a single large-memory GPU, which suits the sovereign, air-gapped deployments in our Kolibri and Armada analysis.
What does Mistral Large 4 cost through the API, and how does that compare with per-seat AI?
Mistral Large 4's preview API is listed at $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. Mistral's model page also shows a launch sale at half those rates.
Token pricing has a property per-seat licensing does not: the bill follows use. A per-seat licence charges the same for an employee who sends a thousand requests a month and one who sends none, so above a small team it is the wrong shape for AI.
To make the gap concrete, take an assumption, not a measurement: an employee who sends 40 requests a working day, each with 2,000 input tokens and 500 output tokens, over 22 working days.
| Line item (assumed usage, list price) | Tokens per month | Cost per user | Cost for 1,000 users |
|---|---|---|---|
| Input (880 requests Γ 2,000 tokens) | 1.76M | $2.39 | $2,394 |
| Output (880 requests Γ 500 tokens) | 0.44M | $1.84 | $1,839 |
| Total at Mistral Large 4 list price | 2.2M | $4.23 | $4,233 |
Hold that $4.23 against any per-seat quote you have, and remember the per-seat figure is charged for the users who never open the tool as well. Your own usage logs, not this assumption, are what the comparison should run on.
Self-hosting changes the shape again: once the weights are out, cost becomes the node you run, not the tokens. Whether that beats the API depends on utilization, which the price-floor analysis works through.
What should an enterprise AI infrastructure plan do with an API-first, weights-later release?
Start the evaluation on the hosted API now, and build it so the same workload can move to self-hosted weights later with nothing rewritten. Mistral Large 4's release order rewards exactly that.
Four decisions hold up whichever way the evaluation goes:
- Evaluate on your own cases during the preview. Mistral's benchmark claims are its own; an evaluation set built from your documents and tasks is the number that matters.
- Size two tiers, not one. A dedicated 8-GPU node for the heavy model and single-GPU capacity for a sparse model like Kolibri cover very different workloads at very different costs.
- Read the licence on weights day. Until the licence is published, plan for self-hosting but do not commit to it.
- Route by workload. Keep the large model for tasks that need it and send routine traffic to the cheaper tier, the pattern visible in Vercel's AI Gateway token-share data.
None of those four steps depends on which model wins. They depend on the layer above the model being able to switch.
Where does ibl.ai fit when frontier models ship as open weights?
On ibl.ai you own all the code and the data. The platform runs model-agnostic across any LLM, so adding Mistral Large 4 through its API today and pointing the same agents at self-hosted weights later is a configuration change.
The platform layer, not the model, carries what has to persist across that move: agent definitions, the audit log, access controls, evaluation sets and data connections. They stay on infrastructure you administer.
ibl.ai deploys anywhere, from your own cloud to on-premise or fully air-gapped, with no per-seat pricing, on Agentic OS. 1.6M+ users across 400+ organizations run it this way, including NVIDIA, MIT, and Syracuse University.
The policy side of the same shift, including how U.S. rules treat open-weight models, is in How Washington Made Sovereign AI the Path of Least Resistance.
Want to run Mistral Large 4 and Kolibri on infrastructure you own?
We deploy the platform as source code you keep, sized to the model tiers you actually need. Book a 30-minute demo or talk to the ibl.ai team.
Sources: Mistral Large 4's release, parameter counts, context, pricing, training hardware, red-teaming and weights timing from Mistral's announcement and model page; Kolibri's parameters, licence and context from its model card; GPU memory from NVIDIA's DGX B200 and H200 pages. Memory and cost figures are our arithmetic on those published numbers.