---
title: "Mistral Large 4 (Le Chonk): The Infrastructure Math"
slug: "mistral-large-4-le-chonk-open-weight-infrastructure-strategy"
author: "ibl.ai Engineering"
date: "2026-10-09 14:00:00"
category: "Premium"
topics: "Mistral Large 4, Le Chonk, open-weight models, mixture of experts, self-hosted LLM, GPU memory sizing, AI infrastructure strategy, model-agnostic AI, sovereign AI"
summary: "Mistral Large 4 launched as a public preview API on 6 October 2026, a 1.05T-parameter mixture-of-experts model with 52B active, with weights promised by the end of October. This post does the memory and cost arithmetic for self-hosting it next to Aleph Alpha's 78B Kolibri, and explains why an API-first, weights-later release rewards a platform that can move a workload between the two."
banner: ""
thumbnail: ""
linkedin: |
  Mistral Large 4, which Mistral very officially calls le Chonk, went into public preview on 6 October.

  The numbers from Mistral's own model page: 1.05T total parameters, 52B active, a 1M-token context, trained from scratch on 3,800 Grace Blackwell GPUs in Mistral's own European datacenters.

  The detail that matters for infrastructure planning is the order of release. Today it is an API, listed at $1.36 per million input tokens and $4.18 per million output. The weights come by the end of the month, after red-teaming with security leaders and state authorities.

  Run the memory arithmetic and the two halves of a sensible plan appear.

  At one byte per parameter, the weights alone are about 1,050 GB. That fits one 8-GPU DGX B200 node at 1,440 GB, with room left for context. It is a dedicated-node decision.

  Aleph Alpha's Kolibri, released three days earlier under Apache 2.0, is 78B total and 3.46B active. About 78 GB at the same precision. One H200-class GPU.

  Those are not competing choices. They are two tiers of the same estate, and the right answer this month is to start a workload on the hosted API and be able to move it onto your own hardware when the weights land, without rewriting anything.

  That is a property of the platform, not the model. On ibl.ai you own all the code and the data, and the platform is model-agnostic, so a model release is a configuration change.

  #iblai #OpenWeights #Mistral #EnterpriseAI #AIInfrastructure #SovereignAI
---

## The Short Answer

**Mistral Large 4 is a 1.05T-parameter mixture-of-experts model with 52B active, released on 6 October 2026 as an API, with weights promised by month's end. Self-hosting it needs a full 8-GPU node; Aleph Alpha's 78B Kolibri fits one H200-class GPU. On ibl.ai you own all the code and the data, model-agnostic, so moving a workload between them is configuration.**

Two European open-weight releases three days apart make a tempting headline about the model layer commoditizing. The more useful reading is narrower: they sit at opposite ends of the hardware range, and one of them is not downloadable yet.

## What is Mistral Large 4 (Le Chonk), and is it open-weight yet?

Mistral Large 4 is Mistral AI's largest model, launched as a **public preview on 6 October 2026**. Mistral's [announcement](https://mistral.ai/news/mistral-large-4/) calls it "Unofficially ML4, very officially: le Chonk."

Mistral's [model page](https://docs.mistral.ai/models/mistral-large-4-0) lists **52B active parameters, 1.05T total**, a 1.6B vision encoder and a **1M-token context**. It is a natively multimodal mixture-of-experts.

It is open-weight by commitment, not yet in fact. Mistral says "We will release the weights by the end of the month," and is red-teaming the model meanwhile with cybersecurity leaders, vetted partners and state authorities.

The announcement does not state the licence the weights will ship under. That is the first thing to read when they arrive, because it decides whether "open-weight" means commercial self-hosting without terms.

Mistral says it trained the model from scratch on **3,800 NVIDIA Grace Blackwell GPUs** in its own European datacenters and serves the preview there. It claims the model significantly outperforms "any open-weight model developed in the US or Europe."

That is Mistral's own benchmark claim, and it is scoped: US and European open models, not every open model. Treat it as a vendor claim until an evaluation on your own workloads agrees.

## How much GPU memory does it take to self-host Mistral Large 4?

Self-hosting Mistral Large 4 takes roughly a full 8-GPU node, because every one of its 1.05T parameters has to sit in memory even though only 52B are active per token.

The arithmetic is simple. At one byte per parameter (FP8), 1.05T parameters is about **1,050 GB** of weights. At two bytes (BF16) it is about **2,100 GB**. Whatever memory is left goes to the context cache and batching.

NVIDIA lists the [DGX B200](https://www.nvidia.com/en-us/data-center/dgx-b200/) at **1,440 GB** of GPU memory across 8 GPUs, and the [H200](https://www.nvidia.com/en-us/data-center/h200/) at **141 GB** per GPU. Here is how both models fit:

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Model</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Total / active</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Weights at FP8</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Weights at BF16</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Smallest fit at FP8</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Mistral Large 4</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">1.05T / 52B</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">~1,050 GB</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">~2,100 GB</td>
      <td style="padding:0.75rem;">One DGX B200 node (1,440 GB), ~390 GB headroom</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Aleph Alpha Kolibri</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">78.1B / 3.46B</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">~78 GB</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">~156 GB</td>
      <td style="padding:0.75rem;">One H200 (141 GB), ~63 GB headroom</td>
    </tr>
  </tbody>
</table>

An 8-GPU H200 node holds **1,128 GB**, which fits Mistral Large 4 at FP8 with under 80 GB to spare. That is enough to load it and not much else, so long-context serving pushes toward B200-class memory or quantization below 8 bits.

The figures are weights only, before runtime overhead, and assume the published parameter counts. They are a sizing floor, not a deployment spec. The same exercise for 2-trillion-parameter models is in [The Open-Weight Tipping Point](/blog/open-weight-tipping-point-two-trillion-parameter-models).

## Why do mixture-of-experts models change the cost of self-hosting?

Mixture-of-experts models split the cost of self-hosting in two: total parameters set how much memory you buy, and active parameters set how much compute each token uses.

Both European releases are sparse in nearly the same proportion. Mistral Large 4 activates about **5%** of its parameters per token (52B of 1.05T). Kolibri, per its [model card](https://huggingface.co/Aleph-Alpha/Kolibri-1), activates about **4.4%** (3.46B of 78.1B).

So per-token compute for Mistral Large 4 is roughly **15 times** Kolibri's (52B against 3.46B active), while its memory bill is roughly **13 times** larger (1.05T against 78.1B total). Neither model is cheap or expensive in the abstract.

What follows for infrastructure is that a mixture-of-experts model is priced by its memory footprint first. A 52B-active model does not run on 52B worth of hardware; it runs on hardware that can hold 1.05T.

That is why the two models land in different tiers. One is a dedicated node you size, power and schedule. The other is a single large-memory GPU, which suits the sovereign, air-gapped deployments in our [Kolibri and Armada analysis](/blog/kolibri-armada-sovereign-ai-model-and-compute-layers).

## What does Mistral Large 4 cost through the API, and how does that compare with per-seat AI?

Mistral Large 4's preview API is listed at **$1.36 per million input tokens** and **$4.18 per million output tokens**, with cached input at $0.14. Mistral's model page also shows a launch sale at half those rates.

Token pricing has a property per-seat licensing does not: the bill follows use. A per-seat licence charges the same for an employee who sends a thousand requests a month and one who sends none, so above a small team it is the wrong shape for AI.

To make the gap concrete, take an assumption, not a measurement: an employee who sends 40 requests a working day, each with 2,000 input tokens and 500 output tokens, over 22 working days.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Line item (assumed usage, list price)</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Tokens per month</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Cost per user</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Cost for 1,000 users</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Input (880 requests × 2,000 tokens)</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">1.76M</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$2.39</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$2,394</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Output (880 requests × 500 tokens)</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">0.44M</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$1.84</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$1,839</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Total at Mistral Large 4 list price</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">2.2M</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>$4.23</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>$4,233</strong></td>
    </tr>
  </tbody>
</table>

Hold that $4.23 against any per-seat quote you have, and remember the per-seat figure is charged for the users who never open the tool as well. Your own usage logs, not this assumption, are what the comparison should run on.

Self-hosting changes the shape again: once the weights are out, cost becomes the node you run, not the tokens. Whether that beats the API depends on utilization, which the [price-floor analysis](/blog/open-weight-models-price-floor-enterprise-ai) works through.

## What should an enterprise AI infrastructure plan do with an API-first, weights-later release?

Start the evaluation on the hosted API now, and build it so the same workload can move to self-hosted weights later with nothing rewritten. Mistral Large 4's release order rewards exactly that.

Four decisions hold up whichever way the evaluation goes:

1. **Evaluate on your own cases during the preview.** Mistral's benchmark claims are its own; an evaluation set built from your documents and tasks is the number that matters.
2. **Size two tiers, not one.** A dedicated 8-GPU node for the heavy model and single-GPU capacity for a sparse model like Kolibri cover very different workloads at very different costs.
3. **Read the licence on weights day.** Until the licence is published, plan for self-hosting but do not commit to it.
4. **Route by workload.** Keep the large model for tasks that need it and send routine traffic to the cheaper tier, the pattern visible in [Vercel's AI Gateway token-share data](/blog/open-weight-models-62-percent-volume-9-percent-spend).

None of those four steps depends on which model wins. They depend on the layer above the model being able to switch.

## Where does ibl.ai fit when frontier models ship as open weights?

On ibl.ai you own all the code and the data. The platform runs model-agnostic across any LLM, so adding Mistral Large 4 through its API today and pointing the same agents at self-hosted weights later is a configuration change.

The platform layer, not the model, carries what has to persist across that move: agent definitions, the audit log, access controls, evaluation sets and data connections. They stay on infrastructure you administer.

ibl.ai deploys anywhere, from your own cloud to on-premise or [fully air-gapped](/service/air-gapped-ai), with no per-seat pricing, on [Agentic OS](/product/agentic-os). 1.6M+ users across 400+ organizations run it this way, including NVIDIA, MIT, and Syracuse University.

The policy side of the same shift, including how U.S. rules treat open-weight models, is in [How Washington Made Sovereign AI the Path of Least Resistance](/blog/open-weight-models-sovereign-ai).

## Want to run Mistral Large 4 and Kolibri on infrastructure you own?

We deploy the platform as source code you keep, sized to the model tiers you actually need. [Book a 30-minute demo](https://cal.com/iblai/30min) or [talk to the ibl.ai team](/contact).

*Sources: Mistral Large 4's release, parameter counts, context, pricing, training hardware, red-teaming and weights timing from Mistral's [announcement](https://mistral.ai/news/mistral-large-4/) and [model page](https://docs.mistral.ai/models/mistral-large-4-0); Kolibri's parameters, licence and context from its [model card](https://huggingface.co/Aleph-Alpha/Kolibri-1); GPU memory from NVIDIA's [DGX B200](https://www.nvidia.com/en-us/data-center/dgx-b200/) and [H200](https://www.nvidia.com/en-us/data-center/h200/) pages. Memory and cost figures are our arithmetic on those published numbers.*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
