---
title: "Open Weights Tied Grok 4.7. What That Does to Procurement"
slug: "open-weight-models-close-the-gap-government-procurement"
author: "Mikel Amigot"
date: "2026-09-24 11:00:00"
category: "Premium"
topics: "open-weight models, government AI procurement, sovereign AI, model-agnostic infrastructure, MiMo-V2.6-Pro, DoWI 8430.01, self-hosted AI"
summary: "Xiaomi's MIT-licensed MiMo-V2.6-Pro scored 46 on Artificial Analysis's Intelligence Index on 22 September 2026 — the same score as Grok 4.7, released a day earlier. DeepSeek V5, meanwhile, has not shipped."
banner: ""
thumbnail: ""
linkedin: |
  On 22 September 2026 Xiaomi released an MIT-licensed open-weight model that tied Grok 4.7 on an independent index.

  MiMo-V2.6-Pro scored 46 on Artificial Analysis's Intelligence Index v4.3.2. Grok 4.7, released the day before, scored 46 too. Artificial Analysis runs all ten evaluations itself under one harness, which is why this comparison means something and most vendor-published tables do not.

  Two corrections worth making, because both are circulating wrong.

  It is not "an agent benchmark" — it is a composite of ten evaluations, of which agentic tasks are 30%. And the widely-shared DeepSeek V5 "leak" is a rumour: DeepSeek's own API documentation still lists deepseek-flash and deepseek-v4-pro, with no V5 anywhere.

  The procurement consequence does not depend on either detail being dramatic.

  → An agency writing a model name into a statement of work is buying the fastest-depreciating component in the system
  → Published API prices: Grok 4.7 at $2 per million input tokens, MiMo-V2.6-Pro at $0.435 — and MIT weights can be run on your own hardware for the cost of the GPU
  → Xiaomi also published the technical report, the RL training code and more than 7,000 task environments, so the result is reproducible rather than asserted
  → DoWI 8430.01, effective 8 September 2026, already orders preference as reuse, then open source, then commercial off-the-shelf

  The durable asset is not the model. It is the platform underneath it: identity, retrieval over agency records, guardrails, evaluation, audit, cost control, and the ability to change the model without re-procuring the system.

  With ibl.ai you own all the code and the data — self-hosted inside your own perimeter, model-agnostic across any LLM, usage-based with no per-seat pricing, deployable anywhere from your own cloud to a fully air-gapped network.

  #iblai #AgenticAI #EnterpriseAI #SovereignAI #OpenWeights #GovTech
---

## The Short Answer

**Xiaomi released MiMo-V2.6-Pro on 22 September 2026 under an MIT licence, and it scores 46 on Artificial Analysis's Intelligence Index — the same score as Grok 4.7, which shipped the day before at $2 per million input tokens. DeepSeek V5 has not shipped. The durable procurement asset is the platform, not the model: with ibl.ai you own all the code and the data.**

A model is the component of an AI system with the shortest half-life. Public-sector contracts are written on the longest timescales. That mismatch is the whole problem.

## Did DeepSeek V5 leak, and does it exist?

No. As of 24 September 2026 there is no DeepSeek V5 — no release note, no model card, no weights, no API identifier.

DeepSeek's own [API documentation](https://api-docs.deepseek.com/news/news) lists `deepseek-flash` and `deepseek-v4-pro` as the current models. Hugging Face's `transformers` library [documents the V4 family](https://huggingface.co/docs/transformers/en/model_doc/deepseek_v4) — V4-Flash, V4-Pro and their Base siblings. Nothing beyond it.

What is real is the V4 line, released 24 April 2026 under an MIT licence, with V4.1-Flash following on 10 September 2026.

So the circulating claim — that V5 "leaked," was built from scratch, and shipped with full training code — is unverified. It is worth saying plainly, because the argument built on top of it does not need it.

Something else shipped that week, and it is better evidence.

## What did Xiaomi's MiMo-V2.6-Pro actually tie Grok 4.7 on?

A composite index, not a single agent benchmark — and the distinction matters more than it sounds.

Xiaomi released the MiMo-V2.6 series on **22 September 2026**. The flagship is a 1.02-trillion-parameter sparse mixture-of-experts model with **42 billion active parameters** and a **1M-token context window**.

It is [published on Hugging Face under an MIT licence](https://forkast.news/xiaomis-mimo-v2-6-ships-open-weights-at-frontier-class-performance-and-the-timing-is-not-an-accident/).

It scored **46 on Artificial Analysis's Intelligence Index v4.3.2** — [tying Grok 4.7](https://venturebeat.com/technology/better-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash), which xAI released one day earlier and which also scores 46.

That index is [ten evaluations in four weighted groups](https://artificialanalysis.ai/methodology/intelligence-benchmarking): agents at 30%, general at 30%, coding at 20%, scientific reasoning at 20%. Agentic tasks are a large share of it, not the whole of it.

The part that makes the comparison worth citing is the harness. Artificial Analysis runs every evaluation itself, on internal copies of the datasets, with identical temperature and output-token settings across models — agentic tasks through its own open-source Stirrup harness.

Most published model comparisons are incommensurable: two vendors, two prompt strategies, two scaffolds, one table. This one is not, which is why it is admissible evidence and a vendor slide usually is not.

Xiaomi also published the technical report, the reinforcement-learning training code, and [more than 7,000 task environments](https://www.unite.ai/xiaomis-new-flagship-model-leads-open-weight-rankings-with-a-score-of-46/). Reproducible, rather than asserted.

## Why does an open-weight model at parity break a multi-year model contract?

Because it removes the two things the contract was priced on: scarcity and switching cost.

Take the published API prices. Grok 4.7 lists at [**$2 per million input tokens and $6 per million output**](https://www.marktechpost.com/2026/09/21/spacexai-releases-grok-4-7/). MiMo-V2.6-Pro lists at [**$0.435 input and $0.870 output**](https://llm-stats.com/models/mimo-v2.6-pro).

That is **4.6x on input and 6.9x on output** for the same index score, derived from the two vendors' own price lists. And because the weights are MIT, an agency with its own GPUs can skip the API entirely and pay only for the hardware.

One correction to how this is usually framed. "Every model contract signed today buys a capability open source will match in months" is a forecast, and a single tie is not a law of nature.

What is demonstrable is narrower and still sufficient: on one day in September 2026, a freely licensed model matched a same-week proprietary release on an independently administered index. That is a pricing risk a contracting officer can reason about, not a prophecy.

The right response is not to bet on open weights. It is to stop writing a bet on any specific model into a document that lives for years.

## What should a government agency actually procure if the model is the commodity?

The layer that does not depreciate — everything between the agency's data and the model's API.

That layer is identity and role-based access tied to the existing directory. Retrieval over agency records with provenance. Guardrails enforced server-side. Evaluation harnesses. Complete audit trails. Budget caps and cost attribution. Operator tooling.

None of it improves when a better model appears, and none of it transfers when you change vendors. It is the expensive, slow, durable part.

A statement of work that names a model is procuring the fastest-depreciating component in the system and calling it the deliverable. A statement of work that specifies the platform — and requires that the model be swappable — procures the part that survives the option years.

This is the same argument as [competence benchmarks over security certifications in government AI procurement](/blog/government-ai-agent-competence-benchmarks-procurement-2026), arriving from the supply side rather than the evaluation side.

It is also why [what government buyers should require from an AI vendor](/blog/what-government-buyers-should-require-from-an-ai-vendor) is a list of platform properties, not model properties.

## Does DoWI 8430.01 already push procurement in this direction?

Partly, and it is worth being precise about how far.

Department of War Instruction 8430.01 was approved on 31 August 2026 and took effect on 8 September 2026. It orders software preference as reuse first, then open source, then commercial off-the-shelf, then new development.

It also states that non-public department information may not be processed by generative AI services that do not reside on department systems.

That second clause is the operative one here. A model reaching parity is worth nothing to an agency that cannot run it where the sensitive data already is — which makes deployment location, not benchmark score, the binding constraint.

To be clear about what the instruction does not say: it contains no requirement for model independence and no requirement that a vendor deliver source code.

Those remain architecture decisions an agency has to make for itself. The full reading is in [DoWI 8430.01 bans external AI hosting, not just training](/blog/dod-dowi-8430-01-software-defined-warfare-sovereign-ai).

The regulatory backdrop has been moving the same way for over a year, including [the federal framework that exempted open-weight models from review entirely](/blog/open-weight-models-sovereign-ai).

## How does ibl.ai deploy for government agencies?

By making the model a configuration value and the platform the thing the agency owns.

With ibl.ai you own all the code and the data.

The platform runs on the agency's own infrastructure with full source code under a perpetual licence.

It is model-agnostic across any LLM — Claude, GPT, Gemini, Llama, Command, an MIT-licensed open-weight model, or your own fine-tune — and switching is a configuration change, not a re-procurement.

Pricing is usage-based with no per-seat pricing, so cost tracks workload rather than headcount. You can deploy anywhere: your own cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

For federal buyers, [the government deployment](/solutions/government) supports IL4/IL5 workloads, NIST 800-53 controls across the stack, and PIV/CAC authentication, with agency data never leaving the agency environment.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY.

*Related reading: [how Washington made sovereign AI the path of least resistance](/blog/open-weight-models-sovereign-ai) — the regulatory half of the same shift · [DoWI 8430.01 bans external AI hosting, not just training](/blog/dod-dowi-8430-01-software-defined-warfare-sovereign-ai) · [government AI procurement's blind spot](/blog/government-ai-agent-competence-benchmarks-procurement-2026)*

*Sources: the MiMo-V2.6 release, MIT licence and Intelligence Index tie from [VentureBeat](https://venturebeat.com/technology/better-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash), [heise](https://www.heise.de/en/news/MiMo-V2-6-Pro-New-model-leads-open-weight-ranking-11462279.html) and [Forkast](https://forkast.news/xiaomis-mimo-v2-6-ships-open-weights-at-frontier-class-performance-and-the-timing-is-not-an-accident/); the training code and 7,000+ environments from [Unite.AI](https://www.unite.ai/xiaomis-new-flagship-model-leads-open-weight-rankings-with-a-score-of-46/); index composition and harness from [Artificial Analysis](https://artificialanalysis.ai/methodology/intelligence-benchmarking); Grok 4.7's release and pricing from [MarkTechPost](https://www.marktechpost.com/2026/09/21/spacexai-releases-grok-4-7/); MiMo pricing from [LLM Stats](https://llm-stats.com/models/mimo-v2.6-pro); DeepSeek's current model list from its [API documentation](https://api-docs.deepseek.com/news/news) and the [transformers V4 model card](https://huggingface.co/docs/transformers/en/model_doc/deepseek_v4).*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
