The Short Answer
Xiaomi released MiMo-V2.6-Pro on 22 September 2026 under an MIT licence, and it scores 46 on Artificial Analysis's Intelligence Index β the same score as Grok 4.7, which shipped the day before at $2 per million input tokens. DeepSeek V5 has not shipped. The durable procurement asset is the platform, not the model: with ibl.ai you own all the code and the data.
A model is the component of an AI system with the shortest half-life. Public-sector contracts are written on the longest timescales. That mismatch is the whole problem.
Did DeepSeek V5 leak, and does it exist?
No. As of 24 September 2026 there is no DeepSeek V5 β no release note, no model card, no weights, no API identifier.
DeepSeek's own API documentation lists deepseek-flash and deepseek-v4-pro as the current models. Hugging Face's transformers library documents the V4 family β V4-Flash, V4-Pro and their Base siblings. Nothing beyond it.
What is real is the V4 line, released 24 April 2026 under an MIT licence, with V4.1-Flash following on 10 September 2026.
So the circulating claim β that V5 "leaked," was built from scratch, and shipped with full training code β is unverified. It is worth saying plainly, because the argument built on top of it does not need it.
Something else shipped that week, and it is better evidence.
What did Xiaomi's MiMo-V2.6-Pro actually tie Grok 4.7 on?
A composite index, not a single agent benchmark β and the distinction matters more than it sounds.
Xiaomi released the MiMo-V2.6 series on 22 September 2026. The flagship is a 1.02-trillion-parameter sparse mixture-of-experts model with 42 billion active parameters and a 1M-token context window.
It is published on Hugging Face under an MIT licence.
It scored 46 on Artificial Analysis's Intelligence Index v4.3.2 β tying Grok 4.7, which xAI released one day earlier and which also scores 46.
That index is ten evaluations in four weighted groups: agents at 30%, general at 30%, coding at 20%, scientific reasoning at 20%. Agentic tasks are a large share of it, not the whole of it.
The part that makes the comparison worth citing is the harness. Artificial Analysis runs every evaluation itself, on internal copies of the datasets, with identical temperature and output-token settings across models β agentic tasks through its own open-source Stirrup harness.
Most published model comparisons are incommensurable: two vendors, two prompt strategies, two scaffolds, one table. This one is not, which is why it is admissible evidence and a vendor slide usually is not.
Xiaomi also published the technical report, the reinforcement-learning training code, and more than 7,000 task environments. Reproducible, rather than asserted.
Why does an open-weight model at parity break a multi-year model contract?
Because it removes the two things the contract was priced on: scarcity and switching cost.
Take the published API prices. Grok 4.7 lists at $2 per million input tokens and $6 per million output. MiMo-V2.6-Pro lists at $0.435 input and $0.870 output.
That is 4.6x on input and 6.9x on output for the same index score, derived from the two vendors' own price lists. And because the weights are MIT, an agency with its own GPUs can skip the API entirely and pay only for the hardware.
One correction to how this is usually framed. "Every model contract signed today buys a capability open source will match in months" is a forecast, and a single tie is not a law of nature.
What is demonstrable is narrower and still sufficient: on one day in September 2026, a freely licensed model matched a same-week proprietary release on an independently administered index. That is a pricing risk a contracting officer can reason about, not a prophecy.
The right response is not to bet on open weights. It is to stop writing a bet on any specific model into a document that lives for years.
What should a government agency actually procure if the model is the commodity?
The layer that does not depreciate β everything between the agency's data and the model's API.
That layer is identity and role-based access tied to the existing directory. Retrieval over agency records with provenance. Guardrails enforced server-side. Evaluation harnesses. Complete audit trails. Budget caps and cost attribution. Operator tooling.
None of it improves when a better model appears, and none of it transfers when you change vendors. It is the expensive, slow, durable part.
A statement of work that names a model is procuring the fastest-depreciating component in the system and calling it the deliverable. A statement of work that specifies the platform β and requires that the model be swappable β procures the part that survives the option years.
This is the same argument as competence benchmarks over security certifications in government AI procurement, arriving from the supply side rather than the evaluation side.
It is also why what government buyers should require from an AI vendor is a list of platform properties, not model properties.
Does DoWI 8430.01 already push procurement in this direction?
Partly, and it is worth being precise about how far.
Department of War Instruction 8430.01 was approved on 31 August 2026 and took effect on 8 September 2026. It orders software preference as reuse first, then open source, then commercial off-the-shelf, then new development.
It also states that non-public department information may not be processed by generative AI services that do not reside on department systems.
That second clause is the operative one here. A model reaching parity is worth nothing to an agency that cannot run it where the sensitive data already is β which makes deployment location, not benchmark score, the binding constraint.
To be clear about what the instruction does not say: it contains no requirement for model independence and no requirement that a vendor deliver source code.
Those remain architecture decisions an agency has to make for itself. The full reading is in DoWI 8430.01 bans external AI hosting, not just training.
The regulatory backdrop has been moving the same way for over a year, including the federal framework that exempted open-weight models from review entirely.
How does ibl.ai deploy for government agencies?
By making the model a configuration value and the platform the thing the agency owns.
With ibl.ai you own all the code and the data.
The platform runs on the agency's own infrastructure with full source code under a perpetual licence.
It is model-agnostic across any LLM β Claude, GPT, Gemini, Llama, Command, an MIT-licensed open-weight model, or your own fine-tune β and switching is a configuration change, not a re-procurement.
Pricing is usage-based with no per-seat pricing, so cost tracks workload rather than headcount. You can deploy anywhere: your own cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.
For federal buyers, the government deployment supports IL4/IL5 workloads, NIST 800-53 controls across the stack, and PIV/CAC authentication, with agency data never leaving the agency environment.
1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
ibl.ai is family-owned and operated from New York, NY.
Related reading: how Washington made sovereign AI the path of least resistance β the regulatory half of the same shift Β· DoWI 8430.01 bans external AI hosting, not just training Β· government AI procurement's blind spot
Sources: the MiMo-V2.6 release, MIT licence and Intelligence Index tie from VentureBeat, heise and Forkast; the training code and 7,000+ environments from Unite.AI; index composition and harness from Artificial Analysis; Grok 4.7's release and pricing from MarkTechPost; MiMo pricing from LLM Stats; DeepSeek's current model list from its API documentation and the transformers V4 model card.