The Short Answer
At VMware Explore on 31 August 2026, Broadcom announced VMware AI Factory β the software-defined foundation of VMware Private AI Cloud β running 150+ open-source models on AMD Instinct MI350 GPUs inside existing VMware Cloud Foundation environments, with no per-token pricing. It removes the technical objection to running AI privately. On ibl.ai you own all the code and the data, so the platform layer above it stays yours regardless of whose hypervisor you standardised on.
The announcement is genuinely significant. It is also worth reading for which dependency it removes and which one it introduces.
What did Broadcom actually ship at VMware Explore 2026?
VMware AI Factory, announced 31 August 2026, is the software-defined foundation underneath VMware Private AI Cloud. The components that matter to a buyer:
Models. More than 150 open-source models run through VMware Cloud Foundation's vLLM-based runtime, including Nemotron 3, Gemma 4, Qwen 3.7-Max and GLM 5.2. Models from NVIDIA, Google DeepMind, NEC, Alibaba and Z.ai are being validated as on-premises model services.
Silicon. A validated path pairing AMD Instinct MI350 Series GPUs with the open ROCm software ecosystem on VMware Cloud Foundation β notable mostly because it is not NVIDIA, in a market where "private AI" has usually meant one vendor's accelerators.
Provisioning. Zero-touch orchestration across vSphere, vSAN, Kubernetes and the GPU operator, so an AI workload is stood up through the same path as any other workload.
Pricing. Explicitly no per-token pricing. Capacity is provisioned, not metered.
Alongside it, Broadcom shipped agent governance into the infrastructure layer β AgentMinder, VMware vDefend and the Avi Load Balancer β which we covered separately in agent governance moving into infrastructure.
Why does provisioning matter more than the model count?
Because 150 models was never the blocker.
Open weights have been downloadable for years. vLLM has been available to anyone willing to run it.
The models were never the hard part for a regulated enterprise β the hard part was everything around them: GPU scheduling, driver stacks, Kubernetes operators, network policy, and a security review for each.
What AI Factory changes is who does that integration work and how it is consumed. An enterprise already running VMware β which is most large enterprises β can provision an AI workload through infrastructure its team already knows, with a vendor relationship it already has.
That is a procurement change disguised as a technology announcement, and procurement is usually what was actually blocking the project.
Which objection does this remove for regulated industries?
The single most common one: we cannot run AI because our data cannot leave our infrastructure.
For a hospital handling PHI, a bank under supervisory examination, an agency with jurisdictional requirements, or a firm holding privileged material, that objection has been load-bearing. It has justified years of not deploying.
It is now much weaker. If validated models run inside your own VMware environment on GPUs you own, the data does not move. Clinical documentation, prior authorisation, coding support, claims review β all of it can run against records that never cross the firewall.
We have argued the specific economics of that case before, for the healthcare revenue cycle, where roughly 65% of denied claims are never appealed while 54% of appeals succeed. The blocker there was never the model quality. It was the architecture around PHI.
| Concern | Vendor-hosted AI | Private AI on infrastructure you run |
|---|---|---|
| Where regulated data sits | Leaves your perimeter | Never moves |
| Cost shape | Per token or per seat, scales with use | Provisioned capacity |
| Model choice | The vendor's models | 150+ open models, swappable |
| Remaining dependency | Model vendor and its cloud | Whoever licenses the platform layer |
What dependency does this introduce?
The bottom row is the one to read twice.
Moving off a model vendor's cloud does not eliminate lock-in. It relocates it.
AI Factory runs on VMware Cloud Foundation, which means the AI strategy of an enterprise adopting it now inherits the licensing, versioning and roadmap of one vendor β a vendor whose licensing changes have been a standing topic for enterprise buyers since 2023.
That is not an argument against it. VMware is genuinely where most enterprise workloads run, and meeting workloads where they already are is exactly right.
It is an argument for keeping the layers separate. The infrastructure layer, the model layer and the platform layer that provides governance, identity, memory and routing are three different decisions.
An architecture that couples them means every future change to one requires renegotiating the others.
Own the platform layer, and the infrastructure underneath becomes a choice you can revisit β including choosing VMware, if that is what your estate runs.
How should you evaluate private AI infrastructure now?
Four questions, in the order they actually bite:
Can the data stay where it is? This is now answerable yes by several stacks. It is table stakes, not a differentiator.
Can you change models without changing platforms? Open weights on a vLLM runtime is a good sign. Verify it extends to models that do not exist yet β that is what model-agnostic has to mean.
What is the cost shape as usage grows? Provisioned capacity behaves fundamentally differently from a per-token meter. A workload that triples in volume should not triple your bill if you own the hardware it runs on.
Which vendor's roadmap does your AI strategy now depend on? If the honest answer is one name, you have improved your data posture and kept your commercial exposure.
The threshold Broadcom crossed is real: private AI stopped being the harder option this week. The question is no longer whether you can run AI on your own infrastructure. It is how much of the stack above it you actually control.
Sources: Broadcom β VMware AI Factory announcement Β· The Next Platform β VMware Intros Private AI Cloud, AI Factory As Workloads Shift To On-Prem