---
title: "Memory Is the Constraint on Which Model You Can Run"
slug: "ai-memory-hbm-bottleneck-enterprise-infrastructure-strategy"
author: "Miguel Amigot"
date: "2026-09-15 11:00:00"
category: "Premium"
topics: "HBM, memory prices, AI infrastructure, self-hosted AI, model-agnostic, quantization, KV cache, enterprise AI"
summary: "WSTS puts 2026 memory revenue above $800 billion, up about 250% year over year, and Huawei has raised Ascend 950DT prices 20–50% on HBM costs. The enterprise lever is not supply. It is which model you run, and where."
banner: ""
thumbnail: ""
linkedin: |
  Memory stopped being a line item this year and became the thing that decides what you can run.

  WSTS now forecasts the global semiconductor market at $1.51 trillion in 2026, with the memory segment alone above $800 billion — up roughly 250% year over year. Memory is more than half the industry this year.

  Reuters reported last week that Huawei's Ascend 950DT now quotes above 250,000 yuan, 20–50% higher than quotes given to customers about two months earlier, with HBM procured through grey-market channels as the cause.

  None of that is a supply problem an enterprise can solve. The large cloud providers locked multi-year agreements that cap their prices; TrendForce says the increases now shift to buyers without them.

  So the useful question is not how to get cheaper HBM. It is which decisions you still control.

  → Model choice. Tencent's Hy4 is 1.56 TB of weights and will not load on an 8xH100 node. NPCI's banking model targets 16 GB and runs on one 80 GB GPU. Roughly 100x apart, both shipped in the last three weeks.
  → Precision. Google's TurboQuant cuts KV-cache memory at least 6x with no accuracy loss.
  → Caching. A cache hit is a fraction of the uncached input price, and it is a configuration decision.
  → Routing. If the platform is model-agnostic, a memory price shock becomes a routing change instead of a renegotiation.

  A per-seat subscription gives you none of these. You inherit whatever the vendor's cost structure does next.

  With ibl.ai you own all the code and the data — self-hosted inside your own perimeter, model-agnostic across any LLM, usage-based with no per-seat pricing, deployable anywhere from your own cloud to a fully air-gapped network.

  #iblai #AgenticAI #EnterpriseAI #AIInfrastructure #HBM #SelfHostedAI
---

## The Short Answer

**Memory, not compute, now decides which AI model an enterprise can run and where. WSTS puts 2026 memory revenue above $800 billion, up about 250% year over year, and Huawei has raised Ascend 950DT prices 20–50% on HBM costs. The lever you control is model choice, quantization and caching, and with ibl.ai you own all the code and the data.**

No enterprise is going to fix the HBM supply chain. Every enterprise can decide what it asks memory to hold.

## How much have memory and AI chip prices actually moved in 2026?

Enough that memory is now more than half the semiconductor industry by revenue.

The World Semiconductor Trade Statistics organization's spring forecast, [reported on 5 June 2026](https://www.digitimes.com/news/a20260605VL208/semiconductor-industry-wsts-growth-forecast-2026.html), puts the global semiconductor market at **USD 1.51 trillion in 2026, a 90% increase**.

[WSTS attributes the revision](https://www.wsts.org/76/Recent-News-Release) overwhelmingly to one segment: **memory, forecast to surge around 250% year over year to more than USD 800 billion**, against Logic at 37% growth. WSTS projects roughly USD 1.9 trillion for 2027.

One correction to the framing that usually travels with this story. A widely circulated claim holds that Nomura projects **$3.7 trillion** in memory revenue by 2030, "larger than the entire semiconductor industry is worth today."

We could not verify that figure against Nomura's own note or any major outlet, so it is not used here. The comparison is also incommensurable: a 2030 forecast set against a present-day industry total measures two different years.

The verified version is more useful anyway. Memory did not overtake a past industry total at some point in the future. It is already over half of a $1.51 trillion industry this year.

Downstream prices moved with it. TrendForce, [on 9 July 2026](https://www.trendforce.com/presscenter/news/20260709-13140.html), forecast **server DRAM contract prices up 13–18% quarter over quarter in 3Q26**.

The same release noted that several U.S. cloud providers hold multi-year agreements capping their prices, so the increases now shift "toward customers without LTAs."

Then the accelerators. [Reuters reported on 10 September 2026](https://www.investing.com/news/stock-market-news/exclusivechinas-ai-chipmakers-raise-prices-as-highbandwidth-memory-shortage-bites-4894950) that Huawei's Ascend 950DT is quoted above 250,000 yuan (about $37,255), **20% to 50% above quotes given to customers roughly two months earlier**.

That last detail corrects a common retelling. The move was reported last week, but the comparison window is about two months, not one.

And Reuters attributes the cause to grey-market HBM sourcing, on three people familiar with the pricing who declined to be identified, rather than inferring it.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">What moved</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Move</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Window</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Reported by</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Worldwide memory revenue</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">+~250% to &gt;$800B</td>
      <td style="padding:0.75rem;">FY2026 forecast</td>
      <td style="padding:0.75rem;">WSTS (May 2026), via DigiTimes 5 Jun</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Server DRAM contract price</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">+13–18% QoQ</td>
      <td style="padding:0.75rem;">3Q26</td>
      <td style="padding:0.75rem;">TrendForce, 9 Jul 2026</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Huawei Ascend 950DT</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">+20–50%, to &gt;¥250,000</td>
      <td style="padding:0.75rem;">~2 months to Sep 2026</td>
      <td style="padding:0.75rem;">Reuters, 10 Sep 2026</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Huawei Ascend 950PR</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">¥60,000 → &gt;¥80,000</td>
      <td style="padding:0.75rem;">2026 to date</td>
      <td style="padding:0.75rem;">Reuters, 10 Sep 2026</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Huawei Ascend 910C</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">¥90,000 → &gt;¥110,000</td>
      <td style="padding:0.75rem;">2026 to date</td>
      <td style="padding:0.75rem;">Reuters, 10 Sep 2026</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Cambricon 690</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">+20–30%</td>
      <td style="padding:0.75rem;">~2 months to Sep 2026</td>
      <td style="padding:0.75rem;">Reuters, 10 Sep 2026</td>
    </tr>
  </tbody>
</table>

## Why does HBM supply decide which model an enterprise can run?

Because memory capacity is a hard gate, and the gap between models on that gate is now about two orders of magnitude.

Two real deployments from the last three weeks make it concrete.

Tencent [open-sourced Hy4 preview under Apache-2.0 on 28 August 2026](/blog/tencent-hy4-open-source-770b-vendor-pricing-leverage): 770B total parameters, 49B active, and roughly **1.56 TB of BF16 weights**. A standard 8×H100 node carries 640 GB of VRAM, so it cannot load the model. Neither can it load the FP8 checkpoint at about 814 GB.

The licence is fully permissive. The hardware requirement is not negotiable by licence.

Two weeks later, [NPCI unveiled FiMI Banking on 10 September 2026](/blog/npci-fimi-4b-model-fits-one-server-regulated-finance), targeting Gemma 4 E4B: 4.5B effective parameters and a **16 GB** BF16 checkpoint. NPCI's paper states that one 80 GB GPU holds the model and supports hundreds of concurrent sessions.

Roughly 100× apart in bytes. Both shipped in the same three-week window. Both are credible for real work in their own class.

That ratio is the whole enterprise decision. A model that fits the memory you already have is deployable next quarter.

A model that needs a node you have not bought, in a market where the large cloud providers hold price-capped supply and you do not, is a procurement project with an unknown close date.

## Which layer of the AI stack can an enterprise actually own?

Not the one absorbing the capital. The buildout numbers make the boundary obvious.

Broadcom's [Q3 FY2026 results, released 2 September 2026](https://www.sec.gov/Archives/edgar/data/0001730168/000173016826000076/avgo-08022026x8kxex99.htm), report **AI semiconductor revenue of $16.7 billion, up 221% year over year**, with Q4 guided to $21.7 billion.

On that call CEO Hock Tan [projected AI semiconductor revenue of roughly $115 billion in FY2027 and $230 billion in FY2028](https://finance.yahoo.com/technology/ai/articles/broadcom-ceo-hock-tan-defends-111113337.html), and named Anthropic as on track to become Broadcom's largest custom chip customer in 2027.

Three precisions the usual summary drops. These are Tan's projections stated on the 2 September call, not reaffirmed contractual targets.

The figures cover custom accelerators **and** networking silicon, not custom chips alone. And Broadcom itself named the customer, which is unusual and is the part worth noticing.

At that scale, the accelerator, the HBM stack and the fab are not layers an enterprise buys its way into. They are layers whose price it inherits.

The layers above are different. Which model runs, at what precision, against which cache, on which hardware, under whose contract — those are decisions made in software, and they are the ones that convert a memory price into an operating cost you can influence.

## What can an enterprise do about memory prices it does not control?

Four things, all of which are configuration rather than procurement.

**Right-size the model.** The 16 GB banking model and the 1.56 TB frontier model are not competing for the same job. Most enterprise agent workloads are closer to the first. Choosing the smaller capable model is the single largest memory decision available.

**Quantize.** Google's TurboQuant [cuts KV-cache memory at least 6× with no accuracy loss](/blog/turboquant-ai-memory-compression-own-infrastructure). KV cache is frequently the binding constraint on a self-hosted deployment, not the weights.

**Cache.** Cache hits are priced at a fraction of uncached input on every major API, and cache design is [an inference-engineering decision rather than an architecture one](/blog/inference-engineering-not-architecture-financial-firms).

**Keep the exit.** If the platform can route to a different model, a price move on any one of them is a configuration change. If it cannot, it is a renegotiation.

This is also where per-seat pricing fails structurally rather than merely costing more. A per-seat subscription prices your bill against headcount while the vendor's cost base moves with memory.

You absorb the pass-through, and you hold none of the four levers above, because none of them is yours to pull.

## How does ibl.ai turn a memory price shock into a routing decision?

By putting every one of those levers on your side of the contract.

With ibl.ai you own all the code and the data.

The platform runs on your own infrastructure with full source code, so model selection, quantization, batching and cache policy are settings you change rather than roadmap items you request.

It is model-agnostic across any LLM, so a 16 GB open-weight model and a frontier API endpoint are both routing targets. It is usage-based with no per-seat pricing, so the bill tracks tokens consumed rather than employees hired.

And it will deploy anywhere, from your own cloud to on-premise, GovCloud or a fully air-gapped network.

The practical effect is that a memory-driven price move becomes a routing change made in an afternoon, instead of a contract you reopen next year.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY.

*Related reading: [Tencent's 770B Hy4 is Apache-2.0, and 1.56 TB of weights](/blog/tencent-hy4-open-source-770b-vendor-pricing-leverage) and [NPCI's bank model fits in 16 GB](/blog/npci-fimi-4b-model-fits-one-server-regulated-finance) — the two ends of the memory gate; [Google's TurboQuant cuts AI memory 6×](/blog/turboquant-ai-memory-compression-own-infrastructure) — the compression lever you control; and [the alpha is in inference engineering, not architecture](/blog/inference-engineering-not-architecture-financial-firms) — why caching and routing beat model selection alone.*

*Sources: the 2026 semiconductor and memory forecasts from [WSTS](https://www.wsts.org/76/Recent-News-Release), dated [5 June 2026 by DigiTimes](https://www.digitimes.com/news/a20260605VL208/semiconductor-industry-wsts-growth-forecast-2026.html); server DRAM contract pricing and the long-term-agreement split from [TrendForce, 9 July 2026](https://www.trendforce.com/presscenter/news/20260709-13140.html); the Huawei and Cambricon price increases from [Reuters, 10 September 2026](https://www.investing.com/news/stock-market-news/exclusivechinas-ai-chipmakers-raise-prices-as-highbandwidth-memory-shortage-bites-4894950); Broadcom's Q3 FY2026 AI revenue from [its 2 September 2026 results release](https://www.sec.gov/Archives/edgar/data/0001730168/000173016826000076/avgo-08022026x8kxex99.htm) and the FY2027/FY2028 projections and Anthropic remark from [reporting on that earnings call](https://finance.yahoo.com/technology/ai/articles/broadcom-ceo-hock-tan-defends-111113337.html). The $3.7 trillion Nomura figure circulating with this story could not be verified and is not used.*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
