---
title: "Decade-Long Compute Bets Face Two Opposite Curves"
slug: "ai-compute-supply-demand-2026-2040-decade-long-bets"
author: "Blanca Amigot"
date: "2026-08-20 10:00:00"
category: "Premium"
topics: "ai compute, capex forecast, inference cost, gpu supply, model-agnostic, ai infrastructure planning, data center"
summary: "Frontier training costs are rising while the cost of a fixed capability has fallen roughly 1,000x in three years. Any decade-long AI infrastructure bet has to survive both curves, and they point in opposite directions."
banner: ""
thumbnail: ""
linkedin: |
  A new market model maps AI compute supply and demand out to 2040, and financial firms are making decade-long infrastructure commitments against it.

  Before signing one, it's worth separating two curves that get conflated constantly — including in most of the commentary about this report.

  Curve 1: frontier training is getting MORE expensive. GPT-4 cost north of $100M to train in 2023. Frontier budgets have gone up since, not down.

  Curve 2: a fixed capability is getting radically cheaper. GPT-4 launched at $30/$60 per million tokens in March 2023. By mid-2026 you get equal-or-better quality from open weights for under $0.50 per million — roughly 1,000x in three years.

  These are not the same trend and they point in opposite directions.

  I keep seeing them merged into "frontier training used to cost $100M, now it runs on one GPU." That's wrong. Training a frontier model on one GPU is not a thing. Running GPT-4-class *inference* cheaply absolutely is — and that's the curve that decides your budget.

  Why it matters for a ten-year bet: McKinsey's base case is $5.2T of AI data center capex, with a $3.7T–$7.9T range. The market is short today and much of the industry is planning for possible oversupply around 2028–2030.

  So you're committing for a decade to a capability tier that reprices by ~10x a year, in a compute market that may invert inside three years.

  The only architecture that survives both curves is one where the model is a component you swap and the compute is yours to redirect. On ibl.ai you own all the code and the data, run it model-agnostic across any LLM, with no per-seat pricing.

  #iblai #AgenticAI #EnterpriseAI #AI #Infrastructure #CIO
---

## The Short Answer

**Frontier training costs are rising while the cost of a fixed capability falls roughly 1,000x in three years, so a decade-long compute commitment must survive two opposite curves. On ibl.ai you own all the code and the data and run it model-agnostic across any LLM, so the model is a swappable component — with no per-seat pricing, and you can deploy anywhere, including fully air-gapped.**

A [new market model for computing and AI in data centers](https://www.globenewswire.com/news-release/2026/08/17/3345987/28124/en/the-global-market-for-computing-and-ai-for-data-centers-report-2026-2040-evaluating-gpus-custom-asics-arm-vs-x86-and-strategies-of-nvidia-amd-google-and-aws.html), published on August 17, 2026, runs bull, base and bear scenarios from 2026 out to 2040 across GPUs, custom ASICs, Arm and x86.

Financial firms and other large buyers are making decade-long infrastructure commitments against exactly this kind of forecast.

Before signing one, it is worth separating two cost curves that are constantly conflated — including in most of the commentary about compute economics. They are not the same trend, and they point in opposite directions.

## Which two curves actually matter?

**Frontier training is getting more expensive.** GPT-4 reportedly cost in excess of $100 million to train in 2023, with some estimates nearer $150 million. Frontier training budgets have risen since, not fallen.

Being at the frontier costs more each cycle, which is why the number of organizations operating there keeps shrinking.

**A fixed capability is getting radically cheaper.** GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output.

By mid-2026, equal-or-better quality is available from open-weight models for **under $0.50 per million tokens** — roughly a 95% decline in two years and close to **1,000x over three** for a fixed capability level.

The first curve describes what it costs to build the best model in the world. The second describes what it costs *you* to do the thing you actually wanted done.

Almost every buyer decision depends on the second, and almost every headline quotes the first.

## Isn't it true that frontier training now runs on one GPU?

No, and the conflation is worth naming because it circulates widely.

Training a frontier model on a single GPU is not a thing and is not close to being a thing.

What has become cheap is **inference at a capability level that was frontier three years ago** — GPT-4-class quality now runs on modest hardware, and single-stream H100 inference at around $0.73 per million tokens drops to roughly $0.18 with batching on the same GPU.

That is a claim about *reproducing yesterday's frontier*, not about *reaching today's*.

The distinction matters for planning because the two support opposite conclusions. If frontier training were collapsing in cost, more organizations would train their own models.

Because it is inference-at-fixed-capability that is collapsing, the rational move is the opposite: stop trying to own the frontier and get very good at consuming whatever it produces, cheaply.

## What do the demand forecasts actually say?

That the range is enormous, which is itself the finding.

McKinsey modelled three scenarios from constrained to accelerated demand, with a base case of **$5.2 trillion** in AI data center capex and a range spanning **$3.7 trillion to $7.9 trillion** depending on the adoption trajectory.

A forecast whose plausible range is more than double from floor to ceiling is not telling you what will happen. It is telling you the uncertainty is irreducible at this horizon.

The near-term picture adds a second complication. The compute market is in genuine shortage today, with windfall pricing — while a substantial part of the industry is simultaneously planning for possible **oversupply around 2028 to 2030**.

Committing for ten years into a market that may invert from shortage to glut inside three is the actual risk being taken, and it is rarely stated that plainly in a business case.

## What does this mean for an infrastructure decision?

That optionality is worth more than any specific forecast, and it has to be designed in rather than negotiated later.

Three consequences follow directly from the two curves.

**Do not lock the model layer.** Capability commoditizes downward at roughly an order of magnitude a year.

A multi-year commitment to one vendor's models is a bet against the single most reliable trend in the sector — the argument we set out in [Model-Agnostic AI: Why Single-Vendor Lock-In Is the Real Risk](/blog/model-agnostic-ai-the-real-risk-is-vendor-lock-in).

**Do not lock the compute location either.** If oversupply arrives around 2028–2030, prices fall for whoever can move. An architecture that runs the same way in your cloud, in someone else's, or on-premise can chase that; one welded to a single provider's managed service cannot.

**Watch the inference-to-training ratio.** For every $1 billion spent training a model, organizations face an estimated $15–20 billion in inference costs over its production lifetime. Training is the headline; inference is the bill.

## Why does per-seat pricing fit this badly?

Because it prices headcount while every underlying cost curve is priced in tokens.

When the cost of serving a fixed capability falls roughly 1,000x in three years and your contract is denominated in employees, none of that decline reaches you. The vendor absorbs it.

That is not a hypothetical — it is what the last three years already did, and per-seat prices have not fallen 1,000x.

A usage-based or owned-compute structure passes the curve through to the buyer. That is the whole argument in [Per-Seat vs Usage-Based AI Pricing](/resources/comparisons/per-seat-vs-usage-based-ai-pricing), and a decade-long horizon makes it sharper rather than softer.

## What survives both curves?

An architecture where the model is a component and the compute is a decision you can revisit.

Concretely: run any model, so a cheaper equivalent can be adopted the week it appears. Own the platform code, so switching models or clouds is engineering rather than a migration project.

Keep deployment portable across your cloud, on-premise and air-gapped, so a market that inverts is an opportunity instead of a stranded commitment.

That is how ibl.ai is built.

The platform is in production with 1.6M+ users from 400+ organizations, ships with the full source code under a perpetual licence, and is model-agnostic by construction — so a ten-year bet reduces to a series of one-year decisions you can actually change.

The related question of what happens when the runtime layer commoditizes too is in [The Agent Runtime Just Commoditized. Now What?](/blog/ai-agent-runtime-commoditized-what-you-actually-pay-for), and the near-term arithmetic in [What Does AI Actually Cost in 2026?](/blog/what-does-ai-actually-cost-in-2026).

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
