The Short Answer
Frontier training costs are rising while the cost of a fixed capability falls roughly 1,000x in three years, so a decade-long compute commitment must survive two opposite curves. On ibl.ai you own all the code and the data and run it model-agnostic across any LLM, so the model is a swappable component β with no per-seat pricing, and you can deploy anywhere, including fully air-gapped.
A new market model for computing and AI in data centers, published on August 17, 2026, runs bull, base and bear scenarios from 2026 out to 2040 across GPUs, custom ASICs, Arm and x86.
Financial firms and other large buyers are making decade-long infrastructure commitments against exactly this kind of forecast.
Before signing one, it is worth separating two cost curves that are constantly conflated β including in most of the commentary about compute economics. They are not the same trend, and they point in opposite directions.
Which two curves actually matter?
Frontier training is getting more expensive. GPT-4 reportedly cost in excess of $100 million to train in 2023, with some estimates nearer $150 million. Frontier training budgets have risen since, not fallen.
Being at the frontier costs more each cycle, which is why the number of organizations operating there keeps shrinking.
A fixed capability is getting radically cheaper. GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output.
By mid-2026, equal-or-better quality is available from open-weight models for under $0.50 per million tokens β roughly a 95% decline in two years and close to 1,000x over three for a fixed capability level.
The first curve describes what it costs to build the best model in the world. The second describes what it costs you to do the thing you actually wanted done.
Almost every buyer decision depends on the second, and almost every headline quotes the first.
Isn't it true that frontier training now runs on one GPU?
No, and the conflation is worth naming because it circulates widely.
Training a frontier model on a single GPU is not a thing and is not close to being a thing.
What has become cheap is inference at a capability level that was frontier three years ago β GPT-4-class quality now runs on modest hardware, and single-stream H100 inference at around $0.73 per million tokens drops to roughly $0.18 with batching on the same GPU.
That is a claim about reproducing yesterday's frontier, not about reaching today's.
The distinction matters for planning because the two support opposite conclusions. If frontier training were collapsing in cost, more organizations would train their own models.
Because it is inference-at-fixed-capability that is collapsing, the rational move is the opposite: stop trying to own the frontier and get very good at consuming whatever it produces, cheaply.
What do the demand forecasts actually say?
That the range is enormous, which is itself the finding.
McKinsey modelled three scenarios from constrained to accelerated demand, with a base case of $5.2 trillion in AI data center capex and a range spanning $3.7 trillion to $7.9 trillion depending on the adoption trajectory.
A forecast whose plausible range is more than double from floor to ceiling is not telling you what will happen. It is telling you the uncertainty is irreducible at this horizon.
The near-term picture adds a second complication. The compute market is in genuine shortage today, with windfall pricing β while a substantial part of the industry is simultaneously planning for possible oversupply around 2028 to 2030.
Committing for ten years into a market that may invert from shortage to glut inside three is the actual risk being taken, and it is rarely stated that plainly in a business case.
What does this mean for an infrastructure decision?
That optionality is worth more than any specific forecast, and it has to be designed in rather than negotiated later.
Three consequences follow directly from the two curves.
Do not lock the model layer. Capability commoditizes downward at roughly an order of magnitude a year.
A multi-year commitment to one vendor's models is a bet against the single most reliable trend in the sector β the argument we set out in Model-Agnostic AI: Why Single-Vendor Lock-In Is the Real Risk.
Do not lock the compute location either. If oversupply arrives around 2028β2030, prices fall for whoever can move. An architecture that runs the same way in your cloud, in someone else's, or on-premise can chase that; one welded to a single provider's managed service cannot.
Watch the inference-to-training ratio. For every $1 billion spent training a model, organizations face an estimated $15β20 billion in inference costs over its production lifetime. Training is the headline; inference is the bill.
Why does per-seat pricing fit this badly?
Because it prices headcount while every underlying cost curve is priced in tokens.
When the cost of serving a fixed capability falls roughly 1,000x in three years and your contract is denominated in employees, none of that decline reaches you. The vendor absorbs it.
That is not a hypothetical β it is what the last three years already did, and per-seat prices have not fallen 1,000x.
A usage-based or owned-compute structure passes the curve through to the buyer. That is the whole argument in Per-Seat vs Usage-Based AI Pricing, and a decade-long horizon makes it sharper rather than softer.
What survives both curves?
An architecture where the model is a component and the compute is a decision you can revisit.
Concretely: run any model, so a cheaper equivalent can be adopted the week it appears. Own the platform code, so switching models or clouds is engineering rather than a migration project.
Keep deployment portable across your cloud, on-premise and air-gapped, so a market that inverts is an opportunity instead of a stranded commitment.
That is how ibl.ai is built.
The platform is in production with 1.6M+ users from 400+ organizations, ships with the full source code under a perpetual licence, and is model-agnostic by construction β so a ten-year bet reduces to a series of one-year decisions you can actually change.
The related question of what happens when the runtime layer commoditizes too is in The Agent Runtime Just Commoditized. Now What?, and the near-term arithmetic in What Does AI Actually Cost in 2026?.