---
title: "Per-Seat AI Is Priced Against a Falling Floor"
slug: "inference-cost-floor-jalapeno-published-benchmarks"
author: "Miguel Amigot"
date: "2026-09-01 14:00:00"
category: "Premium"
topics: "AI inference cost, OpenAI Jalapeño, Broadcom, custom silicon, NVIDIA GB300, self-hosting AI, per-seat pricing, token economics, RTX PRO 6000, vLLM"
summary: "OpenAI published Jalapeño's benchmarks at Hot Chips 2026: 1.5–1.9x throughput per kilowatt and 1.7–3.6x lower latency than NVIDIA's GB200 and GB300, at 700W against 1,400W. Inference costs have fallen roughly 95% in two years, and every per-seat AI licence is priced against a floor that keeps dropping."
banner: ""
thumbnail: ""
linkedin: |
  OpenAI published real benchmarks for its Jalapeño inference chip at Hot Chips 2026. The numbers are worth reading precisely, because the summaries have been rounding them generously.

  Jalapeño, co-developed with Broadcom, delivered 1.5x to 1.9x more throughput per kilowatt than NVIDIA's GB200 and GB300 rack systems, and 1.7x to 3.6x lower end-to-end latency. On highly interactive workloads, 2.1x to 4.1x higher performance.

  The power figure is the one I would put in front of a CFO: 700W, against roughly 1,200W for GB200 and 1,400W for GB300. In a datacentre, watts are the constraint before dollars are.

  Specs: 216GB of HBM4, up to 13.4 MXFP4 PFLOPS, a NUMA-style architecture across 64 memory/core slices. Design to tapeout in nine months, with AI used in the design loop. OpenAI plans to deploy it in its own infrastructure by the end of 2026 while continuing to buy NVIDIA.

  Now the part that matters for buyers.

  Microsoft has Maia. Meta has MTIA. Amazon has Inferentia and Trainium. Google has TPUs. Every hyperscaler is building inference silicon, and the competition pushes the floor down for everyone.

  Meanwhile GPT-4-class output has gone from about $30 per million input tokens in 2023 to under $0.50 from open-weight models in 2026 — roughly 95% in two years, and Stanford's AI Index put one measure of the decline at 280x.

  So consider what a per-seat AI licence at $20-60 per user per month actually is. It is a fixed price written against a cost base that is collapsing underneath it. The vendor absorbs the improvement. You do not.

  Own the infrastructure and every cost reduction is yours automatically. On ibl.ai you own all the code and the data, run it model-agnostic across any LLM, and pay with no per-seat pricing.

  #iblai #AIInference #EnterpriseAI #CustomSilicon #SelfHostedAI #AIEconomics
---


## The Short Answer

**OpenAI published Jalapeño's benchmarks at Hot Chips 2026: 1.5–1.9x more throughput per kilowatt and 1.7–3.6x lower latency than NVIDIA's GB200 and GB300, at 700W against up to 1,400W. Combined with inference costs falling roughly 95% in two years, every per-seat AI licence is priced against a collapsing floor. On ibl.ai you own all the code and the data, so each cost reduction reaches you instead of a vendor's margin.**

The chip is interesting. The pricing consequence is the part that belongs in a budget conversation.

## What did OpenAI actually publish about Jalapeño?

Jalapeño is an inference ASIC OpenAI co-developed with Broadcom, detailed at Hot Chips 2026 in late August. The published comparisons, against NVIDIA's GB200 and GB300 rack systems:

- **1.5x to 1.9x more throughput per kilowatt**
- **1.7x to 3.6x lower end-to-end latency**
- **2.1x to 4.1x higher performance** on highly interactive workloads

Note the ranges. Most coverage has quoted the top of each — "1.9x and 3.6x" — which is the best case, not the result.

The hardware: **216GB of HBM4**, up to **13.4 MXFP4 PFLOPS**, a NUMA-style architecture built around **64 memory/core slices**, drawing **700W**. NVIDIA's comparators draw roughly **1,200W (GB200)** and **1,400W (GB300)**.

OpenAI reports moving from initial design to manufacturing tapeout in **nine months**, using AI in the design loop for implementation exploration, verification and arithmetic circuit optimisation.

Deployment into its own infrastructure is planned by the end of 2026 — alongside, not instead of, continued NVIDIA purchases.

That last detail is the honest one. This is not a company replacing its GPU fleet. It is a company adding a workload-specific accelerator for the workload it runs most.

## Why does the power number matter more than the throughput number?

Because in a datacentre, power is the binding constraint before capital is.

A 700W part delivering more throughput than a 1,400W part is not a 2x improvement on a spreadsheet line — it changes how many accelerators fit in a rack, what cooling is required, and whether an existing facility can host the deployment at all.

For an enterprise evaluating on-premise inference, this is the difference between "we need a new facility" and "this fits in the rack we have." Efficiency gains at the silicon layer show up as *deployability* long before they show up as unit economics.

We flagged the fragmentation of the inference market when [custom silicon started splitting from training hardware](/blog/custom-silicon-enterprise-ai-inference) in June. What is new is not that OpenAI designed a chip — that was known. It is that there are now published numbers to compare against.

## How far have inference costs actually fallen?

Far enough that pricing assumptions from 2024 no longer describe the market.

GPT-4 launched in 2023 at roughly **$30 per million input tokens**. By mid-2026, equal-or-better quality is available from open-weight models at **under $0.50 per million tokens** — a decline of roughly **95% in two years**. Stanford's AI Index measured one version of this decline at **280x**.

Three forces compound:

**Model efficiency.** Smaller models reaching last generation's quality.

**Inference optimisation.** vLLM with PagedAttention delivering multiples of throughput through continuous batching; quantisation (INT4/INT8, FP8) shrinking memory footprints.

**Hardware competition.** Jalapeño, Microsoft's Maia, Meta's MTIA, Amazon's Inferentia and Trainium, Google's TPUs. Every hyperscaler building inference silicon pushes the floor down for everyone, including for the open-weight models an enterprise runs itself.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Cost model</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">1,000 users</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">5,000 users</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Behaviour as usage grows</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Per-seat licence @ $30/user/mo</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$360,000/yr</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$1,800,000/yr</td>
      <td style="padding:0.75rem;">Scales with headcount, used or not</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Per-seat licence @ $60/user/mo</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$720,000/yr</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">$3,600,000/yr</td>
      <td style="padding:0.75rem;">Scales with headcount, used or not</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Self-hosted on owned hardware</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">GPU + power</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">GPU + power</td>
      <td style="padding:0.75rem;"><strong>Scales with actual work; absorbs every cost decline</strong></td>
    </tr>
  </tbody>
</table>

Headcount figures are illustrative at published list prices. The point is the *shape*, not the exact cell: one column grows with how many people you employ, the other with how much work you actually do.

## What does this mean for a per-seat AI contract?

It means you are paying a fixed price against a falling cost base, and the difference accrues to the vendor.

A per-seat licence signed on 2024 cost assumptions is a bet that the underlying cost of inference stays roughly where it was. It has not.

When a model gets 40% cheaper to serve, an organisation running its own infrastructure sees that as a lower bill. An organisation on a per-seat contract sees the same invoice.

This is why per-seat is structurally the wrong shape at scale rather than merely a more expensive option. It prices the wrong thing — people, rather than work — and it insulates you from exactly the improvement you would most want to capture.

The hardware side is now within reach for real workloads. A single **NVIDIA RTX PRO 6000 Blackwell** with **96GB of ECC GDDR7** can self-host a 70-billion-parameter model at FP8.

For organisations at sustained volume, published analyses put break-even on hardware in the region of **9 to 18 months** under continuous load, with self-hosting delivering multiples of cost saving over premium API pricing at high token throughput.

Those break-even figures deserve a caveat that vendor comparisons usually omit: they assume continuous load. Hardware sitting idle overnight has a much worse payback than a spreadsheet showing peak throughput suggests. Model your actual duty cycle, not your peak.

## What should you actually do about it?

**Price the work, not the seats.** Estimate tokens per month for the workflows you are actually automating. If that number is small, a metered API is genuinely the right answer and you should not build a datacentre. If it is large and growing, per-seat pricing is working against you.

**Keep the model swappable.** The entire benefit of a falling floor depends on being able to move to whatever is cheapest and good enough next quarter. An architecture that is model-agnostic across any LLM converts every price drop into your saving. One wired to a single vendor's model converts it into theirs.

**Re-examine contracts written before this year.** A three-year per-seat agreement priced in 2024 is being served by inference that costs a fraction of what it did when the ink dried.

The floor is still falling. The question is only whether your architecture lets you follow it down.

**Sources:** [OpenAI — OpenAI and Broadcom unveil LLM-optimized inference chip](https://openai.com/index/openai-broadcom-jalapeno-inference-chip/) · [Tom's Hardware — Hot Chips 2026: OpenAI's Jalapeño AI ASIC unpacked](https://www.tomshardware.com/tech-industry/artificial-intelligence/hot-chips-2026-openais-jalapeno-ai-asic-unpacked-accelerator-developed-using-ai-achieves-efficiency-and-throughput-gains-against-power-hungry-blackwell) · [Broadcom — OpenAI and Broadcom Unveil LLM-Optimized Intelligence Processor](https://investors.broadcom.com/news-releases/news-release-details/openai-and-broadcom-unveil-llm-optimized-intelligence-processor)

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
