The Short Answer
Kimi K3 reached third on GDPval-AA v2 at roughly a third of frontier pricing, Meta shipped a 30B Apache-2.0 agentic model that runs on one consumer GPU, and Anthropic cut Fable-line cache reads 75% weeks later. Open weights now set the market's price floor β but only buyers who can credibly switch capture the benefit. With ibl.ai you own all the code and the data, model-agnostic across any LLM.
The savings are real. The leverage is bigger, and it belongs only to organizations whose architecture makes switching a configuration change.
What has actually shipped in open weights recently?
Three releases that matter for different reasons:
| Model | Date | Scale | $/M in | $/M out |
|---|---|---|---|---|
| Kimi K3 (Moonshot) | 16 Jul 2026 | 2.8T MoE, 1M context | $3 | $15 |
| Muse Glimmer (Meta) | 10 Aug 2026 | 30B, 128K context | self-hosted | self-hosted |
| GPT-6 Astra (reference) | 3 Sept 2026 | hosted API only | $10 | $50 |
Kimi K3 is the pricing story. Moonshot's 2.8-trillion-parameter mixture-of-experts model with a 1M-token context window benchmarks in genuine frontier territory β third on GDPval-AA v2, behind Claude Fable 5 Max and GPT-5.6 Sol Max β at $3/$15 per million tokens with cached input at $0.30. That is roughly a third of frontier list pricing for results that are not a tier down β with one caveat the same source raises: K3 consumes considerably more tokens per task, so effective cost on identical work lands closer to $810β920 against $525 for GPT-5.6 Terra. List price divided by three is not cost divided by three.
Muse Glimmer is the deployment story. Meta's 30B agentic model is Apache-2.0 and ungated, has a 128K context, and runs under 20GB at 4-bit on a single consumer GPU. It leads on MCP Atlas at 75.5 against Qwen3.6-27B's 62.5 β tool-calling being exactly the workload most agent deployments run constantly. A bundled DFlash drafter for speculative decoding takes an RTX 5090 from 74.9 to 233.4 tokens per second, a 3.1x speedup.
Anthropic's 1 September change is two numbers, and they are easy to conflate. The list price is unchanged at $10/$50 per million tokens; what fell 75% is cache reads, from $1.00 to $0.25. Anthropic separately says that lowers effective bills by about 25% for typical workloads and up to 45% for highly agentic ones, based on four weeks of its own August usage. So the widely-quoted 45% is real β it is an effective-cost figure for context-heavy agent work, not a cut to list pricing.
Does an open-weight model actually pressure closed-model pricing?
The sequence is suggestive, and the mechanism is straightforward.
When a model of genuinely frontier scale can be downloaded, self-hosted and fine-tuned under a permissive licence, the ceiling on what a hosted API can charge for comparable work is set by what it costs to run the open one yourself.
That is sustained structural pressure rather than a promotional cycle, and Anthropic's cache-read cut and Google's decision to hold introductory Flash pricing across three consecutive releases are plausibly responses to it.
The direction is worth stating carefully. Nobody has published a causal account tying a specific price cut to a specific open release, and vendors rarely explain their pricing.
What is observable is that frontier-adjacent open weights arrived, and prices moved down shortly after, repeatedly.
Why does the ability to switch matter more than actually switching?
Because the option has value whether or not it is exercised, and it is the part most enterprises do not have.
A buyer who can credibly self-host negotiates differently. They can evaluate a hosted model against a self-hosted one on their own workload and choose on merit rather than on migration cost.
When a vendor reprices or deprecates a version, they have a response other than absorbing it.
A buyer whose authentication, retrieval, guardrails, evaluation harness and audit logging are welded to one provider's API has none of that. Their published price falls with everyone else's, but their effective cost is set by a relationship they cannot leave.
Prices fall for the whole market; only buyers with somewhere else to go capture the fall.
This is the same asymmetry we described in why vendor lock-in is the real risk in model-agnostic AI β here with an unusually clear price tag attached.
When does self-hosting actually beat a hosted API?
Three conditions, and honesty about them matters more than advocacy.
Sustained, predictable volume. GPU capacity is a fixed cost. High steady throughput amortizes it; bursty low volume does not, and a hosted API is genuinely cheaper there.
Data that cannot leave. For workloads under residency, classification or air-gap constraints, self-hosting is not a cost decision at all β it is the only lawful architecture, and the comparison never happens.
Workloads a smaller model handles well. Most production volume is classification, extraction and summarization, where a 30B model on a single GPU is entirely adequate. Reserving the frontier model for work that needs it is where most real savings come from β which requires routing, which requires model-agnostic infrastructure.
How does ibl.ai make the open-weight floor usable?
By making the model layer the replaceable component rather than the foundation.
With ibl.ai you own all the code and the data.
The platform is deployed on your own infrastructure with full source code access, runs any LLM β hosted frontier models, self-hosted open weights, or both behind one routing policy β is usage-based with no per-seat pricing, and deploys anywhere from your own cloud to on-premise, GovCloud, or a fully air-gapped network.
That combination is what converts a falling market price into a falling bill: you can route each workload to whatever is cheapest and adequate this quarter, re-run your evaluation set against a new model without a migration, and keep the sensitive workloads on hardware you control throughout.
ibl.ai is family-owned and operated from New York, NY.
Related reading: K2 Horizon and the fully open model fleet, and what published inference benchmarks reveal about the cost floor.
Sources: Kimi K3 specifications, benchmark placement and pricing via Solvimon's pricing analysis; Muse Glimmer specifications and benchmarks from Meta AI Research.