---
title: "Base Labs, Marin, Nemotron: The Moat Is Architecture"
slug: "base-labs-baseten-open-source-lab-model-agnostic-moat"
author: "Miguel Amigot"
date: "2026-09-10 18:00:00"
category: "Premium"
topics: "open-source models, open-weight models, Base Labs, model-agnostic architecture, enterprise AI, self-hosted AI, Baseten"
summary: "Base Labs published its manifesto on September 2, 2026, joining Stanford's Marin open lab and NVIDIA's eight-lab Nemotron Coalition. As open models multiply, the durable asset is the architecture that swaps them."
banner: ""
thumbnail: ""
linkedin: |
  Base Labs, a research lab run by the inference platform Baseten, published its manifesto and agenda on September 2, 2026.

  Two corrections worth making. It is not a lab built to train open frontier models — its own mandate is "to build the science of training in the open," and its agenda is continual learning, RL environments and data, a safety stack for open models, and serving cost.

  And it is a commercial platform's research arm. More useful open models mean more inference, and Baseten sells inference. That does not make the research less real. It makes the incentive legible.

  Base Labs is the third shape this has taken since May 2025:

  → Marin, Stanford's open lab, announced May 19, 2025 — experiments declared as code in pull requests, training logs public, Marin 8B Base trained on 12.7 trillion tokens
  → NVIDIA's Nemotron Coalition, March 16, 2026 — eight labs pooling research, data and compute to train an open base model
  → Base Labs, September 2, 2026 — training science and serving economics, published without exception

  The pattern matters more than any one release. If a better open model arrives every few weeks, no single model is a moat. The durable asset is the architecture underneath, where adding a model is a registry entry rather than a migration.

  With ibl.ai you own all the code and the data — self-hosted inside your own perimeter, model-agnostic across any LLM, usage-based with no per-seat pricing, deployable anywhere from your own cloud to a fully air-gapped network.

  #iblai #AgenticAI #EnterpriseAI #OpenSourceAI #OpenWeights #SelfHosted
---

## The Short Answer

**Base Labs is Baseten's research lab, launched with a manifesto and research agenda dated September 2, 2026, and its mandate is to make open models more useful — not to train a frontier model of its own. It joins Stanford's Marin open lab and NVIDIA's eight-lab Nemotron Coalition. With ibl.ai you own all the code and the data, so a new open model is a registry entry rather than a migration.**

A model advantage is a dated artifact. The architecture that lets you adopt the next one is a standing capability.

## What is Base Labs, and what is its actual mandate?

Base Labs is a research lab operated by Baseten, the inference platform company. Its site describes the lab as ["working to advance and democratize open-source intelligence"](https://labs.baseten.co/).

The [manifesto](https://labs.baseten.co/manifesto) states the mandate more precisely: "We exist to build the science of training in the open." Its one procedural rule is that everything gets published — "We publish without exception" — negative results included, in plain English.

The manifesto and the research agenda are both dated **September 2, 2026**.

The framing that circulated with the launch — a research org built solely to advance open-source frontier models — is worth correcting, because the actual agenda is both narrower and more useful than that.

Base Labs is not training a frontier model.

[Its published agenda](https://labs.baseten.co/agenda) covers areas including: blue-sky work on continual learning and the science of model training; a "BaseHub Data Foundry" producing RL environments and mid-training data for open models; "post-post training," including a safety stack for open-source models; and making models affordable through distillation, quantization and speculative decoding.

That is a lab about training methods, data and serving economics, not about topping a leaderboard.

The second correction is about independence. Base Labs is a commercial platform's research arm. It was founded after [Baseten's December 10, 2025 acquisition of Parsed](https://techstartups.com/2025/12/10/baseten-acquires-parsed-to-double-down-on-specialized-ai-over-general-purpose-models/), a reinforcement-learning and post-training startup founded by Mudith Jayasekara and Charles O'Neill.

More useful open models mean more open-model inference, and Baseten sells inference. That does not make the research less real. It makes the incentive legible, which is the more honest way to read it.

## How does Base Labs fit alongside Marin and NVIDIA's Nemotron Coalition?

It is the third distinct shape a dedicated open-model effort has taken since May 2025.

**Marin** is Stanford's open lab, [announced May 19, 2025](https://marin.community/blog/2025/05/19/announcement/) out of CRFM and HAI. Its distinguishing move is procedural: experiments are tracked as GitHub issues, declared as code in pull requests, and training logs are public as runs proceed — mistakes and dead ends included. It released Marin 8B Base and Instruct, with the Base model trained on 12.7 trillion tokens.

**The Nemotron Coalition** is NVIDIA's consortium, [announced March 16, 2026](https://nvidianews.nvidia.com/news/nvidia-launches-nemotron-coalition-of-leading-global-ai-labs-to-advance-open-frontier-models) at GTC. Eight labs — Black Forest Labs, Cursor, LangChain, Mistral AI, Perplexity, Reflection AI, Sarvam and Thinking Machines Lab — pool research, data and compute on DGX Cloud to train an open base model that will underpin the upcoming Nemotron 4 family.

A university lab, a chipmaker-convened consortium, and a platform company's research arm. Different funders, different incentives, different definitions of "open."

What they share is the output shape: weights, methods and data released for anyone to run, rather than access sold through an endpoint.

## If open models keep arriving, why is no single model a moat?

Because each release resets the comparison, and the reset interval is now short enough that a model choice cannot be a strategy.

The pattern shows up on the release side as well as the research side. MBZUAI's Institute of Foundation Models shipped [six Apache-2.0 models from 0.9B to 375B parameters in a single day](/blog/k2-horizon-fully-open-model-fleet-enterprise), with training code and data published for part of the fleet and promised for the rest.

Open weights have also started [setting the price of closed models](/blog/open-weight-models-price-floor-enterprise-ai) rather than trailing them.

An organization that standardized on one model in this environment has bought a dated artifact. The advantage it selected for will be matched, undercut or licensed more permissively within a quarter or two.

The organizations that benefit are the ones for whom the next release is an evaluation task rather than a project. That property is not in the model. It is in the layer beneath it.

## Does an open-model lab's research help you if you cannot run the models yourself?

Only partly, and this is where the enterprise version of the question diverges from the research one.

Published weights and published recipes are genuinely free to read. The constraint is the deployment surface. If you consume a model through a vendor's API, you get each new model when that vendor adds it, priced how that vendor prices it, retired when that vendor retires it.

Baseten's own product line makes the point cleanly. [Baseten for Model Labs](https://www.baseten.co/blog/announcing-baseten-for-model-labs/), announced July 29, 2026, launched with 15 lab partners so that labs can distribute models through one platform instead of building their own infrastructure.

That is good for developers and good for small labs. It is also still a distribution layer someone else owns — the same structural point as [the Nvidia–Hugging Face lock-in question](/blog/nvidia-hugging-face-acquisition-open-weight-model-lock-in): open weights protect you from a model vendor, not from whoever controls how you get them.

## What does an enterprise need so a new open model is a registry entry, not a migration?

Five properties, none of which is a model capability.

- **The model is configuration, not code.** Adding a model means an entry in a registry with its endpoint, context window and cost, not a rewrite of the calling layer.
- **Evaluations run as regression tests.** A candidate model is scored against your own task suite before it is routed any production traffic.
- **Routing is per workload.** A cheap model handles classification, a strong one handles synthesis, and the assignment changes without touching application logic.
- **The data, prompt and tool layer is model-independent.** Retrieval, permissions and tool definitions survive a model change, because they were never coupled to a particular provider's API shape.
- **The right to run the weights inside your perimeter.** A permissive licence is only useful with somewhere to run it that you control.

Get those right and an open-model release is a Tuesday. Get them wrong and every release is a migration you decline to do, which is how organizations end up two model generations behind on infrastructure they cannot change.

## How does ibl.ai make swapping an open model a configuration change?

With ibl.ai you own all the code and the data.

The platform is self-hosted with full source code access, so the weights, the retrieval layer and the audit trail all sit inside your own perimeter.

It is model-agnostic across any LLM — a hosted frontier model, an open-weight model you run on your own GPUs, or a fine-tune you trained yourself — and switching between them is a registry change, not a rebuild.

Pricing is usage-based with no per-seat pricing, which matters directly here: per-seat SaaS charges by headcount, so the bill does not fall when a cheaper open model does the same work. Usage-based billing passes that saving through.

And it deploys anywhere — your own cloud, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity, which is the only configuration in which an open model is genuinely yours to run.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY.

*Related reading: [K2 Horizon and what a fully open model fleet changes](/blog/k2-horizon-fully-open-model-fleet-enterprise) — why a shared architecture across model sizes matters more than any single release, and [the Nvidia–Hugging Face lock-in question](/blog/nvidia-hugging-face-acquisition-open-weight-model-lock-in) on who owns the distribution layer.*

*Sources: the lab's mandate, publishing rule and September 2, 2026 dating from [the Base Labs site](https://labs.baseten.co/), its [manifesto](https://labs.baseten.co/manifesto) and its [research agenda](https://labs.baseten.co/agenda); the Parsed acquisition and founders from [Tech Startups' December 10, 2025 report](https://techstartups.com/2025/12/10/baseten-acquires-parsed-to-double-down-on-specialized-ai-over-general-purpose-models/); Marin's open-lab method, model releases and 12.7T training tokens from [the Marin announcement](https://marin.community/blog/2025/05/19/announcement/); the Nemotron Coalition's eight members and remit from [NVIDIA's March 16, 2026 newsroom release](https://nvidianews.nvidia.com/news/nvidia-launches-nemotron-coalition-of-leading-global-ai-labs-to-advance-open-frontier-models); the 15 launch partners from [Baseten for Model Labs](https://www.baseten.co/blog/announcing-baseten-for-model-labs/).*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
