---
title: "You Cannot Govern a Clinical Model You Cannot Observe"
slug: "clinical-ai-governance-observability-hospitals-own-the-stack"
author: "ibl.ai Engineering"
date: "2026-09-29 11:00:00"
category: "Premium"
topics: "clinical AI, healthcare AI governance, FDA AI devices, hospital IT, model drift, self-hosted AI"
summary: "Anthropic reported that ~950 agents surfaced a novel enzyme system in 21 hours, and the first FDA-approved AI margin-assessment device reached its first operating rooms in August. The capability question is closing. The governance one is not — and even that approved device ships AI updates under a change-control plan the hospital does not hold."
banner: ""
thumbnail: ""
linkedin: |
  Two data points from clinical AI in the last few weeks, and one uncomfortable conclusion.

  Anthropic reported that roughly 950 Claude agent sessions, running 21 hours and 210 million tokens, surveyed 1.9 billion protein clusters, recovered about 200,000 reverse transcriptases and surfaced a previously unknown enzyme system it calls ART — work it says would have taken an expert weeks. Preprint dated 23 September.

  And on 3 March the FDA granted premarket approval to Claire, from Perimeter Medical Imaging AI — the first AI-enabled imaging device approved in the US for intraoperative breast cancer margin assessment, at 88.1% margin accuracy in its pivotal trial. In August, Intermountain Health became the first US health system to put it into operating rooms.

  So capability is not the constraint any more. 75% of health systems are using or planning at least one AI application.

  The harder constraint is one nobody puts in a procurement scorecard: you cannot govern a model you cannot observe.

  And here is the part that surprised me. Claire's own PMA authorises a predetermined change control plan — "planned AI enhancements that can be implemented without additional FDA interaction." That is sensible regulation: the changes were reviewed in advance and are bounded. But it means even an approved class III AI device updates on a schedule the hospital does not set and does not hold the artefact for.

  Managed platforms do offer more than people assume — Azure lets you pin a model version and warns before defaults move. What none of them offers is the thing a morbidity and mortality review needs three years later: the ability to re-run the case exactly as it ran, after that version was retired.

  With ibl.ai you own all the code and the data: the model, the version, the logs and the audit trail sit inside the hospital, where a quality committee can actually inspect them.

  #iblai #ClinicalAI #HealthcareAI #AIGovernance #PatientSafety
---

## The Short Answer

**Clinical AI is arriving faster than hospitals can govern it: 950 agents found a novel enzyme system in 21 hours, and the FDA approved the first AI margin-assessment device in March. Ban or license, both assume you can observe the model. On ibl.ai you own all the code and the data.**

The debate about clinical AI has been about whether the models are good enough. That question is closing. The one underneath it — whether a hospital can see what the model is doing — has barely been asked.

## How fast is clinical capability actually moving?

Fast enough that the recent examples are qualitatively different from the last decade's.

On **23 September 2026** Anthropic released a preprint reporting that roughly **950 Claude agent sessions**, running for **21 hours** and consuming **210 million tokens**, surveyed **1.9 billion protein clusters**, recovered about **200,000 reverse transcriptases**, and surfaced a previously unknown enzyme system it calls **ART** — array-associated reverse transcriptases, found mostly in bacteriophages.

Anthropic presents it as early results from one of its first research programs, and says the work would have taken an expert scientist weeks or months.

Worth stating the limits as clearly as the result. **The function of the system itself is not yet known** — not merely the accessory protein's — and whether it is useful for gene editing is undetermined.

Dario Amodei has acknowledged that a Stanford team previously described a system similar in some ways to this one, so "previously unknown" is doing careful work.

It is not purely computational, though: the preprint reports that ART arrays are highly expressed and appear as discrete units during *Staphylococcus* phage infection.

Expression was confirmed at the bench; function was not. A strong demonstration of search at scale, not a validated therapy.

On **3 March 2026** the FDA granted premarket approval to **Claire**, from Perimeter Medical Imaging AI — the **first AI-enabled imaging device approved in the United States for intraoperative breast cancer margin assessment**.

It images the *excised* lumpectomy specimen during surgery, not tissue in the patient.

Its pivotal trial reported **88.1% margin accuracy** and a statistically significant reduction in patients with residual cancer against standard of care.

Repeat surgery occurs in about **20% of breast-conserving surgeries** in the US, against roughly **300,000** breast cancer surgeries a year.

In **August 2026** Intermountain Health became the first US health system to deploy it commercially, at LDS and American Fork hospitals. Five months from approval to an operating room is fast for a class III device.

## So what is the bottleneck now?

Deployment, and then something harder than deployment.

The deployment part is documented. In a February 2026 survey of 120 US health systems, **75%** were using or planning to use at least one AI application.

The share implementing or planning **three or more** AI solutions rose from 30% to 59% — a **67%** year-on-year increase. Respondents named slow implementation timelines among their challenges.

That is a real drag, and it is the one most people name. But a slow procurement cycle is a solvable problem — budget, staffing, integration work. The constraint underneath it is not.

## What is the constraint underneath it?

You cannot govern a model you cannot observe.

When clinical AI runs on a vendor's cloud, four things are outside the hospital's control at once:

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Question a quality committee must answer</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Vendor cloud</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Self-hosted</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">Which model version produced this recommendation, on this date?</td><td style="padding:0.75rem;">Knowable if you pinned it — several platforms let you</td><td style="padding:0.75rem;">A version you pinned</td></tr>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">When did the weights last change?</td><td style="padding:0.75rem;">On the vendor's schedule, with notice — until the version retires</td><td style="padding:0.75rem;">When you decided</td></tr>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">Has behaviour drifted since validation?</td><td style="padding:0.75rem;">Measurable only from outputs</td><td style="padding:0.75rem;">Measurable against a fixed artefact</td></tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;"><strong>Can you reproduce the decision a year later?</strong></td><td style="padding:0.75rem;"><strong>Only if the vendor retained it</strong></td><td style="padding:0.75rem;"><strong>Yes</strong></td></tr>
  </tbody>
</table>

Be fair to the managed platforms first, because the absolute version of this claim is wrong.

Azure lets a deployment opt out of automatic model upgrades, exposes the current version through the portal and API, gives at least two weeks' notice before a default moves, and keeps the previous major version until its retirement date.

ONC's HTI-1 rule goes further, requiring certified health IT to disclose the update and validation schedule for predictive decision support.

So versions can be pinned and changes can be announced. What ends is the pinning: a retired version is gone, and with it the ability to re-run a case as it ran.

**And the regulated case is not the clean exception either.** Claire's own PMA authorises a predetermined change control plan covering "planned AI enhancements that can be implemented without additional FDA interaction." That is good regulation — the changes were reviewed in advance and are bounded — but it means an approved class III AI device also updates without a fresh submission, on a schedule the hospital does not set.

That is the sense in which "ban it" and "license it" answer a narrower question than they appear to. Both assume you can see what the model is doing well enough to decide whether it is acceptable.

Prohibition does not need observability because nothing is deployed. Licensure absolutely does — and licensing a service whose behaviour changes on a schedule you do not hold licenses a moving target.

## Isn't this an argument against clinical AI generally?

No, and it would be a bad one. Repeat surgery after one in five breast-conserving operations is a real harm, and Claire's trial showed a statistically significant reduction in patients left with residual cancer.

The company puts the re-operation benefit as potential rather than measured. Either way, a hospital that refuses AI on principle is choosing the status quo.

The argument is about where the model runs, not whether it runs. An institution can adopt aggressively and still insist on three things: a version it controls, logs it holds, and the ability to reproduce a decision for a morbidity and mortality review or a malpractice claim.

Those are ordinary clinical governance expectations. They are only difficult when the model is somebody else's service.

## What does a hospital need in place before it scales?

Four things, and none of them is a model choice:

1. **A pinned version** — the artefact you validated is the artefact in production, and changing it is a decision with a date and a signature.
2. **Logs you hold** — inputs, outputs, model version and the clinician who reviewed it, in your systems.
3. **Drift measurement against a fixed baseline** — not vibes, and not the vendor's dashboard.
4. **Reproducibility** — the ability to re-run a case as it ran then, years later, because that is the window in which claims arrive.

## Why does ownership decide all four?

Because every one of them is a property of where the model lives.

On ibl.ai you own all the code and the data. The platform runs under a perpetual licence inside the hospital's own perimeter, so the model version, the inference logs and the audit trail are institutional records rather than a vendor's telemetry.

It is model-agnostic, which is what makes the pinning real: a health system can run an open-weight model entirely inside its own network, upgrade on its own schedule, and keep the prior version available for reproduction.

Pricing is usage-based with no per-seat pricing, and you can deploy anywhere: your cloud, your VPC, on-premise, or fully air-gapped.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

A stethoscope does not get a firmware update. Clinical models will, including the approved ones — so the governance question was never really about the model. It is about who holds the version, the logs, and the ability to reconstruct the day.

*Sources: the enzyme discovery from [Anthropic's announcement](https://www.anthropic.com/news/claude-discovers-novel-enzyme-system), its [preprint](https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf) and [TechCrunch](https://techcrunch.com/2026/09/23/anthropic-says-its-biology-lab-has-already-found-something-big/); Claire's approval, PCCP, indications and the re-excision figure from [Perimeter's approval release](https://www.prnewswire.com/news-releases/perimeter-medical-imaging-ais-claire-becomes-first-fda-approved-ai-enabled-imaging-device-for-breast-cancer-surgery-302703103.html), and the first deployment from [its Intermountain release](https://ir.perimetermed.com/news-events/press-releases/detail/213/intermountain-health-brings-perimeters-claire-to-the); the change-control pathway from [FDA's PCCP guidance](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence); version pinning from [Azure's model-version documentation](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-versions); adoption figures from [Eliciting Insights' 2026 AI Adoption Survey](https://elicitinginsights.com/news/health-systems-accelerate-ai-adoption-with-67-increase-in-multi-solution-deployment-2026/).*

*Related: [AI Agents Do Licensed Work. Liability Law Doesn't Fit.](/blog/ai-agent-liability-clinical-legal-professional-doctrines) — the ban-versus-license split in detail, and why three liability doctrines all miss.*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
