---
title: "Hospital AI Aces Single-Turn Tests. Grade the Actions"
slug: "hospital-ai-agents-single-turn-tests-grade-the-actions"
author: "ibl.ai Engineering"
date: "2026-10-09 15:00:00"
category: "Premium"
topics: "hospital AI agents, healthcare AI evaluation, multi-turn reliability, MedAgentBench, agent testing, clinical AI safety, EHR write access, HIPAA audit controls, AI governance, self-hosted healthcare AI"
summary: "In Stanford's MedAgentBench, the best overall model completed every one-step EHR task but 23.33% of tasks needing three or more steps, and scored lower on tasks that change a record than on tasks that only read one. A redesigned agent from a team including the original authors later reached 96.67% on its multi-tool-call tasks. Reliability belongs to the agent you deploy, so test it by step count and repeated runs, and gate every write to the record."
banner: "/images/blog/hospital-ai-agents-evaluation-gap-short-tests-vs-real-operations.webp"
thumbnail: "/images/blog/hospital-ai-agents-evaluation-gap-short-tests-vs-real-operations.webp"
linkedin: |
  A healthcare AI agent's one-step score says very little about a five-step workflow.

  Stanford's MedAgentBench put models inside a simulated EHR with 300 physician-written tasks. The best overall model completed 100% of one-step tasks and 23.33% of tasks needing three or more steps.

  Then a team including three of the original authors rebuilt the agent around GPT-4.1, with better tools, a plan-first prompt and a memory of past failures. It reached 96.67% on tasks needing three or more tool calls. Same benchmark, different agent.

  In the original study, half the tasks only read the record and the other half change it. The best overall model scored 85.33% on reading and 54.00% on writing.

  A separate study from Microsoft Research and Salesforce Research ran more than 200,000 simulated conversations and found an average 39% drop from single-turn to multi-turn, mostly from unreliability rather than lost capability. When a model takes a wrong turn, it tends not to recover.

  So reliability is a property of the agent you deploy, not the model's name. And the risk a hospital should test for is not only the wrong answer a clinician reads. It is the action nobody reads: a field written to the chart, a code changed, a referral sent.

  ECRI's top health technology hazard for 2026 is the misuse of AI chatbots, a hazard defined by the answers chatbots give. Agents that act are the next category, and the test for them looks different: run the same multi-step task many times and grade the end state of the record, with a human approval gate on every write.

  On ibl.ai you own all the code and the data, so the test sets, the approval gates and the conversation logs live in the hospital's own systems.

  #iblai #HealthcareAI #ClinicalAI #AIGovernance #AgenticAI #PatientSafety
---

## The Short Answer

**An AI agent's one-step score says little about multi-step hospital work: in Stanford's MedAgentBench, the best overall model completed every one-step EHR task but 23.33% of tasks needing three or more steps, and a redesigned agent later reached 96.67% on its three-plus-tool-call tasks. Test the agent you deploy and gate every write. On ibl.ai you own all the code and the data, so tests, approval gates and logs stay yours.**

This note is for CMIOs, CNIOs, compliance officers and the engineers who connect AI agents to an EHR. It uses published benchmark results and regulatory text, and it says plainly where a number comes from an older model.

## Why does a one-step score say little about multi-step healthcare AI work?

Because the best-known medical benchmarks grade answers, and most hospital work is a sequence.

The MedAgentBench v2 authors note that MedQA, PubMedQA and HealthBench test the ability to answer medical questions, not to act inside an EHR. The published evidence shows reliability falling as the sequence gets longer.

In [LLMs Get Lost In Multi-Turn Conversation](https://arxiv.org/abs/2505.06120) (May 2025), researchers from Microsoft Research and Salesforce Research compared each task given fully in one turn against the same task revealed over several turns, across six generation tasks and **more than 200,000 simulated conversations**.

Every top open- and closed-weight model they tested did worse in multi-turn settings, with an **average drop of 39%**. They traced most of it to a large increase in unreliability rather than a loss of aptitude.

Their summary is the line to remember: when a model takes a wrong turn in a conversation, it gets lost and does not recover. Prior authorization and care coordination are long conversations with systems as well as people.

[MedAgentBench](https://arxiv.org/abs/2501.14654), from Stanford and published in NEJM AI, measures a related effect inside a simulated, FHIR-compliant EHR with **300 physician-written tasks** and **100 patient profiles** holding **more than 700,000** data elements.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">MedAgentBench task length</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Claude 3.5 Sonnet v2</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">GPT-4o</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Easy (1 step)</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">100.00%</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">86.67%</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Medium (2 steps)</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">81.67%</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">70.00%</td>
    </tr>
    <tr style="background:#fff7ed; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Hard (3 or more steps)</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>23.33%</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>33.33%</strong></td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Overall</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">69.67%</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">64.00%</td>
    </tr>
  </tbody>
</table>

In the preprint's results, the best overall model was not the best on long tasks, which is itself a finding. A single headline accuracy hides the shape of the curve.

These are early-2025 models in the original, deliberately simple agent. In [MedAgentBench v2](https://psb.stanford.edu/psb-online/proceedings/psb26/chen_eric.pdf), at the Pacific Symposium on Biocomputing 2026, a team including original authors Yixing Jiang, Kameron Black and Jonathan H. Chen rebuilt the agent.

With GPT-4.1, new tools, a plan-first prompt and few-shot examples, it reached **91.0%** overall and **98.0%** with a memory of prior failures. On tasks needing three or more tool calls it reached **96.67%**, and on **300 new** multi-step tasks it scored **88.67%**.

The lesson for a hospital is not that agents are unreliable. It is that the score belongs to the agent and its harness, not the model's name, and that the v2 authors set the clinical bar at **greater than 95%** accuracy.

So measure the agent you will deploy, by task length, and measure it again every time the prompt, the tools or the model change.

## Is a wrong answer or an unauthorized action the bigger risk in hospital AI?

For an agent connected to a record, the action is the larger exposure, because a clinician reads an answer and may never see a write. MedAgentBench separates the two directly.

Half of its 300 tasks only retrieve information. The other half **modify the medical record**. The best overall model scored **85.33%** on retrieval tasks and **54.00%** on tasks that change the record. GPT-4o scored **72.00%** and **56.00%**.

The authors name Gemini 1.5 Pro and Qwen2.5 as exceptions that did better on action tasks, and their own table shows Gemini 2.0 Flash did too. The pattern still holds for most of the 12 models, and they suggest starting with use cases that only read the record.

The hazards that make headlines are about answers. ECRI ranked the [misuse of AI chatbots as the top health technology hazard for 2026](https://www.healthcaredive.com/news/ecri-health-tech-hazards-2026/810223/), citing incorrect diagnoses, unnecessary testing and invented body parts.

Those are failures a reader can catch. An agent that edits a medication field, changes a diagnosis code or sends a referral without review produces a failure that nobody reads, which is why the approval gate matters more than the accuracy score.

## How should a hospital test an AI agent before it touches a patient record?

Run the same task many times and grade the state it leaves behind, not the text it returns. Both ideas come from published agent benchmarks.

[τ-bench](https://arxiv.org/abs/2406.12045) (June 2024) grades agents by comparing the **database state at the end of a conversation** with the annotated goal state. It found that even gpt-4o succeeded on **fewer than 50%** of tasks.

Its **pass^k** metric asks whether an agent succeeds on all of k repeated trials of the same task. In the retail domain, pass^8 fell **below 25%**. That is the number a hospital should care about: does it work every time, not once.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Test</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">What a single-turn eval misses</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Score by step count</strong></td>
      <td style="padding:0.75rem;">The 100% to 23.33% drop between one-step tasks and tasks of three or more steps</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Repeat each task (pass^k)</strong></td>
      <td style="padding:0.75rem;">An agent that succeeds once and fails on the next identical run</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Grade the end state of the record</strong></td>
      <td style="padding:0.75rem;">A fluent reply sitting on top of a wrong write</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Gate every write</strong></td>
      <td style="padding:0.75rem;">The action no clinician reads</td>
    </tr>
  </tbody>
</table>

Voluntary guidance from The Joint Commission and the Coalition for Health AI, [released on September 17, 2025](https://www.aha.org/news/headline/2025-09-18-joint-commission-unveils-first-ai-guidance-partnership-coalition-health-ai) recommends AI policies, local validation and monitoring. It does not prescribe a test method, which leaves the choice to each hospital.

Simulation before go-live is a working pattern outside healthcare too. Nubank screened configurations across **more than 16,000 simulated conversations**, covered in [Nubank Screened 16,000 Simulated Chats Before Going Live](/blog/nubank-simulated-conversations-agent-testing-before-production).

## Why do healthcare AI teams treat the test harness as part of the product?

Because in billing and clinical work, a wrong output is a compliance event, and the harness is the only thing standing between a model change and production. Mature safety-critical software has always been built this way.

As of version 3.42.0 (2023), the SQLite database engine had **155.8 KSLOC** of C source and, [by its own count](https://www.sqlite.org/testing.html), **590 times as much** test code and test scripts, **92,053.1 KSLOC**. The test suite is the larger artifact by more than two orders of magnitude.

Healthcare AI teams treat the harness as part of the product for a different reason: the stakes of one output. Cainex co-founder and CTO Uriah Israel, quoted in [Anthropic's Claude Code guide for startups](https://claude.com/resources/articles/claude-code-guide-for-startups), puts it plainly: "In medical coding, a wrong code isn't a typo."

Cainex's loop has auditors review both the codes and the model's reasoning, then back-tests every candidate change across a golden set plus random samples, surfacing regressions before anything ships.

That loop only works if the hospital or the vendor can run it at will. If the agent, the prompts and the golden set sit in someone else's tenancy, the customer cannot re-test when the model underneath changes.

## What does HIPAA require a hospital to log about AI agent activity?

HIPAA does not mention AI agents, but its audit-controls standard applies to any system holding electronic protected health information. An agent that reads or writes the EHR is activity in such a system.

The Security Rule at [45 CFR 164.312(b)](https://www.law.cornell.edu/cfr/text/45/164.312) requires covered entities and business associates, including the vendors that process ePHI for them, to implement mechanisms that **record and examine activity** in information systems that contain or use ePHI.

Whether a particular agent's logs satisfy that standard is a compliance determination for each organization. What is clear is that a log you have to request from a vendor is weaker evidence than one you hold.

The same reasoning drove two recent pieces on this blog: [You Cannot Govern a Clinical Model You Cannot Observe](/blog/clinical-ai-governance-observability-hospitals-own-the-stack) and [Who Audits the AI Writing Into the Nurse's Flowsheet?](/blog/oracle-health-clinical-agent-who-audits-the-flowsheet).

## Where does ibl.ai fit in testing and gating hospital AI agents?

On ibl.ai you own all the code and the data. The agents, their prompts, the evaluation sets and the conversation history run inside the hospital's perimeter, model-agnostic across any LLM, with no per-seat pricing, so you can deploy anywhere.

[Evals](/docs/os/agent-settings/evals) runs an agent against a benchmark and scores every response, and benchmark items can be seeded from real chat traces. It grades responses; end-state and repeated-run checks against a test copy of the record are built on top of it with your engineering team.

When you switch models, you re-run the same set before anything changes for clinicians.

The [workflow builder](/docs/os/chat-canvas/workflows) lays a multi-step run out as a graph, with a **User approval** node anywhere a person must sign off before the run continues and a **Guardrails** node where output has to be screened.

How a PE-backed healthcare company approaches the data side of the same deployment is in [PE Firms Put AI Agents in Healthcare. The Data Comes First](/blog/private-equity-healthcare-ai-agents-data-comes-first). Healthcare deployments are described on [Medical and Healthcare](/solutions/medical-healthcare).

**1.6M+ users across 400+ organizations** run the platform this way, including NVIDIA, MIT, and Syracuse University.

## Want to test clinical agents on a stack your hospital owns?

We deploy agents, evaluation sets and approval workflows as source code your hospital keeps, in your cloud, on-premise, or fully air-gapped.

[Book a 30-minute demo](https://cal.com/iblai/30min) or [talk to the ibl.ai team](/contact). ibl.ai is family-owned and operated from New York, NY.

*Sources: multi-turn results from [Laban, Hayashi, Zhou and Neville, LLMs Get Lost In Multi-Turn Conversation](https://arxiv.org/abs/2505.06120); task counts, step-level and query/action results from the [MedAgentBench preprint](https://arxiv.org/abs/2501.14654), published in [NEJM AI](https://ai.nejm.org/doi/full/10.1056/AIdbp2500144); the redesigned agent, the benchmark observation and the 95% bar from [MedAgentBench v2](https://psb.stanford.edu/psb-online/proceedings/psb26/chen_eric.pdf); pass^k and end-state grading from [τ-bench](https://arxiv.org/abs/2406.12045); ECRI's 2026 hazard ranking as reported by [Healthcare Dive](https://www.healthcaredive.com/news/ecri-health-tech-hazards-2026/810223/); Joint Commission and CHAI guidance from the [American Hospital Association](https://www.aha.org/news/headline/2025-09-18-joint-commission-unveils-first-ai-guidance-partnership-coalition-health-ai); test-code figures from [SQLite](https://www.sqlite.org/testing.html); the Cainex quote and workflow from [Anthropic](https://claude.com/resources/articles/claude-code-guide-for-startups); the audit-controls standard from [45 CFR 164.312](https://www.law.cornell.edu/cfr/text/45/164.312).*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
