The Short Answer
An AI agent's one-step score says little about multi-step hospital work: in Stanford's MedAgentBench, the best overall model completed every one-step EHR task but 23.33% of tasks needing three or more steps, and a redesigned agent later reached 96.67% on its three-plus-tool-call tasks. Test the agent you deploy and gate every write. On ibl.ai you own all the code and the data, so tests, approval gates and logs stay yours.
This note is for CMIOs, CNIOs, compliance officers and the engineers who connect AI agents to an EHR. It uses published benchmark results and regulatory text, and it says plainly where a number comes from an older model.
Why does a one-step score say little about multi-step healthcare AI work?
Because the best-known medical benchmarks grade answers, and most hospital work is a sequence.
The MedAgentBench v2 authors note that MedQA, PubMedQA and HealthBench test the ability to answer medical questions, not to act inside an EHR. The published evidence shows reliability falling as the sequence gets longer.
In LLMs Get Lost In Multi-Turn Conversation (May 2025), researchers from Microsoft Research and Salesforce Research compared each task given fully in one turn against the same task revealed over several turns, across six generation tasks and more than 200,000 simulated conversations.
Every top open- and closed-weight model they tested did worse in multi-turn settings, with an average drop of 39%. They traced most of it to a large increase in unreliability rather than a loss of aptitude.
Their summary is the line to remember: when a model takes a wrong turn in a conversation, it gets lost and does not recover. Prior authorization and care coordination are long conversations with systems as well as people.
MedAgentBench, from Stanford and published in NEJM AI, measures a related effect inside a simulated, FHIR-compliant EHR with 300 physician-written tasks and 100 patient profiles holding more than 700,000 data elements.
| MedAgentBench task length | Claude 3.5 Sonnet v2 | GPT-4o |
|---|---|---|
| Easy (1 step) | 100.00% | 86.67% |
| Medium (2 steps) | 81.67% | 70.00% |
| Hard (3 or more steps) | 23.33% | 33.33% |
| Overall | 69.67% | 64.00% |
In the preprint's results, the best overall model was not the best on long tasks, which is itself a finding. A single headline accuracy hides the shape of the curve.
These are early-2025 models in the original, deliberately simple agent. In MedAgentBench v2, at the Pacific Symposium on Biocomputing 2026, a team including original authors Yixing Jiang, Kameron Black and Jonathan H. Chen rebuilt the agent.
With GPT-4.1, new tools, a plan-first prompt and few-shot examples, it reached 91.0% overall and 98.0% with a memory of prior failures. On tasks needing three or more tool calls it reached 96.67%, and on 300 new multi-step tasks it scored 88.67%.
The lesson for a hospital is not that agents are unreliable. It is that the score belongs to the agent and its harness, not the model's name, and that the v2 authors set the clinical bar at greater than 95% accuracy.
So measure the agent you will deploy, by task length, and measure it again every time the prompt, the tools or the model change.
Is a wrong answer or an unauthorized action the bigger risk in hospital AI?
For an agent connected to a record, the action is the larger exposure, because a clinician reads an answer and may never see a write. MedAgentBench separates the two directly.
Half of its 300 tasks only retrieve information. The other half modify the medical record. The best overall model scored 85.33% on retrieval tasks and 54.00% on tasks that change the record. GPT-4o scored 72.00% and 56.00%.
The authors name Gemini 1.5 Pro and Qwen2.5 as exceptions that did better on action tasks, and their own table shows Gemini 2.0 Flash did too. The pattern still holds for most of the 12 models, and they suggest starting with use cases that only read the record.
The hazards that make headlines are about answers. ECRI ranked the misuse of AI chatbots as the top health technology hazard for 2026, citing incorrect diagnoses, unnecessary testing and invented body parts.
Those are failures a reader can catch. An agent that edits a medication field, changes a diagnosis code or sends a referral without review produces a failure that nobody reads, which is why the approval gate matters more than the accuracy score.
How should a hospital test an AI agent before it touches a patient record?
Run the same task many times and grade the state it leaves behind, not the text it returns. Both ideas come from published agent benchmarks.
Ο-bench (June 2024) grades agents by comparing the database state at the end of a conversation with the annotated goal state. It found that even gpt-4o succeeded on fewer than 50% of tasks.
Its pass^k metric asks whether an agent succeeds on all of k repeated trials of the same task. In the retail domain, pass^8 fell below 25%. That is the number a hospital should care about: does it work every time, not once.
| Test | What a single-turn eval misses |
|---|---|
| Score by step count | The 100% to 23.33% drop between one-step tasks and tasks of three or more steps |
| Repeat each task (pass^k) | An agent that succeeds once and fails on the next identical run |
| Grade the end state of the record | A fluent reply sitting on top of a wrong write |
| Gate every write | The action no clinician reads |
Voluntary guidance from The Joint Commission and the Coalition for Health AI, released on September 17, 2025 recommends AI policies, local validation and monitoring. It does not prescribe a test method, which leaves the choice to each hospital.
Simulation before go-live is a working pattern outside healthcare too. Nubank screened configurations across more than 16,000 simulated conversations, covered in Nubank Screened 16,000 Simulated Chats Before Going Live.
Why do healthcare AI teams treat the test harness as part of the product?
Because in billing and clinical work, a wrong output is a compliance event, and the harness is the only thing standing between a model change and production. Mature safety-critical software has always been built this way.
As of version 3.42.0 (2023), the SQLite database engine had 155.8 KSLOC of C source and, by its own count, 590 times as much test code and test scripts, 92,053.1 KSLOC. The test suite is the larger artifact by more than two orders of magnitude.
Healthcare AI teams treat the harness as part of the product for a different reason: the stakes of one output. Cainex co-founder and CTO Uriah Israel, quoted in Anthropic's Claude Code guide for startups, puts it plainly: "In medical coding, a wrong code isn't a typo."
Cainex's loop has auditors review both the codes and the model's reasoning, then back-tests every candidate change across a golden set plus random samples, surfacing regressions before anything ships.
That loop only works if the hospital or the vendor can run it at will. If the agent, the prompts and the golden set sit in someone else's tenancy, the customer cannot re-test when the model underneath changes.
What does HIPAA require a hospital to log about AI agent activity?
HIPAA does not mention AI agents, but its audit-controls standard applies to any system holding electronic protected health information. An agent that reads or writes the EHR is activity in such a system.
The Security Rule at 45 CFR 164.312(b) requires covered entities and business associates, including the vendors that process ePHI for them, to implement mechanisms that record and examine activity in information systems that contain or use ePHI.
Whether a particular agent's logs satisfy that standard is a compliance determination for each organization. What is clear is that a log you have to request from a vendor is weaker evidence than one you hold.
The same reasoning drove two recent pieces on this blog: You Cannot Govern a Clinical Model You Cannot Observe and Who Audits the AI Writing Into the Nurse's Flowsheet?.
Where does ibl.ai fit in testing and gating hospital AI agents?
On ibl.ai you own all the code and the data. The agents, their prompts, the evaluation sets and the conversation history run inside the hospital's perimeter, model-agnostic across any LLM, with no per-seat pricing, so you can deploy anywhere.
Evals runs an agent against a benchmark and scores every response, and benchmark items can be seeded from real chat traces. It grades responses; end-state and repeated-run checks against a test copy of the record are built on top of it with your engineering team.
When you switch models, you re-run the same set before anything changes for clinicians.
The workflow builder lays a multi-step run out as a graph, with a User approval node anywhere a person must sign off before the run continues and a Guardrails node where output has to be screened.
How a PE-backed healthcare company approaches the data side of the same deployment is in PE Firms Put AI Agents in Healthcare. The Data Comes First. Healthcare deployments are described on Medical and Healthcare.
1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
Want to test clinical agents on a stack your hospital owns?
We deploy agents, evaluation sets and approval workflows as source code your hospital keeps, in your cloud, on-premise, or fully air-gapped.
Book a 30-minute demo or talk to the ibl.ai team. ibl.ai is family-owned and operated from New York, NY.
Sources: multi-turn results from Laban, Hayashi, Zhou and Neville, LLMs Get Lost In Multi-Turn Conversation; task counts, step-level and query/action results from the MedAgentBench preprint, published in NEJM AI; the redesigned agent, the benchmark observation and the 95% bar from MedAgentBench v2; pass^k and end-state grading from Ο-bench; ECRI's 2026 hazard ranking as reported by Healthcare Dive; Joint Commission and CHAI guidance from the American Hospital Association; test-code figures from SQLite; the Cainex quote and workflow from Anthropic; the audit-controls standard from 45 CFR 164.312.
