The Short Answer
A Nature Medicine Comment published 10 September 2026 reports that Beijing Tsinghua Changgung Hospital's AI-TEC agent clinic was used in just 41 of 1,113 examinations — 3.8% — in one month, rising to 23% after the team cut clicks and manual entry. Diagnostic performance reached about 0.94 AUROC only once 1,426 expert-reviewed images were added. With ibl.ai you own all the code and the data.
The agents did not get smarter between those two months. The environment around them changed.
What did the Nature Medicine AI-TEC paper actually report, and when?
It is a Comment published on 10 September 2026, titled "Initial lessons from real-world implementation of an AI-agent eye clinic in China."
Three points of precision, because the framing circulating this week gets them wrong.
It is a Comment, not a trial. Nature Medicine classifies it under News & Comment. It is a first-person implementation report from the team that built the system — no control arm, no randomization, no patient-outcome endpoint. That is a legitimate and useful genre, and it is not a clinical study.
It is not a robot clinic. AI-TEC — the AI-Agent Augmented Tsinghua Eye Clinic — is a set of specialist agents spanning pre-consultation history-taking, examination triage, image analysis, decision support, education and follow-up, with ophthalmologists in the loop throughout.
The deployment is not new. The agents have been integrated into ophthalmology consultations at Beijing Tsinghua Changgung Hospital since November 2025. The paper is a week old; the experience behind it is about ten months old.
The authors are from Tsinghua's Department of Automation and the Beijing Visual Science and Translational Eye Research Institute, with Tien Yin Wong, Qionghai Dai, Ya Xing Wang and Jiamin Wu as corresponding authors.
Why did clinician use of the AI eye clinic fall to 3.8% of examinations?
Because using it cost clinicians time, and the pilot's own design made that cost visible.
ScienceAlert's account of the paper gives the numbers: 41 of 1,113 examinations — 3.8% — used the available AI-TEC processes in one month five months into the deployment.
The following month, after the team made the system faster with fewer clicks and less manual input, it was 259 of 1,126 examinations, or 23%.
No new model shipped between those two months. The six-fold change came from removing friction.
The authors' own supplementary material explains part of the cost. In the pilot workflow, clinicians first record a diagnosis with no access to AI output, then review the agent's prediction, then state whether it would change their decision, then record a final diagnosis.
That design is rigorous — it measures decision impact, not just algorithmic accuracy. It is also four steps where routine practice has one.
Reported clinician feedback also arrived weeks after the encounter, which means the loop that should have improved the system ran slower than the clinic did.
Did more training data make the eye-clinic model better?
No. Better data did, and the gap between those two is the most transferable finding in the paper.
The system was originally trained on almost 27,000 images drawn from routine care — lower quality and inconsistently labeled. Per ScienceAlert, adding 1,426 expert-reviewed, correctly labeled scans outperformed that larger set.
Reported diagnostic performance rose from about 0.80 to 0.94 AUROC after clinician-verified labels were used, as Medical Daily summarizes it.
This is the correction the "the model worked, the environment didn't" framing needs. The model did not arrive working.
It reached clinical-grade separation only after the institution's own experts curated a small, high-quality, correctly labeled set from inside the clinic — which is a data-layer capability, not a model purchase.
Records generated during routine care are not organized the way data prepared for an AI experiment is. That is true of every hospital, not just this one.
What did AI-TEC's own error analysis say the model was missing?
Context about the patient, in at least one named case.
The supplementary figure attributes misidentified cases to several factors. Severe diabetic retinopathy lesions were "difficult to distinguish from retinal vein occlusion without diabetic history."
Optic disc reflection impaired glaucoma detection. Cataract detection depended on model-specific thresholds.
Read that first one carefully. The image was ambiguous, and the fact that would have disambiguated it — whether this patient has diabetes — exists, in the chart, in another system.
That is not a model-capacity failure. It is a retrieval failure at the data layer, appearing in the error table of a deployed clinical system.
The authors also note that discordance with ground truth did not always mean the model was wrong: some cases involved comorbid or secondary diagnoses requiring clinical correlation. Which is itself an argument that the surrounding record, not the image alone, decides the answer.
What should a health system take from AI-TEC before buying another model?
Four things, none of which is a procurement decision about a model.
- Adoption is an interface property. Two deployments of identical agents differed 3.8% to 23% on clicks and manual entry. A pilot that does not measure use is not measuring anything.
- Local expert-labeled data beats bulk data. 1,426 curated images outperformed nearly 27,000 routine ones. That capability has to live inside the institution, because the experts and the images do.
- Failure modes are often missing facts. When the disambiguating detail sits in a different system, better weights do not help.
- Accuracy metrics do not measure clinical value. The authors are explicit that the transition to AI-native care requires workflow integration, clinician engagement and measurable clinical value — three things a benchmark score does not report.
This is the same conclusion reached from a different direction in healthcare AI's bottleneck was never the model, which argued it from a simulation result and a fragmented record. AI-TEC is the measured version, in a live clinic, with the failure modes named.
It also sharpens the clinician-adoption question: trust matters, and so does the number of clicks between a clinician and the answer.
How does ibl.ai put the agent layer inside a health system's own perimeter?
By running the platform where the record already is, and leaving the parts that determine adoption in the institution's hands.
With ibl.ai you own all the code and the data.
The stack runs on the health system's own infrastructure with full source code access, so protected health information stays inside the perimeter. It is model-agnostic across any LLM, so the institution can switch providers without rewriting the platform.
Billing is usage-based with no per-seat pricing, and you can deploy anywhere — your own cloud, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.
Agents read EHR and ancillary systems in place over HL7 FHIR under role-scoped permissions, with every access audited and access control bound to the institution's existing identity provider.
That matters for exactly the failure AI-TEC reported: an agent that can retrieve a diabetic history at the moment it reads the fundus photograph is answering a different question than one that cannot.
And because the workflow, the connectors and the fine-tuning data are yours, the friction that decided 3.8% versus 23% is something you can fix without waiting for a vendor release. 1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
ibl.ai is family-owned and operated from New York, NY.
Related reading: healthcare AI's bottleneck was never the model — the same argument from a simulation result and a fragmented record, and Sheba is rolling out ChatGPT — the data layer decides — what a literature assistant leaves untouched underneath.
Sources: the article type, title, 10 September 2026 publication date, authorship and the "workflow integration, clinician engagement and measurable clinical value" framing from the Nature Medicine Comment; the pilot workflow and the error attributions from the authors' supplementary information; the 41-of-1,113 and 259-of-1,126 counts and the 27,000 / 1,426 image figures from ScienceAlert; the November 2025 integration date and the 0.80-to-0.94 AUROC figures from Medical Daily. The Comment itself is subscription-only; the counts are quoted as those outlets report them.