---
title: "Duplex Voice Agents Reach the Payer Phone Call. The $31 Billion Figure Is From 2009."
slug: "healthcare-voice-ai-agents-prior-authorization-on-premise"
author: "ibl.ai Engineering"
date: "2026-10-03 09:00:00"
category: "Premium"
topics: "healthcare AI, voice AI agents, prior authorization, revenue cycle, HIPAA, PHI, duplex voice, self-hosted AI, on-premise AI, payer denials, administrative burden"
summary: "Decagon launched Voice 3 and its Chord speech model on 1 October 2026, bringing duplex voice agents that talk while they work. AI voice agents have called payers since 2019, so the news is the removal of hold time, not the arrival of the capability. Meanwhile the $31 billion prior authorization figure everyone quotes is a 2009 estimate covering every physician-plan interaction. The deployment question is control and evidence, not feasibility."
banner: ""
thumbnail: ""
linkedin: |
  Decagon launched Voice 3 and Chord, its first in-house speech model, at Dialogues on 1 October. Duplex architecture: the agent processes incoming audio while it is speaking, and runs the task in parallel with the conversation, so it narrates progress instead of putting you on hold.

  Every healthcare post about this is quoting the same number. "Prior authorization costs U.S. healthcare $31 billion a year."

  That number is from 2009, and it is not prior authorization. Casalino et al. in Health Affairs surveyed nearly 900 physicians and administrators and costed ALL interactions a practice has with health plans: prior authorization, formularies, claims, credentialing, contracting and quality reporting together. 142 hours a year per physician, $68,274 each, $31 billion in total.

  The real figures are current and better sourced.

  The AMA's 2025 survey of 1,000 practicing physicians: 13 hours of physician and staff time every week, 40 prior authorizations per physician per week, 40% of physicians with staff doing nothing else, 94% saying it drives burnout, and phone reported as the most common method.

  One more correction, this time to the framing rather than a number. Voice AI on payer calls is not new. Infinitus has been running AI agents that dial payers, navigate IVR trees and collect authorization requirements since 2019, across more than 500 payors and millions of calls, with human escalation built in.

  So duplex does not make this possible. It removes the hold: the agent can run the lookup while it is still talking, instead of parking the call.

  Which means the real question was never feasibility. It is control and evidence.

  A prior authorization call reads out member ID, diagnosis code, procedure code and clinical justification, and it leaves behind a recording, a transcript and a reasoning trace about a clinical decision. In a denial appeal those are evidence.

  So: whose infrastructure holds them, can you export a specific call's full action log nine months later without filing a request, and can you swap the speech model when a better one ships.

  On ibl.ai you own all the code and the data, so the agent, the recording and the audit trail sit inside your own perimeter rather than behind someone's reporting portal.

  #iblai #HealthcareAI #PriorAuthorization #VoiceAI #RevenueCycle #AIGovernance
---

## The Short Answer

**Duplex voice agents that listen while they speak are now shipping, and prior authorization is the obvious healthcare target. But AI agents have called payers since 2019, so duplex removes hold time rather than enabling the workload. And the $31 billion figure circulating with these announcements is a 2009 estimate of all physician-plan interactions. On ibl.ai you own all the code and the data.**

The technology is real. The business case is real. The number almost everyone is quoting is not, and neither is the novelty.

## Where does the "$31 billion prior authorization" figure actually come from?

From a study published in 2009, measuring something else.

Lawrence Casalino and colleagues published [What Does It Cost Physician Practices To Interact With Health Insurance Plans?](https://www.healthaffairs.org/doi/10.1377/hlthaff.28.4.w533) in *Health Affairs* 28, no. 4.

Surveying nearly **900** physicians and medical group administrators, they found physicians spent **142 hours a year** on health-plan interactions, costing practices **$31 billion annually**, or **$68,274 per physician per year**.

Two things get lost when that becomes "prior authorization costs U.S. healthcare $31B."

It is **17 years old**. And the study's own scope is *every* interaction a practice has with a health plan: prior authorization, pharmaceutical formularies, claims, credentialing, contracting, and quality data reporting, together.

It is also the cost **to physician practices**, not to the U.S. healthcare system, which is a different and much larger denominator.

The best-sourced recent figure is narrower in scope and larger in size.

Howell, Yin and Robinson's [Quantifying The Economic Burden Of Drug Utilization Management](https://www.healthaffairs.org/doi/10.1377/hlthaff.2021.00036) in *Health Affairs* puts the annual cost of **drug** utilization management at **$93.3 billion**: $6.0 billion for payers, $24.8 billion for manufacturers, **$26.7 billion in physician time**, and $35.8 billion borne by patients.

The authors call that a lower bound, because parts of the cycle could not be quantified.

None of this weakens the case for automating the work. It just means the case should be made with the physician-time figures, which are current, specific, and attributable.

## What do the current prior authorization time figures actually say?

The AMA's **2025 Prior Authorization Physician Survey** is the live source, drawn from **1,000 practicing physicians** (400 primary care, 600 specialists).

Its headline findings:

<table>
<thead>
<tr><th>Finding</th><th>Figure</th></tr>
</thead>
<tbody>
<tr><td>Time spent on prior authorization each week</td><td><strong>13 hours</strong> of physician <em>and staff</em> time</td></tr>
<tr><td>Prior authorizations completed per physician per week</td><td><strong>40</strong></td></tr>
<tr><td>Physicians with staff working <em>exclusively</em> on prior auth</td><td><strong>40%</strong></td></tr>
<tr><td>Physicians saying it contributes to burnout</td><td><strong>94%</strong></td></tr>
<tr><td>Physicians reporting denials rose over five years</td><td><strong>74%</strong></td></tr>
<tr><td>Physicians agreeing medical-necessity denials get qualified clinician review</td><td><strong>24%</strong></td></tr>
<tr><td>Physicians who <em>always</em> appeal an adverse decision</td><td><strong>21%</strong></td></tr>
</tbody>
</table>

One correction worth making precisely, because it changes the math: the 13 hours is **physician and staff time combined**, not 13 hours per physician. Briefs routinely restate it as per-physician, which inflates the labour pool by a large multiple.

The 40 prior authorizations per week is the per-physician figure. The AMA also reports phone as the most commonly used method for completing them, which is why the channel matters.

A fuller treatment of where that time actually goes is in [Prior Authorization and the Time-to-Patient Bottleneck](/blog/prior-authorization-ai-agents-time-to-patient-bottleneck).

## What did Decagon's Voice 3 and Chord launch with?

Decagon introduced **Voice 3** and **Chord** at its Dialogues event on **1 October 2026**, per the company's [Dialogues 2026 announcement](https://decagon.ai/blog/dialogues-2026).

Chord is, in Decagon's words, "the first voice model from Decagon Labs, post-trained specifically for customer conversations."

Two properties matter for the healthcare case.

The **duplex architecture** means the agent processes incoming audio while it is speaking. It talks through a listener saying "mhm" but yields on a real interruption.

And conversation runs **in parallel with task execution**, so the agent narrates what it is doing during a long-running lookup instead of placing the caller on hold.

Both matter on a payer line, which is mostly waiting, transferring, and being interrupted. A turn-taking agent handles that by going quiet while it works; a duplex one does not have to.

Decagon's [Voice product page](https://decagon.ai/product/voice) states support for **70+ languages** with automatic detection and language switching, on the same page that says the product is built on Chord.

Note that the figure is published for Decagon Voice generally; neither the Dialogues announcement nor the press release attaches a language count to Voice 3 specifically.

Decagon builds for customer experience, not healthcare specifically, and nothing here is a claim about HIPAA posture. The capability is what is newly available; where it runs is a separate decision, and the rest of this post is about that decision.

## Has anyone actually run voice agents on payer calls before?

Yes, for years, and any post treating this as a new capability is wrong.

Infinitus Systems, founded in 2019, runs AI voice agents that call commercial and government payors and pharmacy benefit managers. Its own description: the agent "knows which number to call, can navigate IVR, ask necessary questions, push back on or correct bad data."

Among the roughly 150 data points it collects per call are **authorization requirements and status**. The company cites **"millions of calls"** and **over 500 payors**, and the agent "can escalate to a human operator if needed."

That last clause matters, because it kills a tempting argument. A voice agent on a payer call does not have to run without human oversight; vendors already build mid-call escalation, redaction and post-call review into the product.

So what does duplex change? Narrower than the announcements suggest, and real. Turn-taking agents must stop talking to work, which on a payer line means hold time and dropped context. Running conversation and task execution in parallel removes that.

This is a throughput improvement to an established workflow, not a new frontier.

## What actually decides where a prior authorization voice agent should run?

Not feasibility, and not accuracy. What crosses the wire, and what is left behind afterwards.

A text chatbot answering patient questions can be scoped to publish-safe content. A voice agent working a prior authorization call cannot.

To do the job it must read out the member ID, the diagnosis code, the procedure code, the clinical justification, and often the patient's name and date of birth.

That is protected health information, spoken aloud, in **both** directions, in real time, on a line that also produces an audio recording, a transcript, and a model's reasoning trace about a clinical decision.

Those artifacts are discoverable, and in a denial appeal they are evidence. Established vendors hold them under a BAA, which is a legitimate arrangement and not the same thing as holding them yourself.

So three questions decide the deployment, and none of them are about whether it works.

1. **Whose infrastructure holds the recording and the transcript?** If it is a vendor's, your BAA is doing work your perimeter should be doing.
2. **Can you produce the full action log of what the agent said on a specific call, nine months later, without filing a request?** In an appeal or an audit, a vendor report is not the same artifact as a log you hold.
3. **Can you change the model without renegotiating anything?** Speech models are improving on a monthly cadence, and a stack welded to one provider's voice model inherits that provider's roadmap.

The general case for keeping this class of workload inside the perimeter is in [Self-Hosted AI Agents for Healthcare](/blog/self-hosted-ai-agents-for-healthcare).

## What do physicians think automation is doing to denials?

They are worried about it, and the AMA measured the worry rather than the deployment.

**74%** of physicians report denials have increased over the past five years, and **60%** say they are concerned that augmented intelligence increases or will increase prior authorization denial rates.

That is a survey of concern. It is not evidence about what any payer currently runs, and it should not be read as one.

What the same survey does establish is a trust gap with numbers attached. Only **24%** of physicians agree that medical-necessity denials are reviewed by a licensed and qualified clinician, a commitment insurers made in their 2025 reform pledge.

And only **21%** always appeal an adverse decision. Among those who do not, **49%** say they expect the appeal to fail based on past experience.

Put those beside the workload figures: **13 hours** of combined physician and staff time each week, **40%** of physicians with staff assigned to nothing else, **40** authorizations per physician per week, mostly by phone.

The result is a process where most adverse decisions are never contested, largely because contesting them costs staff time that practices do not have.

That is the gap automation on the provider side actually addresses, and it is an appeal-capacity argument rather than a hold-music one.

Which makes the procurement question sharper, not softer. If the appeal record becomes the thing a machine produces, the question is who holds it.

Our analysis of the denial side of this is in [Healthcare AI, Revenue Cycle, and Prior Authorization Denials](/blog/healthcare-ai-revenue-cycle-prior-authorization-denials).

## What should a health system check before putting a voice agent on a payer call?

Five checks, ordered by how much they tell you per hour spent.

1. **Ask where the audio lands.** Not where the model runs, where the recording is written and for how long. These are frequently different vendors with different retention.
2. **Ask for a specific call's complete action log**, as a test, before signing. If the answer is a dashboard rather than an export you can hold, you have a reporting relationship, not an audit trail.
3. **Price it against the real time figure**, which is 13 hours of combined physician and staff time per week, not 13 physician-hours. The honest number is still large and it survives scrutiny.
4. **Decide the escalation threshold before deployment, not after.** Duplex makes mid-call handoff to a human smooth, which makes it easy to never specify when it must happen.
5. **Confirm you can swap the speech model.** If Chord is better this quarter and something else is better in two, that should be a configuration change.

Nothing on that list is a question about voice quality, and voice quality is what every demo is about.

## Where does ibl.ai fit in healthcare voice agents?

On ibl.ai you own all the code and the data. The agent runtime, the recordings, the transcripts and the audit log deploy inside your own perimeter, model-agnostic across any LLM, with no per-seat pricing.

For this workload that matters for one narrow reason. Every control a prior authorization voice agent needs is a property of the infrastructure it runs in, and infrastructure you do not administer is infrastructure you can ask about but not configure.

Model-agnostic is the second half of it. Speech models are moving fast enough that the right posture is to be able to adopt the next one without a procurement cycle.

And because pricing is by usage rather than per seat, a workload whose whole point is to stop consuming staff headcount is not billed as though it still did.

**1.6M+ users across 400+ organizations** run the platform this way, including NVIDIA, MIT, and Syracuse University.

## Want a voice agent your compliance team actually administers?

We deploy the agent runtime, the recording store and the audit trail as source code you keep, in your cloud, on-premise, GovCloud, or fully air-gapped. [Book a 30-minute demo](https://cal.com/iblai/30min) or [talk to the ibl.ai team](/contact). ibl.ai is family-owned and operated from New York, NY.

*Sources: Voice 3, the Chord speech model, the duplex architecture and parallel task execution are from [Decagon's Dialogues 2026 announcement](https://decagon.ai/blog/dialogues-2026) and its [Business Wire release](https://www.financialcontent.com/article/bizwire-2026-10-1-decagon-unveils-ai-for-the-era-of-personal-agents-at-dialogues) of 1 October 2026. Neither states a supported-language count, so the "70+ languages" figure appearing in [third-party coverage](https://www.unite.ai/decagon-introduces-voice-3-agent-with-chord-speech-model/) is omitted here. The $31 billion, $68,274-per-physician and 142-hours figures are from [Casalino et al., Health Affairs 28, no. 4, 2009](https://www.healthaffairs.org/doi/10.1377/hlthaff.28.4.w533), summarised by [the Commonwealth Fund](https://www.commonwealthfund.org/publications/journal-article/2009/may/what-does-it-cost-physician-practices-interact-health). The $93.3 billion drug utilization management breakdown is from [Howell, Yin and Robinson, Health Affairs, 2021](https://www.healthaffairs.org/doi/10.1377/hlthaff.2021.00036). The 13 hours, 40 authorizations per week, 40% exclusive-staff, 94% burnout, 74% denial-increase, 60% AI-concern, 24% clinician-review, 21% always-appeal and 49% expect-failure figures are from the [AMA 2025 Prior Authorization Physician Survey](https://www.ama-assn.org/system/files/prior-authorization-survey.pdf), with corroborating coverage at [Healthcare Dive](https://www.healthcaredive.com/news/physicians-skeptical-insurer-pledges-reform-prior-authorization-ama/820126/). The prior history of voice agents on payer calls, the 500+ payors, the "millions of calls" figure, IVR navigation, authorization requirements among ~150 data points per call and human escalation are from [Infinitus Systems](https://www.infinitus.ai/solutions/benefit-verification/). The 70+ languages figure is from [Decagon's Voice product page](https://decagon.ai/product/voice).*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
