ibl.ai Agentic AI Blog

Insights on building and deploying agentic AI systems. Our blog covers AI agent architectures, LLM infrastructure, MCP servers, enterprise deployment strategies, and real-world implementation guides. Whether you are a developer building AI agents, a CTO evaluating agentic platforms, or a technical leader driving AI adoption, you will find practical guidance here.

Topics We Cover

Featured Research and Reports

We analyze key research from leading institutions and labs including Google DeepMind, Anthropic, OpenAI, Meta AI, McKinsey, and the World Economic Forum. Our content includes detailed analysis of reports on AI agents, foundation models, and enterprise AI strategy.

For Technical Leaders

CTOs, engineering leads, and AI architects turn to our blog for guidance on agent orchestration, model evaluation, infrastructure planning, and building production-ready AI systems. We provide frameworks for responsible AI deployment that balance capability with safety and reliability.

Back to Blog

Duplex Voice Agents Reach the Payer Phone Call. The $31 Billion Figure Is From 2009.

ibl.ai EngineeringOctober 3, 2026
Premium

Decagon launched Voice 3 and its Chord speech model on 1 October 2026, bringing duplex voice agents that talk while they work. AI voice agents have called payers since 2019, so the news is the removal of hold time, not the arrival of the capability. Meanwhile the $31 billion prior authorization figure everyone quotes is a 2009 estimate covering every physician-plan interaction. The deployment question is control and evidence, not feasibility.

The Short Answer

Duplex voice agents that listen while they speak are now shipping, and prior authorization is the obvious healthcare target. But AI agents have called payers since 2019, so duplex removes hold time rather than enabling the workload. And the $31 billion figure circulating with these announcements is a 2009 estimate of all physician-plan interactions. On ibl.ai you own all the code and the data.

The technology is real. The business case is real. The number almost everyone is quoting is not, and neither is the novelty.

Where does the "$31 billion prior authorization" figure actually come from?

From a study published in 2009, measuring something else.

Lawrence Casalino and colleagues published What Does It Cost Physician Practices To Interact With Health Insurance Plans? in Health Affairs 28, no. 4.

Surveying nearly 900 physicians and medical group administrators, they found physicians spent 142 hours a year on health-plan interactions, costing practices $31 billion annually, or $68,274 per physician per year.

Two things get lost when that becomes "prior authorization costs U.S. healthcare $31B."

It is 17 years old. And the study's own scope is every interaction a practice has with a health plan: prior authorization, pharmaceutical formularies, claims, credentialing, contracting, and quality data reporting, together.

It is also the cost to physician practices, not to the U.S. healthcare system, which is a different and much larger denominator.

The best-sourced recent figure is narrower in scope and larger in size.

Howell, Yin and Robinson's Quantifying The Economic Burden Of Drug Utilization Management in Health Affairs puts the annual cost of drug utilization management at $93.3 billion: $6.0 billion for payers, $24.8 billion for manufacturers, $26.7 billion in physician time, and $35.8 billion borne by patients.

The authors call that a lower bound, because parts of the cycle could not be quantified.

None of this weakens the case for automating the work. It just means the case should be made with the physician-time figures, which are current, specific, and attributable.

What do the current prior authorization time figures actually say?

The AMA's 2025 Prior Authorization Physician Survey is the live source, drawn from 1,000 practicing physicians (400 primary care, 600 specialists).

Its headline findings:

FindingFigure
Time spent on prior authorization each week13 hours of physician and staff time
Prior authorizations completed per physician per week40
Physicians with staff working exclusively on prior auth40%
Physicians saying it contributes to burnout94%
Physicians reporting denials rose over five years74%
Physicians agreeing medical-necessity denials get qualified clinician review24%
Physicians who always appeal an adverse decision21%

One correction worth making precisely, because it changes the math: the 13 hours is physician and staff time combined, not 13 hours per physician. Briefs routinely restate it as per-physician, which inflates the labour pool by a large multiple.

The 40 prior authorizations per week is the per-physician figure. The AMA also reports phone as the most commonly used method for completing them, which is why the channel matters.

A fuller treatment of where that time actually goes is in Prior Authorization and the Time-to-Patient Bottleneck.

What did Decagon's Voice 3 and Chord launch with?

Decagon introduced Voice 3 and Chord at its Dialogues event on 1 October 2026, per the company's Dialogues 2026 announcement.

Chord is, in Decagon's words, "the first voice model from Decagon Labs, post-trained specifically for customer conversations."

Two properties matter for the healthcare case.

The duplex architecture means the agent processes incoming audio while it is speaking. It talks through a listener saying "mhm" but yields on a real interruption.

And conversation runs in parallel with task execution, so the agent narrates what it is doing during a long-running lookup instead of placing the caller on hold.

Both matter on a payer line, which is mostly waiting, transferring, and being interrupted. A turn-taking agent handles that by going quiet while it works; a duplex one does not have to.

Decagon's Voice product page states support for 70+ languages with automatic detection and language switching, on the same page that says the product is built on Chord.

Note that the figure is published for Decagon Voice generally; neither the Dialogues announcement nor the press release attaches a language count to Voice 3 specifically.

Decagon builds for customer experience, not healthcare specifically, and nothing here is a claim about HIPAA posture. The capability is what is newly available; where it runs is a separate decision, and the rest of this post is about that decision.

Has anyone actually run voice agents on payer calls before?

Yes, for years, and any post treating this as a new capability is wrong.

Infinitus Systems, founded in 2019, runs AI voice agents that call commercial and government payors and pharmacy benefit managers. Its own description: the agent "knows which number to call, can navigate IVR, ask necessary questions, push back on or correct bad data."

Among the roughly 150 data points it collects per call are authorization requirements and status. The company cites "millions of calls" and over 500 payors, and the agent "can escalate to a human operator if needed."

That last clause matters, because it kills a tempting argument. A voice agent on a payer call does not have to run without human oversight; vendors already build mid-call escalation, redaction and post-call review into the product.

So what does duplex change? Narrower than the announcements suggest, and real. Turn-taking agents must stop talking to work, which on a payer line means hold time and dropped context. Running conversation and task execution in parallel removes that.

This is a throughput improvement to an established workflow, not a new frontier.

What actually decides where a prior authorization voice agent should run?

Not feasibility, and not accuracy. What crosses the wire, and what is left behind afterwards.

A text chatbot answering patient questions can be scoped to publish-safe content. A voice agent working a prior authorization call cannot.

To do the job it must read out the member ID, the diagnosis code, the procedure code, the clinical justification, and often the patient's name and date of birth.

That is protected health information, spoken aloud, in both directions, in real time, on a line that also produces an audio recording, a transcript, and a model's reasoning trace about a clinical decision.

Those artifacts are discoverable, and in a denial appeal they are evidence. Established vendors hold them under a BAA, which is a legitimate arrangement and not the same thing as holding them yourself.

So three questions decide the deployment, and none of them are about whether it works.

  1. Whose infrastructure holds the recording and the transcript? If it is a vendor's, your BAA is doing work your perimeter should be doing.
  2. Can you produce the full action log of what the agent said on a specific call, nine months later, without filing a request? In an appeal or an audit, a vendor report is not the same artifact as a log you hold.
  3. Can you change the model without renegotiating anything? Speech models are improving on a monthly cadence, and a stack welded to one provider's voice model inherits that provider's roadmap.

The general case for keeping this class of workload inside the perimeter is in Self-Hosted AI Agents for Healthcare.

What do physicians think automation is doing to denials?

They are worried about it, and the AMA measured the worry rather than the deployment.

74% of physicians report denials have increased over the past five years, and 60% say they are concerned that augmented intelligence increases or will increase prior authorization denial rates.

That is a survey of concern. It is not evidence about what any payer currently runs, and it should not be read as one.

What the same survey does establish is a trust gap with numbers attached. Only 24% of physicians agree that medical-necessity denials are reviewed by a licensed and qualified clinician, a commitment insurers made in their 2025 reform pledge.

And only 21% always appeal an adverse decision. Among those who do not, 49% say they expect the appeal to fail based on past experience.

Put those beside the workload figures: 13 hours of combined physician and staff time each week, 40% of physicians with staff assigned to nothing else, 40 authorizations per physician per week, mostly by phone.

The result is a process where most adverse decisions are never contested, largely because contesting them costs staff time that practices do not have.

That is the gap automation on the provider side actually addresses, and it is an appeal-capacity argument rather than a hold-music one.

Which makes the procurement question sharper, not softer. If the appeal record becomes the thing a machine produces, the question is who holds it.

Our analysis of the denial side of this is in Healthcare AI, Revenue Cycle, and Prior Authorization Denials.

What should a health system check before putting a voice agent on a payer call?

Five checks, ordered by how much they tell you per hour spent.

  1. Ask where the audio lands. Not where the model runs, where the recording is written and for how long. These are frequently different vendors with different retention.
  2. Ask for a specific call's complete action log, as a test, before signing. If the answer is a dashboard rather than an export you can hold, you have a reporting relationship, not an audit trail.
  3. Price it against the real time figure, which is 13 hours of combined physician and staff time per week, not 13 physician-hours. The honest number is still large and it survives scrutiny.
  4. Decide the escalation threshold before deployment, not after. Duplex makes mid-call handoff to a human smooth, which makes it easy to never specify when it must happen.
  5. Confirm you can swap the speech model. If Chord is better this quarter and something else is better in two, that should be a configuration change.

Nothing on that list is a question about voice quality, and voice quality is what every demo is about.

Where does ibl.ai fit in healthcare voice agents?

On ibl.ai you own all the code and the data. The agent runtime, the recordings, the transcripts and the audit log deploy inside your own perimeter, model-agnostic across any LLM, with no per-seat pricing.

For this workload that matters for one narrow reason. Every control a prior authorization voice agent needs is a property of the infrastructure it runs in, and infrastructure you do not administer is infrastructure you can ask about but not configure.

Model-agnostic is the second half of it. Speech models are moving fast enough that the right posture is to be able to adopt the next one without a procurement cycle.

And because pricing is by usage rather than per seat, a workload whose whole point is to stop consuming staff headcount is not billed as though it still did.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

Want a voice agent your compliance team actually administers?

We deploy the agent runtime, the recording store and the audit trail as source code you keep, in your cloud, on-premise, GovCloud, or fully air-gapped. Book a 30-minute demo or talk to the ibl.ai team. ibl.ai is family-owned and operated from New York, NY.

Sources: Voice 3, the Chord speech model, the duplex architecture and parallel task execution are from Decagon's Dialogues 2026 announcement and its Business Wire release of 1 October 2026. Neither states a supported-language count, so the "70+ languages" figure appearing in third-party coverage is omitted here. The $31 billion, $68,274-per-physician and 142-hours figures are from Casalino et al., Health Affairs 28, no. 4, 2009, summarised by the Commonwealth Fund. The $93.3 billion drug utilization management breakdown is from Howell, Yin and Robinson, Health Affairs, 2021. The 13 hours, 40 authorizations per week, 40% exclusive-staff, 94% burnout, 74% denial-increase, 60% AI-concern, 24% clinician-review, 21% always-appeal and 49% expect-failure figures are from the AMA 2025 Prior Authorization Physician Survey, with corroborating coverage at Healthcare Dive. The prior history of voice agents on payer calls, the 500+ payors, the "millions of calls" figure, IVR navigation, authorization requirements among ~150 data points per call and human escalation are from Infinitus Systems. The 70+ languages figure is from Decagon's Voice product page.

Why does owning the AI stack matter?

ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing β€” so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.

  • You own all the code and the data

    Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform β€” the stack itself is yours.

  • Model-agnostic

    Run any LLM β€” Claude, GPT, Gemini, Llama, Command, or your own fine-tune β€” and switch providers without rewriting the platform.

  • No per-seat pricing

    Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.

  • Deploy anywhere

    Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY β€” a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

See the ibl.ai AI Operating System in Action

Discover how leading universities and organizations are transforming education with the ibl.ai AI Operating System. Explore real-world implementations from Harvard, MIT, Stanford, and users from 400+ institutions worldwide.

View Case Studies
Work with our team

Pilots, deployment, and full ownership

Most enterprise engagements are one-time, not subscriptions. You integrate ibl.ai with your own data, deploy it on your own infrastructure, and the engineering hours scale with the work β€” so the price tracks the scope, not your headcount.

Start here

Pilot

from $15K

fixed scope Β· fixed timeline

A time-boxed proof of value on your real data β€” not a slide deck.

Best for: Teams that want to see ibl.ai working before committing.

  • Deployed on your infrastructure or our cloud
  • 1–2 production agents wired to a slice of your data
  • One integration (LMS / SIS / SSO / data source)
  • Weekly working sessions with our engineers
  • Pilot fee credits toward a full engagement
Scope a pilot
Most common

Integration & Deployment

$25K – $80K

one-time Β· not a subscription

Full deployment integrated with your data and systems. Engineering hours scale with scope.

Best for: Organizations rolling ibl.ai out across a department, campus, or business unit.

  • Platform deployed in your VPC, on-prem, or air-gapped
  • Integrated with your data + identity (SSO / SAML)
  • Multiple custom agents built to your workflows
  • Engineering hours proportional to scope
  • You own the data Β· run any LLM you choose
Plan a deployment
Full ownership

Codebase Transfer + Custom AI Engineering

Custom quote

perpetual license Β· you own the stack

We transfer the full source code. You own and self-host the entire platform β€” outright.

Best for: Organizations and enterprises that benefit from perpetual ownership and sovereignty.

  • Complete source-code transfer + perpetual license
  • Dedicated AI engineering team on your roadmap
  • Custom agents, models, and integrations to spec
  • Air-gapped capable Β· zero vendor lock-in
  • Family-owned, New York–based long-term partner
Talk about ownership
You own the code and data Run any LLM β€” Claude, GPT, Gemini, Llama Family-owned & operated from New York, NY