---
title: "Hallucination Is Provably Inevitable. Stop Procuring an Error Rate."
slug: "llm-hallucination-inevitability-government-ai-procurement"
author: "ibl.ai Engineering"
date: "2026-10-02 11:00:00"
category: "Premium"
topics: "LLM hallucination, AI procurement, government AI, RAG, retrieval augmented generation, AI governance, abstention, audit trail, sovereign AI"
summary: "A line of computability-theory papers argues that LLM hallucination is inevitable rather than a defect awaiting a fix, which means a contract specifying zero hallucination is purchasing an impossibility. The same research names the alternative, and it is an architecture rather than a threshold. Procurement should specify grounding, citation, abstention and an audit trail."
banner: ""
thumbnail: ""
linkedin: |
  There is a claim going around that researchers "just proved" LLM hallucinations can never be fixed. The research is real and the date is not, so both halves are worth separating.

  The paper people are describing is "Hallucination as a Computational Boundary: A Hierarchy of Inevitability and the Oracle Escape" by Wang Xi and colleagues. It formalizes an LLM as a probabilistic Turing machine and argues hallucination is inevitable at diagonalization, incomputability and information-theoretic boundaries.

  It was submitted in August 2025 and revised in December 2025. It is not this week's news, and it is not alone: Banerjee and colleagues made a related argument using Gödel's first incompleteness theorem, and a third paper reduces hallucination detection to the halting problem.

  Now the part that gets dropped, which is the useful part. Read the title again. "The Oracle Escape."

  The same paper that argues inevitability proposes the mitigation: model retrieval as an oracle machine. Inevitable is not the same as unmitigable. "Can never be fixed" overstates what the authors wrote, in a paper whose own title offers the way out.

  Why this matters commercially: any AI contract clause that specifies a maximum hallucination rate or a zero-tolerance threshold on citations is buying something no vendor controls. A vendor who signs one has either not read the literature or has decided the clause is unenforceable.

  The specification that works is architectural, not statistical:

  → Every answer grounded in approved sources, with the retrieved passage attached
  → A real abstention path, so "I cannot answer from the record" is a success state
  → A complete audit trail of query, retrieval and response
  → Model portability, because the error floor differs per model and you will want to move

  You cannot buy a model that does not hallucinate. You can buy an architecture where a hallucination is visible, attributable and cheap to catch.

  On ibl.ai you own all the code and the data, so the retrieval layer and the audit trail are infrastructure you hold.

  #iblai #AIGovernance #GovTech #AIProcurement #RAG
---

## The Short Answer

**Several computability-theory papers argue LLM hallucination is inevitable rather than a bug awaiting a fix, so a contract specifying a hallucination rate of zero buys an impossibility. The leading paper also proposes the mitigation in its own title: retrieval modelled as an oracle. On ibl.ai you own all the code and the data, so grounding, abstention and the audit trail are yours to specify.**

Two things need separating here, because the claim circulating is stronger and newer than the research it rests on.

## What does the research actually say about hallucination being unfixable?

The paper being described is **"Hallucination as a Computational Boundary: A Hierarchy of Inevitability and the Oracle Escape"**, by Wang Xi, Quan Shi, Zenghui Ding, Jianqing Gao and Xianjun Yang.

It formalizes a large language model as a **probabilistic Turing machine** and constructs what the authors call a computational necessity hierarchy. On that basis it argues hallucination is inevitable at three boundaries: **diagonalization, incomputability, and information theory**.

It is not an isolated result. Banerjee and colleagues argued in **"LLMs Will Always Hallucinate, and We Need to Live With This"** that hallucination follows from the mathematical structure of these models, drawing on Gödel's first incompleteness theorem.

A third paper reduces hallucination *detection* to the halting problem, which is undecidable.

So the direction of the research is real, and it is a body of work rather than a single viral finding.

## Is this new research from this week?

No, and this is the first correction worth making plainly.

The paper was submitted to arXiv on **10 August 2025** and revised on **8 December 2025**. At the time of writing it is well over a year old in its first version.

Descriptions calling it new, or saying it "went viral this week", are attaching recency to something that has been in the literature for a while.

The engagement counts circulating alongside it are a measure of attention, not of evidence, and a post on X is not a source for a claim about what a paper proves.

None of this weakens the finding. It does change what you should conclude from the fact that it is being discussed now, which is that the procurement implications are finally being noticed, not that the mathematics just landed.

## Does the research really say hallucination can never be fixed?

This is the second correction, and it is the one that matters for anyone making a decision.

Read the paper's title to the end. **"The Oracle Escape."**

The same authors who argue inevitability propose the mitigation: model **retrieval-augmented generation as an oracle machine**, and continuous learning as an internalized oracle. The paper is structured as a limit *and* a way to work within it.

"Inevitable" and "unmitigable" are different claims. A system that cannot be made perfect can still be made accountable, bounded and cheap to check. That is the ordinary condition of every safety-critical system ever fielded.

So "researchers proved hallucinations can NEVER be fixed" overstates a paper that offers an escape in its own title. The defensible version is narrower and more useful: **no model-level accuracy threshold is purchasable, so stop writing contracts that specify one.**

## Why do government AI contracts specifying accuracy thresholds fail?

Because they name a quantity no vendor controls and no test can certify for the queries that matter.

The pattern to watch for, stated here as illustrative shapes rather than quotations from any particular solicitation, looks like this:

- A maximum hallucination rate, expressed as a percentage
- A minimum factual accuracy figure on domain-specific queries
- A zero-tolerance clause on citations to statute or regulation

Each one has the same three problems. The threshold is unachievable in the limit the research describes. It is measured against a benchmark that is not the agency's actual query distribution. And it places the burden of proof on a number rather than on an inspectable record.

There is a practical failure mode too. A vendor who accepts a zero-hallucination clause has either not read the literature or has decided the clause is unenforceable. Neither is the counterparty an agency wants.

A related procurement blind spot, where security certifications crowd out any test of whether the system can do the work, is in [Government AI Procurement's Blind Spot](/blog/government-ai-agent-competence-benchmarks-procurement-2026).

## What should an agency specify instead of an accuracy threshold?

Specify the architecture that makes an error visible, attributable and correctable. Four requirements, all testable at acceptance.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f8f9fa; border-bottom:2px solid #e5e7eb;">
      <th style="padding:0.75rem; text-align:left;">Requirement</th>
      <th style="padding:0.75rem; text-align:left;">How you test it at acceptance</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Grounding in approved sources</strong></td>
      <td style="padding:0.75rem;">Every response carries the retrieved passage and its source document, from a corpus the agency controls</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Citation at passage level</strong></td>
      <td style="padding:0.75rem;">A reviewer can open the cited statute or policy and read the sentence relied on, without a second search</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>An abstention path</strong></td>
      <td style="padding:0.75rem;">Ask questions the corpus cannot answer; "not answerable from the record" must be a success, not a penalty</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Complete audit trail</strong></td>
      <td style="padding:0.75rem;">Query, retrieval set, model version and response are reproducible months later, by the agency, unaided</td>
    </tr>
  </tbody>
</table>

Add a fifth that is commercial rather than technical: **model portability**. The error floor differs between models and moves with every release, so the ability to switch without rewriting the system is how an agency responds to that without re-procuring.

Abstention deserves emphasis because it is the requirement most often missing. A system rewarded only for answering will answer, and the research above says the answer will sometimes be wrong. A system permitted to decline converts an invisible error into a visible gap.

The same audit-trail argument applied to regulatory filings is in [When Compliance AI Hallucinates, Who Audits the Filing?](/blog/compliance-ai-hallucinations-audit-trail).

## Where does ibl.ai fit when hallucination cannot be eliminated?

On ibl.ai you own all the code and the data. The retrieval layer, the evaluation harness and the audit log run inside your own perimeter, model-agnostic across any LLM, with no per-seat pricing.

The reason that follows from the research rather than being bolted onto it: if hallucination is a property of the architecture, then your real control is the layer around the model. Grounding, citation, abstention and logging are all that layer.

A vendor-hosted system can offer each of those as a feature. It cannot offer them as something you administer, and at acceptance the difference is whether you can reproduce last quarter's answer yourself.

**1.6M+ users across 400+ organizations** run the platform this way, including NVIDIA, MIT, and Syracuse University.

## Want a grounded AI stack you can audit yourself?

We deploy the retrieval layer, abstention policy and audit trail as source code you keep, in your cloud, on-premise, GovCloud, or fully air-gapped. [Book a 30-minute demo](https://cal.com/iblai/30min) or [talk to the ibl.ai team](/contact). ibl.ai is family-owned and operated from New York, NY.

*Sources: the paper's title, authors, probabilistic-Turing-machine formalization, the three inevitability boundaries, the submission date of 10 August 2025, the 8 December 2025 revision and the oracle-escape mitigation are from [arXiv:2508.07334](https://arxiv.org/abs/2508.07334). The Gödel-based argument is from [arXiv:2409.05746](https://arxiv.org/abs/2409.05746). The reduction of hallucination control to the halting problem is from [arXiv:2506.06382](https://arxiv.org/abs/2506.06382).*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
