---
title: "Scaffolding Moves Accuracy 28 Points. Model Choice Is the Smaller Variable."
slug: "scaffold-effects-ai-accuracy-architecture-not-model-choice"
author: "ibl.ai Engineering"
date: "2026-10-08 09:00:00"
category: "Premium"
topics: "AI accuracy, agent scaffolding, GAIA benchmark, model selection, retrieval augmented generation, clinical AI evaluation, OpenEvidence, enterprise AI architecture, evaluations, ontology layer, elicitation gap"
summary: "A pre-registered controlled comparison published in June 2026 found that scaffold choice alone moves measured agent accuracy by as much as 28 percentage points inside a single model. The published clinical record on OpenEvidence points the same way from two directions: retrieval scaffolding drove fabricated references to zero across 4,979 citations, and a 100-question board-style pilot still capped at 41%. Benchmarks measure the scaffold as much as the model, and the scaffold is the part you own."
banner: ""
thumbnail: ""
linkedin: |
  There is a number circulating this week: a 55% accurate base model taken to 96.7% by adding five layers of scaffolding, usually credited to a clinical decision-support company. We went looking for the primary source and found it somewhere else entirely. It is Findustry AI's chargeback-response leaderboard, scoring accept-or-contest decisions on credit-card chargebacks with GPT-5.5 underneath. Not a medical benchmark. The harness run also has proprietary tool outputs injected, so it sees information the bare model never gets, which makes it a product demo rather than a controlled measurement. Here is what is actually published, which makes the point better.

  A pre-registered controlled comparison, "Scaffold Effects on GAIA" (arXiv, June 2026), held the tasks and conditions fixed and varied only the scaffold: ReAct, a planner-actor-rater multi-agent design, and planner-then-executor, across five models from three providers, three attempts per question.

  Scaffold choice alone moved measured accuracy by as much as 28 percentage points inside a single model.

  Read that again, because it is the whole procurement argument. A published capability score is a joint measurement of a model and the harness somebody wrapped around it. When you compare vendors by benchmark, you are often comparing their scaffolds and attributing the result to their models.

  The clinical record says the same thing from two directions. An npj Health Systems study of OpenEvidence examined all 4,979 references returned to 150 standardized prompts across five specialties and confirmed none as fabricated, with three abstracts carrying author attribution errors, 0.06%. Retrieval scaffolding solved the failure mode everybody feared.

  And a medRxiv pilot put the same platform against 100 board-style subspecialty questions and found a 41% ceiling for Deep Consult, 34% for Quick Consult.

  Both are true. They measure different things. Scaffolding fixed sourcing and did not fix hard multi-step reasoning, which is exactly what you would want to know before deploying either one.

  The practical conclusion for a buyer: stop shopping for a model and start owning the scaffold. The retrieval layer, the typed relationships over your own systems, the guardrails, the evaluation harness. That is where the 28 points live, and it is the asset that survives the next model release.

  With ibl.ai you own all the code and the data, run it model-agnostic across any LLM, and pay by usage with no per-seat pricing.

  #iblai #EnterpriseAI #AIEvaluation #AgenticAI #AIArchitecture #RAG
---

## The Short Answer

**Measured AI accuracy is a property of the scaffold as much as the model. A pre-registered controlled comparison found scaffold choice alone moving accuracy up to 28 percentage points inside one model. So buy the layer you can own, not the benchmark: with ibl.ai you own all the code and the data, run it model-agnostic across any LLM, and pay with no per-seat pricing.**

The question "which model is most accurate" has an uncomfortable answer. It depends on what you wrapped around it, and the wrapping moves the number further than the model swap does.

## Why does the same AI model score differently on the same benchmark?

Because the published score measures the model and its harness together, and nobody separates them.

[Scaffold Effects on GAIA: A Controlled Comparison](https://arxiv.org/abs/2606.08529), submitted to arXiv on 7 June 2026 by Jason Starace, was built to measure exactly that confound. It is pre-registered, which matters: the hypothesis was fixed before the runs.

The design holds tasks and conditions constant and varies only the scaffold. Three scaffolds were tested: ReAct, a planner-actor-rater multi-agent design, and planner-then-executor.

Five models across three providers, on GAIA validation Levels 1 and 2, with three attempts per question.

The finding, in the author's words: "Scaffold choice alone moves measured accuracy by as much as 28 percentage points within a single model."

The paper names the thing this conflation hides as the *elicitation gap*: the distance between what a model can do and what its scaffold lets it do. Published agent capability scores sit somewhere inside that gap without telling you where.

For a procurement team, that is a direct instruction. A vendor demo that beats another vendor's demo may be a better harness around a similar model, and the harness is the part you can build, buy, or own.

## What does "scaffolding" actually mean in an enterprise deployment?

It means the five things around the model that decide whether its answer is usable, none of which arrive with a model subscription.

The layers are not exotic, and they map one to one onto the architecture we build with clients:

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Scaffold layer</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">What it decides</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Who has to build it</th>
    </tr>
  </thead>
  <tbody>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Domain knowledge</strong></td>
      <td style="padding:0.75rem;">Whether the corpus the model reasons over is yours or the internet's</td>
      <td style="padding:0.75rem;">You, from your own systems of record</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Retrieval</strong></td>
      <td style="padding:0.75rem;">Whether a claim is grounded in a document that exists</td>
      <td style="padding:0.75rem;">You, over your data layer</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Task instruction</strong></td>
      <td style="padding:0.75rem;">Whether the model is doing your job or a generic one</td>
      <td style="padding:0.75rem;">Your domain experts, in plain language</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Guardrails</strong></td>
      <td style="padding:0.75rem;">What the model is not allowed to say or do</td>
      <td style="padding:0.75rem;">You, against your own policy</td>
    </tr>
    <tr style="background:#f0f9ff;">
      <td style="padding:0.75rem;"><strong>Evaluation</strong></td>
      <td style="padding:0.75rem;">Whether any of the above is working this week</td>
      <td style="padding:0.75rem;">You, on your own workloads</td>
    </tr>
  </tbody>
</table>

Four of the five columns on the right say the same word. That is the finding, restated as an org chart.

The GAIA comparison only varied the orchestration pattern and still produced a 28-point spread. Vary retrieval quality and corpus coverage as well, and the model becomes the smallest term in the expression.

## Does scaffolding fix clinical accuracy? What the published OpenEvidence record shows

It fixed the failure mode everyone was afraid of, and did not fix the one that decides deployment. Both results are published, and they are about the same product.

On sourcing, the scaffolding worked.

[Reference quality of OpenEvidence across five medical specialties](https://www.nature.com/articles/s44401-026-00142-8), published in npj Health Systems on 2 September 2026, examined all 4,979 references returned to 150 standardized prompts across oncology, cardiology, rheumatology, psychiatry and infectious diseases, 30 prompts each.

No reference was confirmed as fabricated. Three ASCO meeting abstracts carried author attribution errors, 0.06% of the set, and all three were confirmed to exist with the correct titles and years.

For anyone who watched general-purpose chatbots invent citations, that is what a working retrieval layer looks like in a number.

On hard reasoning, the same scaffolding did not carry.

[The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios](https://www.medrxiv.org/content/10.64898/2025.11.29.25341091v1.full), posted to medRxiv on 4 December 2025 by Jagarapu, Babata, Chamarthi and Hoyt, ran 100 board-style questions drawn from the MedXpertQA dataset.

Maximum accuracy was 41% for Deep Consult and 34% for Quick Consult. Inter-rater concordance was 77% and 72%, with Cohen's kappa at 0.74 and 0.69, so the grading was consistent enough to trust the ceiling.

The authors' conclusion is the operative one: low accuracy combined with inconsistent performance on complex cases argues against deployment without expert oversight.

Read together, the two studies are a specification, not a contradiction. The scaffold you build determines which failure modes you eliminate, and you only learn which ones remain by evaluating on the work you actually do.

## Where does the "55% to 96.7%" scaffolding number come from?

From a credit-card chargeback leaderboard, not from a medical benchmark, and not from the clinical platform it is usually credited to.

The claim circulating in early October 2026 is that a 55% accurate base model was taken to 96.7% purely by adding five scaffolding layers, and it is widely attached to a clinical decision-support company. That attribution is wrong.

The figures come from the [Findustry AI chargebacks performance leaderboard](https://findustryai.com/benchmark/), which scores accept-or-contest decisions on credit-card chargeback responses. The base model is GPT-5.5: **55.0%** bare, **96.7%** inside Findustry's vertical harness.

The pair also is not new. It sits in that page's findings section under a chart dated **28 May 2026**, and the current leaderboard revision does not rank GPT-5.5 at all. A figure going around as this week's news is four and a half months old.

The misattribution appears to come from a newsletter that placed the Findustry figure and a separate claim about a clinical platform in adjacent sentences. The two were then read as one.

Three caveats belong with the number even now that it has a source.

It is the vendor's own leaderboard, self-published and not independently audited, and it publishes no methodology section.

On scale it says only that building a benchmark means curating "dozens or even hundreds of test scenarios," which is a description of the craft rather than a disclosure of what these runs measured.

Most importantly, the harness run has proprietary tool outputs injected, so it sees information the bare model never receives. The two runs are therefore not the same task, which makes this a product demonstration rather than a controlled measurement of scaffolding.

That is why this post leads with the pre-registered 28-point result instead. A figure that flatters the vendor quoting it gets the primary source first, and then gets read carefully once found.

## What should a buyer measure instead of benchmark scores?

Your own workloads, continuously, with the results stored where you can audit them.

Benchmarks are useful for the thing they measure, which is relative capability under somebody else's harness on somebody else's tasks. GAIA Levels 1 and 2 are not your prior-authorization queue or your student records reconciliation.

The replacement is an evaluation harness pointed at your real traffic: [Evaluations](https://ibl.ai/docs/os/agent-settings/evals) with LLM-as-judge scoring plus human annotation and export, run against the cases your staff escalate.

Scoring judgment rather than generation is its own discipline, covered in [generation is commoditized, judgment is the new frontier](https://ibl.ai/blog/jev-judge-model-evaluation-not-generation-enterprise).

Three questions make a vendor's accuracy claim checkable. Which scaffold produced this number. Can we re-run it on our data. Do we keep the harness if we leave.

The last one decides the other two. An accuracy result you cannot reproduce in your own environment is a marketing asset, not an engineering input.

## What does owning the scaffold mean in practice?

It means the 28 points live on your side of the contract.

The model layer converges and reprices on somebody else's schedule, which is the argument in [the model is the commodity, the context layer is the moat](https://ibl.ai/blog/context-layer-is-the-moat-not-the-model).

The scaffold does not converge, because it is made of your taxonomy, your policies, your escalation rules and your evaluation set.

Concretely, on ibl.ai that scaffold is: connectors exposing your systems in place over MCP, scoped to each caller's role; an ontology of typed relationships so the model knows your SIS "student" and your CRM "contact" describe overlapping realities; skills written in plain Markdown by the people who do the work; guardrails and a sandboxed runtime; and evaluations with a full audit trail.

You own all the code and the data, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing, so you can deploy anywhere, including on-premise and fully air-gapped.

That last property is what makes the scaffold an asset rather than a dependency. When the next model lands at a lower price, you re-point the harness and keep every layer you paid to build.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

## The Edge

Every AI vendor comparison you have been handed is a comparison of scaffolds presented as a comparison of models, and the only pre-registered controlled measurement of that confound puts its size at up to 28 percentage points inside a single model.

That reframes the buying decision. The question is not which model scores highest, because the harness around it moves the score further than the swap does. The question is who ends up owning the harness, since that is where the accuracy was actually manufactured.

A model is a line item you can re-negotiate next quarter. A scaffold built from your own data, policies and evaluation sets is the only part of an AI deployment that compounds, and the only part a vendor can quietly keep.

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
