The Short Answer
Measured AI accuracy is a property of the scaffold as much as the model. A pre-registered controlled comparison found scaffold choice alone moving accuracy up to 28 percentage points inside one model. So buy the layer you can own, not the benchmark: with ibl.ai you own all the code and the data, run it model-agnostic across any LLM, and pay with no per-seat pricing.
The question "which model is most accurate" has an uncomfortable answer. It depends on what you wrapped around it, and the wrapping moves the number further than the model swap does.
Why does the same AI model score differently on the same benchmark?
Because the published score measures the model and its harness together, and nobody separates them.
Scaffold Effects on GAIA: A Controlled Comparison, submitted to arXiv on 7 June 2026 by Jason Starace, was built to measure exactly that confound. It is pre-registered, which matters: the hypothesis was fixed before the runs.
The design holds tasks and conditions constant and varies only the scaffold. Three scaffolds were tested: ReAct, a planner-actor-rater multi-agent design, and planner-then-executor.
Five models across three providers, on GAIA validation Levels 1 and 2, with three attempts per question.
The finding, in the author's words: "Scaffold choice alone moves measured accuracy by as much as 28 percentage points within a single model."
The paper names the thing this conflation hides as the elicitation gap: the distance between what a model can do and what its scaffold lets it do. Published agent capability scores sit somewhere inside that gap without telling you where.
For a procurement team, that is a direct instruction. A vendor demo that beats another vendor's demo may be a better harness around a similar model, and the harness is the part you can build, buy, or own.
What does "scaffolding" actually mean in an enterprise deployment?
It means the five things around the model that decide whether its answer is usable, none of which arrive with a model subscription.
The layers are not exotic, and they map one to one onto the architecture we build with clients:
| Scaffold layer | What it decides | Who has to build it |
|---|---|---|
| Domain knowledge | Whether the corpus the model reasons over is yours or the internet's | You, from your own systems of record |
| Retrieval | Whether a claim is grounded in a document that exists | You, over your data layer |
| Task instruction | Whether the model is doing your job or a generic one | Your domain experts, in plain language |
| Guardrails | What the model is not allowed to say or do | You, against your own policy |
| Evaluation | Whether any of the above is working this week | You, on your own workloads |
Four of the five columns on the right say the same word. That is the finding, restated as an org chart.
The GAIA comparison only varied the orchestration pattern and still produced a 28-point spread. Vary retrieval quality and corpus coverage as well, and the model becomes the smallest term in the expression.
Does scaffolding fix clinical accuracy? What the published OpenEvidence record shows
It fixed the failure mode everyone was afraid of, and did not fix the one that decides deployment. Both results are published, and they are about the same product.
On sourcing, the scaffolding worked.
Reference quality of OpenEvidence across five medical specialties, published in npj Health Systems on 2 September 2026, examined all 4,979 references returned to 150 standardized prompts across oncology, cardiology, rheumatology, psychiatry and infectious diseases, 30 prompts each.
No reference was confirmed as fabricated. Three ASCO meeting abstracts carried author attribution errors, 0.06% of the set, and all three were confirmed to exist with the correct titles and years.
For anyone who watched general-purpose chatbots invent citations, that is what a working retrieval layer looks like in a number.
On hard reasoning, the same scaffolding did not carry.
The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios, posted to medRxiv on 4 December 2025 by Jagarapu, Babata, Chamarthi and Hoyt, ran 100 board-style questions drawn from the MedXpertQA dataset.
Maximum accuracy was 41% for Deep Consult and 34% for Quick Consult. Inter-rater concordance was 77% and 72%, with Cohen's kappa at 0.74 and 0.69, so the grading was consistent enough to trust the ceiling.
The authors' conclusion is the operative one: low accuracy combined with inconsistent performance on complex cases argues against deployment without expert oversight.
Read together, the two studies are a specification, not a contradiction. The scaffold you build determines which failure modes you eliminate, and you only learn which ones remain by evaluating on the work you actually do.
Where does the "55% to 96.7%" scaffolding number come from?
From a credit-card chargeback leaderboard, not from a medical benchmark, and not from the clinical platform it is usually credited to.
The claim circulating in early October 2026 is that a 55% accurate base model was taken to 96.7% purely by adding five scaffolding layers, and it is widely attached to a clinical decision-support company. That attribution is wrong.
The figures come from the Findustry AI chargebacks performance leaderboard, which scores accept-or-contest decisions on credit-card chargeback responses. The base model is GPT-5.5: 55.0% bare, 96.7% inside Findustry's vertical harness.
The pair also is not new. It sits in that page's findings section under a chart dated 28 May 2026, and the current leaderboard revision does not rank GPT-5.5 at all. A figure going around as this week's news is four and a half months old.
The misattribution appears to come from a newsletter that placed the Findustry figure and a separate claim about a clinical platform in adjacent sentences. The two were then read as one.
Three caveats belong with the number even now that it has a source.
It is the vendor's own leaderboard, self-published and not independently audited, and it publishes no methodology section.
On scale it says only that building a benchmark means curating "dozens or even hundreds of test scenarios," which is a description of the craft rather than a disclosure of what these runs measured.
Most importantly, the harness run has proprietary tool outputs injected, so it sees information the bare model never receives. The two runs are therefore not the same task, which makes this a product demonstration rather than a controlled measurement of scaffolding.
That is why this post leads with the pre-registered 28-point result instead. A figure that flatters the vendor quoting it gets the primary source first, and then gets read carefully once found.
What should a buyer measure instead of benchmark scores?
Your own workloads, continuously, with the results stored where you can audit them.
Benchmarks are useful for the thing they measure, which is relative capability under somebody else's harness on somebody else's tasks. GAIA Levels 1 and 2 are not your prior-authorization queue or your student records reconciliation.
The replacement is an evaluation harness pointed at your real traffic: Evaluations with LLM-as-judge scoring plus human annotation and export, run against the cases your staff escalate.
Scoring judgment rather than generation is its own discipline, covered in generation is commoditized, judgment is the new frontier.
Three questions make a vendor's accuracy claim checkable. Which scaffold produced this number. Can we re-run it on our data. Do we keep the harness if we leave.
The last one decides the other two. An accuracy result you cannot reproduce in your own environment is a marketing asset, not an engineering input.
What does owning the scaffold mean in practice?
It means the 28 points live on your side of the contract.
The model layer converges and reprices on somebody else's schedule, which is the argument in the model is the commodity, the context layer is the moat.
The scaffold does not converge, because it is made of your taxonomy, your policies, your escalation rules and your evaluation set.
Concretely, on ibl.ai that scaffold is: connectors exposing your systems in place over MCP, scoped to each caller's role; an ontology of typed relationships so the model knows your SIS "student" and your CRM "contact" describe overlapping realities; skills written in plain Markdown by the people who do the work; guardrails and a sandboxed runtime; and evaluations with a full audit trail.
You own all the code and the data, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing, so you can deploy anywhere, including on-premise and fully air-gapped.
That last property is what makes the scaffold an asset rather than a dependency. When the next model lands at a lower price, you re-point the harness and keep every layer you paid to build.
ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
The Edge
Every AI vendor comparison you have been handed is a comparison of scaffolds presented as a comparison of models, and the only pre-registered controlled measurement of that confound puts its size at up to 28 percentage points inside a single model.
That reframes the buying decision. The question is not which model scores highest, because the harness around it moves the score further than the swap does. The question is who ends up owning the harness, since that is where the accuracy was actually manufactured.
A model is a line item you can re-negotiate next quarter. A scaffold built from your own data, policies and evaluation sets is the only part of an AI deployment that compounds, and the only part a vendor can quietly keep.