ibl.ai Agentic AI Blog

Insights on building and deploying agentic AI systems. Our blog covers AI agent architectures, LLM infrastructure, MCP servers, enterprise deployment strategies, and real-world implementation guides. Whether you are a developer building AI agents, a CTO evaluating agentic platforms, or a technical leader driving AI adoption, you will find practical guidance here.

Topics We Cover

Featured Research and Reports

We analyze key research from leading institutions and labs including Google DeepMind, Anthropic, OpenAI, Meta AI, McKinsey, and the World Economic Forum. Our content includes detailed analysis of reports on AI agents, foundation models, and enterprise AI strategy.

For Technical Leaders

CTOs, engineering leads, and AI architects turn to our blog for guidance on agent orchestration, model evaluation, infrastructure planning, and building production-ready AI systems. We provide frameworks for responsible AI deployment that balance capability with safety and reliability.

Back to Blog

Scaffolding Moves Accuracy 28 Points. Model Choice Is the Smaller Variable.

ibl.ai EngineeringOctober 8, 2026
Premium

A pre-registered controlled comparison published in June 2026 found that scaffold choice alone moves measured agent accuracy by as much as 28 percentage points inside a single model. The published clinical record on OpenEvidence points the same way from two directions: retrieval scaffolding drove fabricated references to zero across 4,979 citations, and a 100-question board-style pilot still capped at 41%. Benchmarks measure the scaffold as much as the model, and the scaffold is the part you own.

The Short Answer

Measured AI accuracy is a property of the scaffold as much as the model. A pre-registered controlled comparison found scaffold choice alone moving accuracy up to 28 percentage points inside one model. So buy the layer you can own, not the benchmark: with ibl.ai you own all the code and the data, run it model-agnostic across any LLM, and pay with no per-seat pricing.

The question "which model is most accurate" has an uncomfortable answer. It depends on what you wrapped around it, and the wrapping moves the number further than the model swap does.

Why does the same AI model score differently on the same benchmark?

Because the published score measures the model and its harness together, and nobody separates them.

Scaffold Effects on GAIA: A Controlled Comparison, submitted to arXiv on 7 June 2026 by Jason Starace, was built to measure exactly that confound. It is pre-registered, which matters: the hypothesis was fixed before the runs.

The design holds tasks and conditions constant and varies only the scaffold. Three scaffolds were tested: ReAct, a planner-actor-rater multi-agent design, and planner-then-executor.

Five models across three providers, on GAIA validation Levels 1 and 2, with three attempts per question.

The finding, in the author's words: "Scaffold choice alone moves measured accuracy by as much as 28 percentage points within a single model."

The paper names the thing this conflation hides as the elicitation gap: the distance between what a model can do and what its scaffold lets it do. Published agent capability scores sit somewhere inside that gap without telling you where.

For a procurement team, that is a direct instruction. A vendor demo that beats another vendor's demo may be a better harness around a similar model, and the harness is the part you can build, buy, or own.

What does "scaffolding" actually mean in an enterprise deployment?

It means the five things around the model that decide whether its answer is usable, none of which arrive with a model subscription.

The layers are not exotic, and they map one to one onto the architecture we build with clients:

Scaffold layer What it decides Who has to build it
Domain knowledge Whether the corpus the model reasons over is yours or the internet's You, from your own systems of record
Retrieval Whether a claim is grounded in a document that exists You, over your data layer
Task instruction Whether the model is doing your job or a generic one Your domain experts, in plain language
Guardrails What the model is not allowed to say or do You, against your own policy
Evaluation Whether any of the above is working this week You, on your own workloads

Four of the five columns on the right say the same word. That is the finding, restated as an org chart.

The GAIA comparison only varied the orchestration pattern and still produced a 28-point spread. Vary retrieval quality and corpus coverage as well, and the model becomes the smallest term in the expression.

Does scaffolding fix clinical accuracy? What the published OpenEvidence record shows

It fixed the failure mode everyone was afraid of, and did not fix the one that decides deployment. Both results are published, and they are about the same product.

On sourcing, the scaffolding worked.

Reference quality of OpenEvidence across five medical specialties, published in npj Health Systems on 2 September 2026, examined all 4,979 references returned to 150 standardized prompts across oncology, cardiology, rheumatology, psychiatry and infectious diseases, 30 prompts each.

No reference was confirmed as fabricated. Three ASCO meeting abstracts carried author attribution errors, 0.06% of the set, and all three were confirmed to exist with the correct titles and years.

For anyone who watched general-purpose chatbots invent citations, that is what a working retrieval layer looks like in a number.

On hard reasoning, the same scaffolding did not carry.

The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios, posted to medRxiv on 4 December 2025 by Jagarapu, Babata, Chamarthi and Hoyt, ran 100 board-style questions drawn from the MedXpertQA dataset.

Maximum accuracy was 41% for Deep Consult and 34% for Quick Consult. Inter-rater concordance was 77% and 72%, with Cohen's kappa at 0.74 and 0.69, so the grading was consistent enough to trust the ceiling.

The authors' conclusion is the operative one: low accuracy combined with inconsistent performance on complex cases argues against deployment without expert oversight.

Read together, the two studies are a specification, not a contradiction. The scaffold you build determines which failure modes you eliminate, and you only learn which ones remain by evaluating on the work you actually do.

Where does the "55% to 96.7%" scaffolding number come from?

From a credit-card chargeback leaderboard, not from a medical benchmark, and not from the clinical platform it is usually credited to.

The claim circulating in early October 2026 is that a 55% accurate base model was taken to 96.7% purely by adding five scaffolding layers, and it is widely attached to a clinical decision-support company. That attribution is wrong.

The figures come from the Findustry AI chargebacks performance leaderboard, which scores accept-or-contest decisions on credit-card chargeback responses. The base model is GPT-5.5: 55.0% bare, 96.7% inside Findustry's vertical harness.

The pair also is not new. It sits in that page's findings section under a chart dated 28 May 2026, and the current leaderboard revision does not rank GPT-5.5 at all. A figure going around as this week's news is four and a half months old.

The misattribution appears to come from a newsletter that placed the Findustry figure and a separate claim about a clinical platform in adjacent sentences. The two were then read as one.

Three caveats belong with the number even now that it has a source.

It is the vendor's own leaderboard, self-published and not independently audited, and it publishes no methodology section.

On scale it says only that building a benchmark means curating "dozens or even hundreds of test scenarios," which is a description of the craft rather than a disclosure of what these runs measured.

Most importantly, the harness run has proprietary tool outputs injected, so it sees information the bare model never receives. The two runs are therefore not the same task, which makes this a product demonstration rather than a controlled measurement of scaffolding.

That is why this post leads with the pre-registered 28-point result instead. A figure that flatters the vendor quoting it gets the primary source first, and then gets read carefully once found.

What should a buyer measure instead of benchmark scores?

Your own workloads, continuously, with the results stored where you can audit them.

Benchmarks are useful for the thing they measure, which is relative capability under somebody else's harness on somebody else's tasks. GAIA Levels 1 and 2 are not your prior-authorization queue or your student records reconciliation.

The replacement is an evaluation harness pointed at your real traffic: Evaluations with LLM-as-judge scoring plus human annotation and export, run against the cases your staff escalate.

Scoring judgment rather than generation is its own discipline, covered in generation is commoditized, judgment is the new frontier.

Three questions make a vendor's accuracy claim checkable. Which scaffold produced this number. Can we re-run it on our data. Do we keep the harness if we leave.

The last one decides the other two. An accuracy result you cannot reproduce in your own environment is a marketing asset, not an engineering input.

What does owning the scaffold mean in practice?

It means the 28 points live on your side of the contract.

The model layer converges and reprices on somebody else's schedule, which is the argument in the model is the commodity, the context layer is the moat.

The scaffold does not converge, because it is made of your taxonomy, your policies, your escalation rules and your evaluation set.

Concretely, on ibl.ai that scaffold is: connectors exposing your systems in place over MCP, scoped to each caller's role; an ontology of typed relationships so the model knows your SIS "student" and your CRM "contact" describe overlapping realities; skills written in plain Markdown by the people who do the work; guardrails and a sandboxed runtime; and evaluations with a full audit trail.

You own all the code and the data, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing, so you can deploy anywhere, including on-premise and fully air-gapped.

That last property is what makes the scaffold an asset rather than a dependency. When the next model lands at a lower price, you re-point the harness and keep every layer you paid to build.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

The Edge

Every AI vendor comparison you have been handed is a comparison of scaffolds presented as a comparison of models, and the only pre-registered controlled measurement of that confound puts its size at up to 28 percentage points inside a single model.

That reframes the buying decision. The question is not which model scores highest, because the harness around it moves the score further than the swap does. The question is who ends up owning the harness, since that is where the accuracy was actually manufactured.

A model is a line item you can re-negotiate next quarter. A scaffold built from your own data, policies and evaluation sets is the only part of an AI deployment that compounds, and the only part a vendor can quietly keep.

Why does owning the AI stack matter?

ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.

  • You own all the code and the data

    Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.

  • Model-agnostic

    Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.

  • No per-seat pricing

    Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.

  • Deploy anywhere

    Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

Related Articles

The AI Harness Thesis: Orchestration Beats Model Selection

Enterprises spend their AI strategy debating which model to buy. The model is the commodity — it is replaced every few months and its price falls. The harness around it (retrieval, validation, routing, memory) is the durable asset, and it only compounds if you own it.

ibl.ai EngineeringJuly 29, 2026

University of Oxford: Who Should Develop Which AI Evaluations?

The memo proposes a framework for assigning AI evaluation development to various actors—government, contractors, third-party organizations, and AI companies—by using four approaches and nine criteria that balance risk, method requirements, and conflicts of interest, while advocating for a market-based ecosystem to support high-quality evaluations.

Jeremy WeaverFebruary 11, 2025

Agents That Only Read Are Demos. Write Access Is the Whole Problem.

Ampersand raised a $15M Series A led by Bessemer Venture Partners on 6 October 2026 to build integration infrastructure that lets agents read from and write into systems of record like Salesforce and NetSuite, arguing that without that access enterprise AI is vaporware. The round is a price signal on the layer between the agent and production data, and it raises a question every buyer should ask before signing: who owns that layer.

ibl.ai EngineeringOctober 8, 2026

Hallucination Is Provably Inevitable. Stop Procuring an Error Rate.

A line of computability-theory papers argues that LLM hallucination is inevitable rather than a defect awaiting a fix, which means a contract specifying zero hallucination is purchasing an impossibility. The same research names the alternative, and it is an architecture rather than a threshold. Procurement should specify grounding, citation, abstention and an audit trail.

ibl.ai EngineeringOctober 2, 2026

See the ibl.ai AI Operating System in Action

Discover how leading universities and organizations are transforming education with the ibl.ai AI Operating System. Explore real-world implementations from Harvard, MIT, Stanford, and users from 400+ institutions worldwide.

View Case Studies
Work with our team

Pilots, deployment, and full ownership

Most enterprise engagements are one-time, not subscriptions. You integrate ibl.ai with your own data, deploy it on your own infrastructure, and the engineering hours scale with the work — so the price tracks the scope, not your headcount.

Start here

Pilot

from $15K

fixed scope · fixed timeline

A time-boxed proof of value on your real data — not a slide deck.

Best for: Teams that want to see ibl.ai working before committing.

  • Deployed on your infrastructure or our cloud
  • 1–2 production agents wired to a slice of your data
  • One integration (LMS / SIS / SSO / data source)
  • Weekly working sessions with our engineers
  • Pilot fee credits toward a full engagement
Scope a pilot
Most common

Integration & Deployment

$25K – $80K

one-time · not a subscription

Full deployment integrated with your data and systems. Engineering hours scale with scope.

Best for: Organizations rolling ibl.ai out across a department, campus, or business unit.

  • Platform deployed in your VPC, on-prem, or air-gapped
  • Integrated with your data + identity (SSO / SAML)
  • Multiple custom agents built to your workflows
  • Engineering hours proportional to scope
  • You own the data · run any LLM you choose
Plan a deployment
Full ownership

Codebase Transfer + Custom AI Engineering

Custom quote

perpetual license · you own the stack

We transfer the full source code. You own and self-host the entire platform — outright.

Best for: Organizations and enterprises that benefit from perpetual ownership and sovereignty.

  • Complete source-code transfer + perpetual license
  • Dedicated AI engineering team on your roadmap
  • Custom agents, models, and integrations to spec
  • Air-gapped capable · zero vendor lock-in
  • Family-owned, New York–based long-term partner
Talk about ownership
You own the code and data Run any LLM — Claude, GPT, Gemini, Llama Family-owned & operated from New York, NY