ibl.ai Agentic AI Blog

Insights on building and deploying agentic AI systems. Our blog covers AI agent architectures, LLM infrastructure, MCP servers, enterprise deployment strategies, and real-world implementation guides. Whether you are a developer building AI agents, a CTO evaluating agentic platforms, or a technical leader driving AI adoption, you will find practical guidance here.

Topics We Cover

Featured Research and Reports

We analyze key research from leading institutions and labs including Google DeepMind, Anthropic, OpenAI, Meta AI, McKinsey, and the World Economic Forum. Our content includes detailed analysis of reports on AI agents, foundation models, and enterprise AI strategy.

For Technical Leaders

CTOs, engineering leads, and AI architects turn to our blog for guidance on agent orchestration, model evaluation, infrastructure planning, and building production-ready AI systems. We provide frameworks for responsible AI deployment that balance capability with safety and reliability.

Back to Blog

Nubank Screened 16,000 Simulated Chats Before Going Live

ibl.ai EngineeringSeptember 29, 2026
Premium

Nubank screened open-weight model configurations across more than 16,000 simulated conversations before putting one in front of customers, and published the results. The paper's buried finding is not the simulation count β€” it is that a bank serving 140 million customers in a regulated market chose an open-weight model, and simulation is what made that choice defensible.

The Short Answer

Nubank screened open-weight model configurations across more than 16,000 simulated conversations before deploying one to customers, and the selected model raised self-service rate 8.82 percentage points. On ibl.ai you own all the code and the data and run the platform model-agnostic, so the same screening-then-switching loop is a configuration change inside your own perimeter rather than a renegotiation with a vendor who controls which models you may use.

Nubank's engineering team and Guardrails AI, whose Snowglobe simulator they used, published the method as Screen Before You Serve on 24 September 2026.

It is worth reading in full, because it is one of very few public documents that shows the evaluation record behind a production agent rather than its launch announcement.

Most coverage of the paper stopped at the headline number. The more useful finding sits one layer down, and this page states it plainly: simulation is what let a regulated bank justify an open-weight model.

How did Nubank test its AI agent before putting it in front of customers?

With synthetic customers instead of real ones. Nubank used the Snowglobe simulator on its Card Delivery agent and its successor Card Management β€” described in the paper as Nubank's highest-volume chat-support agent in Brazil.

The mechanism matters more than the count. Synthetic customers react to what the agent says, and simulated tool outputs stand in for production backends, so a multi-step agentic workflow can run end to end without touching a real system of record.

That is what makes the volume affordable. An agent that must call live services to be tested can only be tested as often as those services tolerate. An agent whose tools are mocked can be tested thousands of times overnight.

The paper is explicit about why this was necessary rather than merely convenient. Manual end-to-end testing gives limited coverage, and live experiments expose customers to failures that, in the authors' words, can erode trust.

Do simulated conversations actually predict production behavior?

In Nubank's case, closely enough to act on. Across 4 deployed versions, the paper reports that simulated and production version-level binary evaluator scores show high correlation.

That correlation is the entire load-bearing claim, and it is the one to interrogate before copying the method. A simulation that does not track production is not a cheap experiment β€” it is a confident wrong answer produced at volume.

Four versions is a small sample, and the paper presents it as evidence from one company's agents in one language and one domain, not as a general law. Treat it as a reason to measure your own correlation, not a reason to assume it.

Why did Nubank end up choosing an open-weight model?

Because simulation made the comparison affordable enough to run properly. The paper says the team screened open-weight configurations in over 16,000 simulated conversations β€” the number most write-ups quote, usually without the word "open-weight" that gives it its significance.

The paper names both sides, which is what makes the finding checkable: the open-weight candidate that won was Qwen3.5-122B-A10B with reasoning enabled, and the incumbent it replaced was GPT-5.2.

A bank serving a customer base the paper puts at 140 million does not select a model from a public leaderboard. Leaderboard scores are measured on someone else's tasks, in someone else's language, against someone else's definition of success.

What a risk committee can act on is a screening record in the bank's own domain: its policies, its tools, its customers' phrasing. Sixteen thousand simulated conversations is that record.

This is the part that generalizes beyond fintech. The barrier to using an open-weight model in a regulated setting is rarely capability β€” it is the absence of evidence in your own context. Simulation manufactures that evidence before a customer is ever involved.

What did the two live A/B tests actually measure?

Two different things, and they are routinely merged into one claim. The paper reports them separately, and so should anyone citing it:

Experiment What changed Result
Simulation-guided iteration Agent design, refined against simulated conversations +36.69 tNPS
Open-weight model selection The model, chosen by screening 16,000+ conversations +8.82pp self-service rate
no significant tNPS change

The second row carries a qualifier the summaries drop. The selected model raised self-service rate to the highest level observed at Nubank with no statistically significant change in tNPS β€” customers resolved more on their own, and did not report liking it more or less.

That is a good outcome and an honest one. It is not the same as "simulation doubled customer satisfaction," a sentence that merges the two experiments and overstates both.

What can simulation not screen for?

Anything your synthetic customer does not think to do. A simulated customer is generated from a persona, so it explores the space of behaviors someone anticipated β€” which is a large space, and a bounded one.

It also cannot screen the failure mode where the agent is correct and the backend is not.

Mocked tools return what the mock was told to return; the paper is clear that avoiding production backends is the point, and that necessarily excludes production backend behavior from the test.

And it cannot tell you whether your evaluator is measuring the right thing. A binary evaluator that correlates between simulation and production tells you the two environments agree, not that the metric captures a good conversation.

None of this argues against the method. It argues for the paper's own framing: simulation is a screening step that makes live experiments rarer and better-aimed, not a replacement for them.

How do you run simulation-first deployment on a stack you own?

The loop Nubank describes has a hard prerequisite that the paper does not need to state and a buyer does: swapping the model has to be cheap. Screening open-weight configurations is only worth doing if the winner can actually be deployed.

That is where most agent platforms foreclose the experiment. A per-seat subscription to a vendor's own models does not have a configuration setting for "run the open-weight model that won our screening" β€” the model is the product.

ibl.ai is built the other way round.

You own all the code and the data, the platform is model-agnostic across any LLM, and there is no per-seat pricing β€” so screening candidates and switching to the winner happens inside your perimeter, on infrastructure you control, without a commercial conversation.

The agent sandbox is part of the same story: agents run in a Linux virtual machine with no network at all by default, opened only to hosts you allowlist, which is what makes running unvetted configurations against internal tools a contained experiment rather than a risk.

We wrote about that for districts in Letting a K-12 AI Agent Run Code Without Letting Data Out.

Simulation is also one of four layers we see in every organization that gets agents into production, which is the subject of Forward-Deployed Engineering: The Four-Layer Agent Stack.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

Want to screen agent configurations in your own environment?

We can stand up a simulation-and-evaluation loop against your own policies and tools, on infrastructure you own. Book a 30-minute demo or talk to the ibl.ai team β€” ibl.ai is family-owned and operated from New York, NY.

Why does owning the AI stack matter?

ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing β€” so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.

  • You own all the code and the data

    Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform β€” the stack itself is yours.

  • Model-agnostic

    Run any LLM β€” Claude, GPT, Gemini, Llama, Command, or your own fine-tune β€” and switch providers without rewriting the platform.

  • No per-seat pricing

    Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.

  • Deploy anywhere

    Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY β€” a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

See the ibl.ai AI Operating System in Action

Discover how leading universities and organizations are transforming education with the ibl.ai AI Operating System. Explore real-world implementations from Harvard, MIT, Stanford, and users from 400+ institutions worldwide.

View Case Studies
Work with our team

Pilots, deployment, and full ownership

Most enterprise engagements are one-time, not subscriptions. You integrate ibl.ai with your own data, deploy it on your own infrastructure, and the engineering hours scale with the work β€” so the price tracks the scope, not your headcount.

Start here

Pilot

from $15K

fixed scope Β· fixed timeline

A time-boxed proof of value on your real data β€” not a slide deck.

Best for: Teams that want to see ibl.ai working before committing.

  • Deployed on your infrastructure or our cloud
  • 1–2 production agents wired to a slice of your data
  • One integration (LMS / SIS / SSO / data source)
  • Weekly working sessions with our engineers
  • Pilot fee credits toward a full engagement
Scope a pilot
Most common

Integration & Deployment

$25K – $80K

one-time Β· not a subscription

Full deployment integrated with your data and systems. Engineering hours scale with scope.

Best for: Organizations rolling ibl.ai out across a department, campus, or business unit.

  • Platform deployed in your VPC, on-prem, or air-gapped
  • Integrated with your data + identity (SSO / SAML)
  • Multiple custom agents built to your workflows
  • Engineering hours proportional to scope
  • You own the data Β· run any LLM you choose
Plan a deployment
Full ownership

Codebase Transfer + Custom AI Engineering

Custom quote

perpetual license Β· you own the stack

We transfer the full source code. You own and self-host the entire platform β€” outright.

Best for: Organizations and enterprises that benefit from perpetual ownership and sovereignty.

  • Complete source-code transfer + perpetual license
  • Dedicated AI engineering team on your roadmap
  • Custom agents, models, and integrations to spec
  • Air-gapped capable Β· zero vendor lock-in
  • Family-owned, New York–based long-term partner
Talk about ownership
You own the code and data Run any LLM β€” Claude, GPT, Gemini, Llama Family-owned & operated from New York, NY