The Short Answer
Nubank screened open-weight model configurations across more than 16,000 simulated conversations before deploying one to customers, and the selected model raised self-service rate 8.82 percentage points. On ibl.ai you own all the code and the data and run the platform model-agnostic, so the same screening-then-switching loop is a configuration change inside your own perimeter rather than a renegotiation with a vendor who controls which models you may use.
Nubank's engineering team and Guardrails AI, whose Snowglobe simulator they used, published the method as Screen Before You Serve on 24 September 2026.
It is worth reading in full, because it is one of very few public documents that shows the evaluation record behind a production agent rather than its launch announcement.
Most coverage of the paper stopped at the headline number. The more useful finding sits one layer down, and this page states it plainly: simulation is what let a regulated bank justify an open-weight model.
How did Nubank test its AI agent before putting it in front of customers?
With synthetic customers instead of real ones. Nubank used the Snowglobe simulator on its Card Delivery agent and its successor Card Management β described in the paper as Nubank's highest-volume chat-support agent in Brazil.
The mechanism matters more than the count. Synthetic customers react to what the agent says, and simulated tool outputs stand in for production backends, so a multi-step agentic workflow can run end to end without touching a real system of record.
That is what makes the volume affordable. An agent that must call live services to be tested can only be tested as often as those services tolerate. An agent whose tools are mocked can be tested thousands of times overnight.
The paper is explicit about why this was necessary rather than merely convenient. Manual end-to-end testing gives limited coverage, and live experiments expose customers to failures that, in the authors' words, can erode trust.
Do simulated conversations actually predict production behavior?
In Nubank's case, closely enough to act on. Across 4 deployed versions, the paper reports that simulated and production version-level binary evaluator scores show high correlation.
That correlation is the entire load-bearing claim, and it is the one to interrogate before copying the method. A simulation that does not track production is not a cheap experiment β it is a confident wrong answer produced at volume.
Four versions is a small sample, and the paper presents it as evidence from one company's agents in one language and one domain, not as a general law. Treat it as a reason to measure your own correlation, not a reason to assume it.
Why did Nubank end up choosing an open-weight model?
Because simulation made the comparison affordable enough to run properly. The paper says the team screened open-weight configurations in over 16,000 simulated conversations β the number most write-ups quote, usually without the word "open-weight" that gives it its significance.
The paper names both sides, which is what makes the finding checkable: the open-weight candidate that won was Qwen3.5-122B-A10B with reasoning enabled, and the incumbent it replaced was GPT-5.2.
A bank serving a customer base the paper puts at 140 million does not select a model from a public leaderboard. Leaderboard scores are measured on someone else's tasks, in someone else's language, against someone else's definition of success.
What a risk committee can act on is a screening record in the bank's own domain: its policies, its tools, its customers' phrasing. Sixteen thousand simulated conversations is that record.
This is the part that generalizes beyond fintech. The barrier to using an open-weight model in a regulated setting is rarely capability β it is the absence of evidence in your own context. Simulation manufactures that evidence before a customer is ever involved.
What did the two live A/B tests actually measure?
Two different things, and they are routinely merged into one claim. The paper reports them separately, and so should anyone citing it:
| Experiment | What changed | Result |
|---|---|---|
| Simulation-guided iteration | Agent design, refined against simulated conversations | +36.69 tNPS |
| Open-weight model selection | The model, chosen by screening 16,000+ conversations | +8.82pp self-service rate no significant tNPS change |
The second row carries a qualifier the summaries drop. The selected model raised self-service rate to the highest level observed at Nubank with no statistically significant change in tNPS β customers resolved more on their own, and did not report liking it more or less.
That is a good outcome and an honest one. It is not the same as "simulation doubled customer satisfaction," a sentence that merges the two experiments and overstates both.
What can simulation not screen for?
Anything your synthetic customer does not think to do. A simulated customer is generated from a persona, so it explores the space of behaviors someone anticipated β which is a large space, and a bounded one.
It also cannot screen the failure mode where the agent is correct and the backend is not.
Mocked tools return what the mock was told to return; the paper is clear that avoiding production backends is the point, and that necessarily excludes production backend behavior from the test.
And it cannot tell you whether your evaluator is measuring the right thing. A binary evaluator that correlates between simulation and production tells you the two environments agree, not that the metric captures a good conversation.
None of this argues against the method. It argues for the paper's own framing: simulation is a screening step that makes live experiments rarer and better-aimed, not a replacement for them.
How do you run simulation-first deployment on a stack you own?
The loop Nubank describes has a hard prerequisite that the paper does not need to state and a buyer does: swapping the model has to be cheap. Screening open-weight configurations is only worth doing if the winner can actually be deployed.
That is where most agent platforms foreclose the experiment. A per-seat subscription to a vendor's own models does not have a configuration setting for "run the open-weight model that won our screening" β the model is the product.
ibl.ai is built the other way round.
You own all the code and the data, the platform is model-agnostic across any LLM, and there is no per-seat pricing β so screening candidates and switching to the winner happens inside your perimeter, on infrastructure you control, without a commercial conversation.
The agent sandbox is part of the same story: agents run in a Linux virtual machine with no network at all by default, opened only to hosts you allowlist, which is what makes running unvetted configurations against internal tools a contained experiment rather than a risk.
We wrote about that for districts in Letting a K-12 AI Agent Run Code Without Letting Data Out.
Simulation is also one of four layers we see in every organization that gets agents into production, which is the subject of Forward-Deployed Engineering: The Four-Layer Agent Stack.
1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
Want to screen agent configurations in your own environment?
We can stand up a simulation-and-evaluation loop against your own policies and tools, on infrastructure you own. Book a 30-minute demo or talk to the ibl.ai team β ibl.ai is family-owned and operated from New York, NY.