Nubank’s AI Agents Got 36 Points Better Without Touching Real Customers

Nubank’s AI Agents Got 36 Points Better Without Touching Real Customers

A fintech serving 140 million people couldn’t afford to A/B test its way to better AI — one bad agent version and you’ve torched customer trust at continental scale. So they built a flight simulator for chatbots. The numbers are real, the lift is large, and every CX team building on LLMs should read this before their next deployment.

What happened

Figure 4 : P2 – In all versions, both simulated groups are closer to production than the off-topic control; Snowglobe is no farther from production than the baseline in three versions. Cosine distance
Figure 4 : P2 – In all versions, both simulated groups are closer to production than the off-topic control; Snowglobe is no farther from production than the baseline in three versions. Cosine distance

Nubank’s engineering team built a simulation workflow called Snowglobe that lets them test agentic workflows end-to-end before a single real customer sees them. Synthetic customers respond to agent messages, and simulated tool outputs replace live production backend calls, enabling full multi-turn conversations in a sandbox. They applied this to Card Management — their highest-volume chat support agent in Brazil — across 4 successive deployed versions, and found that simulated binary evaluator scores correlated highly with real production scores, validating that the simulator is predictive rather than decorative. Using simulation-guided iteration, they ran a live A/B test that produced a +36.69 point increase in transactional net promoter score (tNPS) — a number so large it would be implausible from a single prompt tweak. In a separate experiment, they screened over 16,000 simulated conversations across open-weight model configurations, and the winning model delivered a +8.82 percentage point increase in self-service rate — Nubank’s highest SSR ever observed — with no statistically significant tNPS regression. The LLM-as-judge pattern implicitly underpins the binary evaluators used to score those simulated runs.

Cold read

The headline numbers are impressive but structurally compound: the +36.69 tNPS gain bundled simulation plus multiple rounds of iteration across 4 versions, so you cannot isolate how much of the lift came from the simulation framework itself versus the engineering hours it freed up. “High correlation” between simulated and production scores is reported without sharing the actual correlation coefficient or confidence intervals — that’s a qualitative claim dressed as a quantitative one. The simulator’s fidelity depends entirely on how realistic the synthetic customers are; in a regulated fintech context where real users have specific, sometimes irrational, edge-case behaviors, a synthetic customer distribution is always a convenience sample. The 16,000-conversation model screen is more credible methodology, but the result (best open-weight model wins a live test) tells us nothing about whether your domain’s simulation would be equally predictive — Nubank has years of production conversation data to calibrate synthetics against, a cold-start operator does not. Finally, this is a single company, single language (Brazilian Portuguese), single product category: generalization to, say, an English-language SaaS or healthcare context is unproven.

What it means for you

  • Signal maturity: 4/5 — Production results at real scale with real metrics, but single-company evidence
  • Who gets hurt: Teams at regulated-industry startups (insurance, lending, healthcare) who are still doing manual QA or small-scale live tests to iterate on their CX agents — they’re leaving both velocity and NPS points on the table
  • What breaks if this is true: The standard “move slow because we can’t risk live failures” justification for slow AI agent iteration collapses; safety and speed stop being a tradeoff if simulation fidelity holds
  • Why it might not land: Building a faithful simulator requires enough historical conversation data to construct believable synthetic users — startups under ~1M annual support conversations will struggle to build the calibration layer that makes Snowglobe’s predictions reliable
  • Watch for: Third-party tooling (Guardrails AI, LangSmith, ContextQA) announcing simulation-as-a-service features that package this pattern for teams without Nubank’s data moat — that’s the signal this approach is crossing the chasm

Forecast as of 2026-09-25

By Q3 2027, at least two well-funded CX AI platforms (e.g., Sierra, Intercom, or a YC-backed challenger) will ship a named “pre-deployment simulation” product citing this methodology — and at least one will publish a case study claiming >20-point NPS lift, whether or not their simulator fidelity actually warrants it.


Source: Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale — Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Conceição Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath. https://arxiv.org/abs/2609.30137v1

Similar Posts