← All speakers

Aman Gupta is a principal machine learning engineer at Nubank who builds production AI agents and foundation models for financial services. His work combines simulation-driven agent evaluation, model alignment, and large-scale optimization to make customer-support automation more reliable and measurable.

Gupta studied computer science at BITS Pilani, worked on infrastructure automation and security at Amazon, and earned a research-focused computer science master’s degree at Carnegie Mellon University. He subsequently worked on computer vision and autonomous systems at Apple before spending six years at LinkedIn, where he became a senior staff machine learning engineer and senior manager in Core AI. There, he led work on ranking-model compression and optimization for products including Feed, Ads, and Jobs, contributing to the GDMix personalization framework and DuaLip, an open-source solver for large recommendation and allocation problems.

He joined Nubank in April 2025, extending that background into financial foundation models and customer-facing agents. His research includes nuFormer, which learns from financial transaction histories, and AlphaPO, an ICML 2025 preference-optimization method addressing likelihood displacement and over-optimization.

  • Evaluation-driven customer-support agents. Gupta’s first-authored research covers deployments spanning card delivery, debt management, credit-limit support, card management, and product explanations. One card-delivery deployment improved transactional customer satisfaction by 37 percentage points and self-service rates by 29 percentage points over earlier agents.
  • Simulation-driven agent evaluation. Working with Snowglobe’s Shreya Rajpal, Gupta helped apply synthetic customer personas, realistic mocked tools, and multi-turn conversations to test Nubank agents before release. Their AI Engineer session describes a reported 20-fold acceleration in agent iteration, earlier regression detection, and faster comparisons between open-source and frontier models.
  • Human-calibrated improvement loops. Gupta combines simulated interactions and production traces with expert-reviewed graders, automated prompt optimization, and controlled deployment. His production-agent evaluation playbook emphasizes inter-rater reliability and offline-to-online correlation: evaluation scores matter only when they track real customer outcomes.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Nubank uses simulated customers and consistent mocked tool state to test agent changes before live experiments, shortening the wait for useful evaluation signal.

  • Improving support without waiting on every live experiment
    0:32 ↗
  • An agent example is a trajectory
    2:21 ↗
  • Two ways to obtain data, two different costs
    4:49 ↗
  • Remove the wait for feedback
    5:46 ↗
  • Run the real agent inside a simulated environment
    7:20 ↗
  • Maria Souza: a persona with consistent account state
    8:42 ↗
  • Evaluate real and simulated conversations together
    10:10 ↗
  • Check whether simulation tracks production
    11:14 ↗
  • Screen candidates while preserving self-service
    11:55 ↗
  • Shortlist models before testing them live
    13:33 ↗
  • Continuous improvement depends on trustworthy signal
    14:26 ↗

References