← All speakers

Bio, Work & Ideas

Preetika Bhateja

Conference affiliation: Product Manager · Google · 2026

Preetika Bhateja is a product manager and former Google Cloud strategic cloud engineer specializing in production-grade agent evaluation. Her work on generative advertising confronts a consequential reliability problem: ensuring AI-generated creative is accurate, brand-safe, and preserves legally required disclosures.

In 2021, she coauthored Google Cloud guidance on serverless data-pipeline orchestration, integrating Workflows, Dataflow, Cloud Storage, BigQuery, and Cloud Functions into recoverable, observable processing pipelines. In 2022, she coauthored a guide to automating BigQuery dataset snapshots for backup and recovery.

At AI Engineer World’s Fair 2026, Bhateja was identified as a Google product manager affiliated with YouTube Ads. Her joint conference session with Daniel Bump examined how evaluation systems mature alongside advertising agents.

Her approach to reliable advertising agents

  • Human-calibrated evaluation rubrics: Establish agreement about successful outputs, document edge cases, and require explanations that distinguish product ambiguity from genuine agent failures.
  • LLM-judge calibration: Compare automated judgments against expert-reviewed examples and track disagreement instead of trusting unexplained pass-or-fail scores.
  • Agent trace analysis: Inspect intermediate decisions for failures invisible to aggregate metrics. Bhateja described an agent that detected a mandatory advertising disclaimer and then removed it despite explicit instructions to preserve it.
  • Multidimensional creative evaluation: Score accuracy, brand safety, and other requirements separately; refresh test cases with production data and define which regressions must prevent launch.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Reliable agents need more than better prompts: they need focused tools, shared rating criteria, trace inspection, and launch decisions grounded in recurring behavior.

  • Reliability starts with the agent’s foundation
    0:27 ↗
  • Define success, then learn from the outputs
    2:52 ↗
  • Start small and agree on the rubric
    5:52 ↗
  • Collect explanations for each quality dimension
    8:18 ↗
  • Calibrate automated judges against people
    9:38 ↗
  • The agent detected the disclaimer—and removed it
    11:02 ↗
  • Keep a test set outside the iteration loop
    12:21 ↗
  • Improve the evaluator and the agent together
    13:01 ↗
  • Optimize patterns, not isolated runs
    14:45 ↗
  • Let evaluation mature with the product
    15:51 ↗
  • Define launch criteria before deciding to launch
    17:15 ↗
  • How much judging should be automated?
    18:25 ↗

References