← All speakers

Sean Cai is an AI data-quality researcher, investor, and author of State of Data, an investigation into how models acquire the expertise needed for economically valuable work. He studies the gap between polished benchmark results and reliable performance inside complicated businesses.

Cai studied at Cornell University, became a General Catalyst fellow, and worked with Hummingbird Ventures and Weekend Fund after building live-videoconferencing technology and vertical AI products. His analysis of browser agents identified a recurring obstacle: systems that can navigate software interfaces still struggle with extended tasks, specialized knowledge, and enterprise reliability. He subsequently joined Costanoa Ventures, helping founders with integrations, internal tools, product metrics, and customer acquisition. His work has also included data-quality research and a reinforcement-learning residency at Prime Intellect.

Through State of Data and his analysis of human-data markets, Cai traces how artificial-intelligence training is becoming an industry organized around human expertise, evaluation, and access to authentic professional workflows.

What distinguishes his thinking

  • Type 1 and Type 2 data: Type 1 captures real professional activity, such as software commits and workflow recordings; Type 2 consists of examples manufactured under artificial conditions. Cai argues that advanced capabilities depend on the context and constraints embedded in actual work.
  • Enterprise workflow data: Finished spreadsheets and database records conceal the decisions, corrections, tool use, and changing context that produced them. Capturing those trajectories continuously gives companies training material that remains relevant as models improve.
  • Task verifiability: Cai evaluates whether work decomposes into checkable steps, whether experts agree on correctness, and whether authentic verified examples are plentiful. Software benefits from tests and public development histories; finance, healthcare, and law involve scarcer evidence and more contested judgments.
  • Benchmark gaming: Vendors can manufacture tests and then sell data designed to improve those same scores. Cai favors deterministic checks, model-assisted grading, and detailed rubrics for demanding tasks such as recurring-revenue reconciliation, leveraged-buyout valuation, and investment comparisons.
  • Antikythera mechanisms: Cai uses this name for systems that translate messy business processes into practical evaluations and reinforcement-learning environments. His writing on enterprise reinforcement learning argues that durable value comes from maintaining these pipelines as organizations, workflows, and underlying models change.

At AI Engineer World’s Fair 2026, Cai described working with companies to turn operational expertise into useful training infrastructure and enterprise reinforcement-learning systems.

Read the topics behind these talks

1 conference talk

▶ Watch ↗

AI Engineer World's Fair 202618:22

State of Data

Read the full talk →

Key ideas

Scroll to read ↓

Turning generalist models into professional experts requires real workflows, credible evaluations, and infrastructure that can keep learning as the underlying models change.

  • From annotation rooms to expertise
    0:38 ↗
  • Compute cannot compensate for every missing input
    3:14 ↗
  • Capture the work, not just the finished file
    4:15 ↗
  • Why some domains mature sooner
    5:52 ↗
  • When the benchmark supplier also sells the training data
    8:04 ↗
  • A score includes the conditions that produced it
    9:12 ↗
  • Similar finance scores, opposite weaknesses
    10:14 ↗
  • Reading demand before the application arrives
    11:58 ↗
  • Useful workflow evidence—and a robotics counterexample
    13:05 ↗
  • Moving the boundary toward dependent work
    14:18 ↗
  • Owning applications without owning the foundation model
    14:44 ↗
  • The infrastructure for repeated adaptation
    16:29 ↗
  • The business becomes the learning pipeline
    17:04 ↗

References