← All organizations

Local commerce, delivery and merchant software

DoorDash

DoorDash operates a local commerce marketplace connecting consumers with restaurants, grocery stores and retailers. Merchants use its services for delivery, pickup and their own online stores. Ask DoorDash adds conversational discovery: customers can describe a meal or turn recipe links and cookbook photos into grocery carts. At its 2026 introduction, restaurant search and grocery shopping were available in select areas on iOS.

Founded in 2013, DoorDash’s founders include CEO Tony Xu, Andy Fang and Stanley Tang. Its consumer-memory platform turns shopping behavior into semantic descriptions of preferences, supplying language context to LLMs and embeddings and graph features to conventional recommendation models. Its personalized carousel architecture uses that consumer context to generate titles and search intents offline. When a customer opens a store, vector search and structured category retrieval find items for those carousels. Keeping LLM generation outside the live request avoids its per-page latency and spreads generation costs across the refresh interval.

DoorDash completed its acquisition of SevenRooms in 2025, extending its merchant offering into in-person hospitality through reservations, customer relationship management, guest experience and marketing software. Integration into the Commerce Platform was planned at closing. DoorDash reported operating in over 30 countries in that announcement; SevenRooms served more than 13,000 venues globally, spanning dining, hotels, nightlife and entertainment.

Explore the recordings

The supplied DoorDash archive contains one recording. It offers two useful paths: organizing cross-functional evaluation work and building the platform that supports it. The guide below describes the presenters’ recorded account; it does not establish DoorDash’s current practices or product capabilities.

Start with the evaluation process and its owners

Watch AI Evals for Cross-Functional Teams for a practical account of how evaluation responsibilities can span strategy and operations, product, and engineering. Follow the presenters’ loop from captured traces and sessions through sampling, human annotation, golden datasets, judge calibration, and monitoring. This path is useful for teams deciding who sets the quality bar, translates requirements into rubrics, runs reviews, and maintains evaluation infrastructure. The speakers distinguish session-level judgments from trajectory-based evaluation, showing why applications may need different review workflows.

Then examine the infrastructure for self-service

Revisit the same recording for the presenters’ shift from UI-first tooling toward stable APIs and workflow-first operations. Their account separates telemetry access through APIs, an SDK, and MCP from annotation, dataset review, and judge-management workflows. Operators building specialized annotation interfaces with coding agents provide a useful example of how stable platform primitives can support varied use cases. The judge-calibration discussion covers baseline scores and comparisons before promotion. The presenters report faster iteration and lower per-annotation spending, but supply no numerical spending reduction; treat those as recorded outcomes, not verified current performance.

1 talk

Newest first

2 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Company sources · checked 2026-09-01