AI Engineer World's Fair 2026
AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash
Read the talk
AI Evals for Cross-Functional Teams
DoorDash’s GenAI platform moved evaluation beyond an engineering harness by giving domain experts stable APIs, task-specific annotation workflows, golden datasets, and a reviewable way to calibrate LLM judges.
From a talk by Nachiket Paranjape and Swaroop Chitlur Haridas
At a glance
Ideas worth remembering
Evaluation is cross-functional because engineering can provide traces, datasets, and judges, but domain experts must define and apply the product’s quality criteria.
The recurring loop is: trace, sample, annotate, review, create a golden dataset, calibrate, monitor, and repeat.
Stable APIs let the platform team maintain shared capabilities while operators use coding agents to build annotation UIs suited to menus, images, manual tests, and other tasks.
Self-service judge calibration remains a review process: prompt owners can compare the original and calibrated prompts before promoting a new judge.
Prompt ownership can remain flexible while teams learn; DoorDash has seen strategy and operations, product, and engineering each own it in different groups.
DoorDash reports lower per-annotation costs and faster iteration for thousands of rows each week, but the talk does not quantify the savings or speedup.
One platform must support different kinds of quality
DoorDash’s GenAI platform team supports product teams with shared infrastructure. Its initial platform choices—an LLM gateway for switching models, an agent gateway that centralizes tool connections, authentication, and agent identity, plus open-weights model hosting—help teams balance three competing forces: accuracy, latency, and cost. The speakers treat those forces as relevant to agents as well as individual models. Evaluation became the fourth pillar because none of the other primitives can tell a product team whether the resulting behavior is actually good enough.
The evaluation target changes by product. A consumer discovery or shopping assistant may need a judgment over an entire session. Personalization work may need human judgment at much greater scale. A multi-agent system may need trajectory-based evaluation, where the sequence of actions matters rather than only the final answer. A common platform therefore cannot assume that every team will grade the same artifact with the same interface.
The first organizational discovery was that engineers were not the only people qualified to define or assess quality. Strategy and operations staff, product managers, and labeling partners carried the domain knowledge needed for the judgments. DoorDash first emphasized UIs so non-engineers could contribute, then added stable APIs so engineers could build without waiting for the platform team. Coding agents pushed the design toward what the speakers call workflow-first: strategy and operations staff and product managers could increasingly run the work themselves, not merely submit input through a fixed screen.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Quality becomes a division of work
This changes evaluation from an engineering harness into a cross-functional operating process. Traces, datasets, and scoring mechanisms are technical objects, but the criteria encoded in them come from people who understand the product and its users. The platform’s job is to carry that knowledge through the system without pretending engineering can infer every team’s quality bar.
The speakers divide the work into four parallel responsibilities:
- Strategy and operations: set priorities and decide what quality bar the product should meet.
- Product: translate those goals into rubrics and workflows that people can apply.
- Operations: run the annotation process and produce human judgments.
- Engineering: provide APIs, telemetry, datasets, and automated judges.
The handoffs matter. A broad business goal must become a rubric before annotators can apply it consistently, and those annotations must become structured data before engineering can use them to calibrate or monitor an automated judge.
DoorDash then turns these responsibilities into a recurring loop: capture and inspect behavior, sample it down to a set humans will really review, annotate it with domain expertise, review the annotations, promote suitable examples into a golden dataset, calibrate against that reference, monitor behavior over time, and repeat. Sampling keeps the human workload manageable; the golden dataset gives later measurement an explicit human reference instead of an assumed definition of quality.
Set priorities and the target quality bar.
The platform combines business priorities, operational judgment, and engineering machinery rather than assigning the whole evaluation process to one function.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Two platform surfaces support one quality loop
At the platform level, DoorDash separates telemetry from workflow. The telemetry layer holds traces, scores, and observations and exposes them through MCP, an SDK, and APIs. The workflow layer is where strategy, operations, and product teams set annotation tasks, review golden datasets, create judges, and calibrate them. One layer records and exposes what the system did; the other organizes the human and automated work used to judge it.
Stable APIs connect those surfaces. Scores and datasets use platform-owned APIs, and the platform’s own UIs sit on the same foundation. This is the important meaning of API-first here: a team does not need to abandon shared data and evaluation primitives merely because it needs a different interface. SDK users and custom UIs still operate over a common access plane.
The lifecycle begins with evidence from production behavior: capture traces and sessions and calculate available scores. Humans then add judgment, context, and domain knowledge. Those labeled examples support judge calibration, after which monitoring produces the next collection of traces. The talk gives no required sampling method, sample size, annotation-agreement threshold, or rule for admitting an example into the golden dataset. Those remain consequential design choices for each implementation.
Capture sessions, traces, scores, and observations.
Production behavior becomes manageable human review, then reference data for calibration and another round of monitoring.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Operators build the interface the task needs
Annotation is where a large collection of agent sessions becomes an investigation: which cases went well, which failed, and what happened inside each session? The obstacle is interface variety. Image review, manual testing, restaurant-menu grading, and other tasks do not necessarily need the same controls or layout. A central platform team trying to design every task-specific UI would become a queue for the rest of the organization.
DoorDash splits responsibility instead. The platform team maintains the APIs, strategy and operations decide what should be annotated, and annotators perform the labeling. Because the common data operations already exist behind stable APIs, strategy and operations partners can use coding agents to create the task-specific annotation interface themselves. The presenters mention coding tools including Codex and Claude Code, but the architectural point does not depend on a particular agent: generated interfaces call shared platform capabilities rather than recreating the evaluation backend.
The restaurant-menu example is deliberately modest. The interface only needs to present the material and collect the required annotation cleanly. Its value is not technical novelty; it puts workflow design in the hands of operators who understand the task. Image annotation and manual testing can use different UIs while retaining the same underlying platform pattern.
This freedom also creates an unstated governance question. The talk does not explain how generated UI code is reviewed, secured, maintained, or checked for faithful implementation of a rubric. The example establishes that stable APIs make decentralized interface creation possible; it does not establish that every generated annotation tool will be correct merely because it looks clean and completes the task.
Expose shared scores, datasets, and data operations.
The platform team maintains common APIs while operators choose the data and generate an interface suited to the annotation task.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Judge calibration must be self-serve and reviewable
Once human annotations have become a golden dataset, a team can use them to improve an LLM judge. The described process starts with a simple judge prompt stating what should be measured. The team runs that judge over traces to establish baseline scores, applies a prompt-optimization loop, and promotes the resulting prompt when the partner team is satisfied. Calibration here changes the judge’s instructions against reference data; the speakers do not describe training model weights.
DoorDash packages this process behind a self-service UI because judge calibration is still new and evolving, even if the basic idea sounds straightforward. A product manager or operator chooses the exposed configuration and runs the loop without repeated back-and-forth with engineering. The example supports multiple model providers, making model selection part of the workflow rather than a platform-team intervention.
Self-service alone would hide too much. Prompt optimization can otherwise feel like a closed box, so the platform shows the original system prompt beside the calibrated prompt. The prompt owner can inspect what changed before adopting it. The presenters report significant improvement in one example, but provide neither the prompts nor a numerical metric, so the stronger supported lesson is about reviewability: optimization should produce a proposed artifact that a domain owner can examine, not merely a higher unexplained score.
The talk also leaves important calibration details open. It does not specify the optimization objective, training and validation split, acceptance threshold, repeated-run variance, or safeguards against overfitting the golden dataset. A side-by-side prompt view can help people understand the proposed change, but visual inspection alone does not prove that a judge will generalize to new production cases.
Who performs that review varies. Strategy and operations own the judge prompt in some teams; a product manager or engineering owns it elsewhere. DoorDash treats this variation as evidence that teams and organizational design are still learning. The platform therefore supports different owners instead of prematurely standardizing one function as the permanent home of judge prompts.
Human-reviewed reference examples used for calibration.
Golden data drives prompt optimization, but the prompt owner still inspects the proposed change before promotion.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Lower costs come from shorter dependency loops
The platform itself evolved through this work. DoorDash moved from an initial UI emphasis toward API-first and workflow-first operation while reusing existing internal infrastructure. The self-service goal is organizational as much as technical: teams should not need the central platform group to operate every annotation or calibration cycle.
The speakers report a substantial reduction in per-annotation spending at a workload of thousands of rows each week. They also connect self-service annotation and judge calibration to faster iteration. The talk supplies no before-and-after cost, percentage reduction, or measured cycle time, so these remain qualitative operating results rather than quantified performance claims.
The ending returns to the loop rather than the interface: inspect traces and sessions, sample to a size humans can handle, annotate with domain knowledge, build the golden dataset, and use it to calibrate workflows, agents, and LLM judges. Then repeat. The durable mechanism is not any single UI or judge prompt. It is a system that repeatedly turns observed behavior into human judgment, reviewable automation, and another production measurement cycle.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The official recording page includes the video, timestamped transcript, chapter navigation, and a reading version of the talk.
- Nachiket Paranjape on XReference
Speaker profile supplied with the recording, including posts about DoorDash’s evaluation platform and cross-functional AI quality work.
Related talks
- The maturity phases of running evals
Extends the same progression from expert human judgment to validated automated judges and production-informed evaluation systems.
- Shipping AI That Works: An Evaluation Framework for PMs
Develops the product-manager side of shared prompt and evaluation ownership through repeatable experiments and human-labeled data.
- Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
Provides a complementary production example involving representative simulations, LLM judges, launch gates, and continual evaluation of multi-turn agents.
Read the complete timestamped transcript
- 0:01
[music]
- 0:13
Good afternoon everyone. Thanks for uh
- 0:16
coming for a post lunch uh talk. Always
- 0:19
appreciate that. Um my name is Farup and
- 0:22
here's my teammate Nachiket. Uh we are
- 0:25
uh here behalf of the Door Dash Genai
- 0:28
platform team. Um and we kind of wanted
- 0:31
to share our eval journey. Uh it started
- 0:35
as uh uh you know eval is another
- 0:38
engineering thing but then it slowly we
- 0:40
realized it evolved into a cross
- 0:42
functional effort and we kind of want to
- 0:45
share our story here. So what is this
- 0:48
team? This team is a gen platform team.
- 0:50
Uh we are a horizontal team that helps
- 0:52
all other product teams. So product
- 0:54
teams at Door Dash build on top of the
- 0:57
infrastructure and the primitives that
- 0:58
we provide. Um and we see our uh USP and
- 1:03
the value that we provide is that we
- 1:05
help product teams balance these three
- 1:07
forces which is accuracy, latency and
- 1:10
cost. Um initially we applied this in
- 1:13
terms of models but if you think about
- 1:14
it it also applies to agents. Um and the
- 1:18
way we achieve this is we have uh
- 1:20
primitives and building blocks. Um so
- 1:23
for example we have an LLM gateway where
- 1:25
you can easily switch between different
- 1:27
models uh and try the latest and
- 1:29
greatest. Uh we have an agent gateway
- 1:32
where you can connect to tools uh and
- 1:35
other agents uh and we help solve
- 1:38
authentication uh agent identity and
- 1:40
other things in a central place uh which
- 1:43
our security team can bless. Um
- 1:45
similarly we pair the LLM gateway with
- 1:48
open weights models hosting. Uh, of
- 1:50
course cost is a number one concern
- 1:52
these days. Uh, and we uh kind of
- 1:55
invested in open weights models uh and
- 1:57
have seen significant impact uh already.
- 2:00
Um, and maybe we'll talk about that in a
- 2:03
future conference. Uh, the fourth pillar
- 2:05
is eval and that's the part that we
- 2:07
would want to share today. Um
- 2:11
when we started talking to product teams
- 2:13
internally at Door Dash uh there were
- 2:16
varying uh distinct needs across teams.
- 2:20
We had a consumer discovery and shopping
- 2:22
assistant team. Uh for those who
- 2:24
attended Ragago talk earlier today uh
- 2:26
you will uh see the need for session
- 2:29
level quality judgments. um uh then the
- 2:32
personalization ML then you needed uh a
- 2:35
way to scale up human judgment and with
- 2:38
multi- aent systems we needed trajectory
- 2:40
based evals now the question is how do
- 2:43
you cater to all these different needs
- 2:47
under a common platform
- 2:50
and uh as we spoke to these teams we
- 2:52
realized like um we needed to empower
- 2:56
the people who are the domain experts
- 2:58
and in our case that was strategy and
- 3:00
operations folks it as product managers
- 3:02
uh it was even labeling partners uh and
- 3:05
not only engineers so we kind of started
- 3:07
with like okay we have to be UI first
- 3:09
and this was the guidance we had from
- 3:11
Andy Fang our co-founder as well um so
- 3:14
we had UIs for non-engineers to
- 3:16
contribute uh then we kind of evolved to
- 3:19
also being API first so that engineers
- 3:22
can also build and not be blocked on the
- 3:24
central platform and they can build
- 3:26
their own uh uh uh systems uh and Then
- 3:30
of course with the coding agents now we
- 3:33
have become workflow first where we kind
- 3:35
of empower SNO and PMs to also being
- 3:37
able to uh navigate uh the platform and
- 3:41
uh run operations as well. Um so with
- 3:45
that context I'll hand it off to Nachig
- 3:47
to talk about uh how we went about
- 3:49
delivering this. Cool. Thanks Harup. Um
- 3:52
and thanks everyone for joining us. I
- 3:54
know France is playing right now and I
- 3:56
promise you this will be better than
- 3:58
that. I'm kidding. Um so as Faroo was
- 4:01
saying uh Evals is not just an
- 4:03
engineering harness it is a cross
- 4:05
functional effort across different
- 4:07
pillars across different uh teams uh
- 4:10
that actually helps us add all the
- 4:13
domain specific knowledge into our uh
- 4:16
into the quality of the AI itself. So
- 4:18
from your traces to your data sets uh
- 4:21
from you know scoring mechanisms uh this
- 4:24
is all basically a team sport. we all
- 4:27
have to play uh and help improve the
- 4:30
quality of AI.
- 4:33
So going a little bit deeper into the
- 4:36
same aspect uh we have different uh
- 4:38
teams uh at Door Dash who help us
- 4:40
actually improve the quality of AI. So
- 4:42
you're going to have your strategy and
- 4:43
operations folks who are going to set
- 4:45
priorities, set the quality bar that you
- 4:47
want to aim for. You're going to have
- 4:48
your product people who are going to
- 4:50
translate uh these requirements into
- 4:53
rubrics workflows. You're going to have
- 4:55
your operations teams running uh
- 4:57
annotations. You're going to have your
- 4:58
engineering teams like us uh providing
- 5:01
APIs, telemetry, data sets, judges, all
- 5:04
you know the the cool things. Um and
- 5:08
combining all these together is is what
- 5:11
a recipe is for actually making sure
- 5:14
that you are shipping quality AI
- 5:16
products through an eval platform.
- 5:19
So we've tried to boil this down uh into
- 5:22
sort of you know like a a continuous
- 5:24
iteration loop. Uh so right from tracing
- 5:28
uh you know having a tracing solution
- 5:30
viewing your sessions your traces to
- 5:33
sampling them down you uh you know to a
- 5:35
very small uh set that you actually want
- 5:38
to look at uh annotating these with the
- 5:40
domain specific expertise that you bring
- 5:42
in with the different teams I mentioned
- 5:45
reviewing those uh then creating those
- 5:47
golden data sets which are going to be
- 5:50
you you know your uh golden data sets
- 5:53
and that that you want to measure or
- 5:54
calibrate against uh and then of course
- 5:56
like you know monitoring this over a
- 5:58
period of time and then you know rinse
- 6:00
and repeat uh go through the whole loop
- 6:02
again. So this is in our experience has
- 6:05
been you know like a good sort of
- 6:06
continuous loop uh for you know shipping
- 6:09
quality AI
- 6:12
at the plat on on the platform level uh
- 6:14
we we have two surfaces uh so we have
- 6:17
the telemetry layer uh where we have all
- 6:19
our traces our scores uh observations
- 6:22
that is also sort of the plane where
- 6:25
users are able to access these traces
- 6:27
using an MCP using an SDK uh using our
- 6:30
APIs and then we have the workflow This
- 6:33
is where a lot of our strat ops, our
- 6:34
product teams operate on the platform.
- 6:37
So this is where all the annotation
- 6:38
tasks are set. Uh you know this is where
- 6:41
they review their golden data sets, uh
- 6:43
create their judges, calibrate their
- 6:45
judges and so on.
- 6:47
So maybe today we'll go through you know
- 6:50
these sort of four different uh modules
- 6:53
or pillars of our platform uh step by
- 6:55
step. Uh so again first one uh tracing
- 6:58
and sampling uh which is actually
- 7:00
capturing what your agents what your
- 7:03
LLMs are actually uh you know outputting
- 7:06
for the lack of better words uh and
- 7:08
actually viewing those.
- 7:10
Now in order to also power this uh whole
- 7:14
platform we have I think as far
- 7:16
mentioned we have gone in an API first
- 7:18
uh approach. Uh what that has allowed us
- 7:20
to do is have these table APIs that
- 7:23
actually uh you know and then you know
- 7:25
build UIs uh on top of that. Uh so all
- 7:29
our scores our data sets uh these are
- 7:31
all powered by very stable APIs uh that
- 7:34
our team owns. Uh so all your API access
- 7:38
uh including you know like an SDK access
- 7:40
is basically powered by this single uh
- 7:42
plane.
- 7:44
Um again
- 7:47
going back uh and you know like just
- 7:49
refreshing your memory. Uh step one
- 7:51
capture your traces uh capture your
- 7:54
sessions uh measure your scores. Uh then
- 7:57
you want to start uh almost you know
- 7:59
like adding all your judgment your
- 8:02
context your domain knowledge uh and
- 8:05
then calibrating your judges is what we
- 8:07
have seen as the whole uh life cycle.
- 8:13
Step two is on the annotation side. Uh
- 8:15
so you obviously are capturing a lot of
- 8:17
your uh agentic behavior, your sessions,
- 8:19
your traces, but you actually want to
- 8:21
see what are some places where things
- 8:24
went well and what are some places where
- 8:26
things did not go well. This is where
- 8:28
you can actually titrate your your you
- 8:31
know and actually look in inside what's
- 8:33
actually happening uh at the session
- 8:35
level and annotate these data sets. Um
- 8:39
and as Surup mentioned, we have a lot of
- 8:41
use cases. we have we we talked to
- 8:43
multiple different teams who have uh
- 8:46
various uh ways of annotating uh their
- 8:49
data sets uh and it's it's almost hard
- 8:52
for a platform team to you know build
- 8:55
like a UI specific uh for each use case
- 8:58
uh and and you know to give you an
- 8:59
example uh it's usually going to be an
- 9:02
annotator who's going to annotate these
- 9:05
data sets so the platform team is you
- 9:07
know in charge of the APIs we have a
- 9:09
strategy and of person who's actually
- 9:12
deciding what to annotate and then you
- 9:13
have an annotator who's actually going
- 9:15
to annotate uh your data set. So we took
- 9:18
this approach uh everybody uh has uh you
- 9:21
know access to coding agents uh and we
- 9:23
actually doubled down on that API first
- 9:25
approach. So because we had these APIs
- 9:27
we were actually uh able to enable our
- 9:30
statops teams to use something like a
- 9:32
codeex or a claw code and v code their
- 9:35
own annotation UIs. Uh so we had
- 9:39
different use cases. Uh I think we had a
- 9:41
talk from Ragav before. Uh we had image
- 9:44
annotation use cases. We had some uh you
- 9:46
know manual testing use cases. What
- 9:49
stood out to us was the underlying
- 9:51
patterns were similar. So if we are are
- 9:54
API first uh we can actually enable our
- 9:57
our our partners to simply v code these
- 9:59
UIs for annotation. So it's it's like a
- 10:02
very simple example then you know of of
- 10:05
a vibe coded UI looks pretty clean does
- 10:07
the job uh and you get you know the
- 10:10
annotation that you eat this is
- 10:11
basically like a menu from a restaurant
- 10:13
uh it's it's you know nothing crazy uh
- 10:16
but the point I want to make here is
- 10:18
that what helped us was to give this
- 10:22
workflow in the hands of the operators
- 10:24
so that they can actually build their
- 10:26
own vcoded annotation UIs. Uh so moving
- 10:29
on once you have these annotation UIs
- 10:31
you obviously want to you know calibrate
- 10:33
your your your judge prompts you
- 10:35
obviously have some LM as a judge uh
- 10:37
metric that you're tracking you want to
- 10:39
now start improving that with these
- 10:41
golden data sets
- 10:43
u in order to do that uh you know we
- 10:46
have a pretty simple process uh you
- 10:48
you're going to start with you know some
- 10:49
judge prompt take a look at you know
- 10:51
what exactly do you want to measure from
- 10:53
the output uh g you know have have us
- 10:56
have something simple you're going to
- 10:58
have your baseline scores uh where
- 11:00
you're going to simply run those LLM
- 11:02
judges on your traces and then you're
- 11:04
going to have that optimization loop. Uh
- 11:06
so we use uh the JPEA library which is a
- 11:09
pretty commonly used library out there
- 11:11
for prompt optimization. Uh and once you
- 11:14
know the the iteration loop is complete
- 11:16
uh our partner teams are happy they're
- 11:18
going to then elevate that judge prompt
- 11:21
as their LLM as a judge. Now even while
- 11:24
doing that uh LLM as a judge as a
- 11:26
concept the whole prompt calibration
- 11:28
concept might be uh straightforward to a
- 11:31
lot of folks but it is still like a
- 11:32
pretty new and evolving field. Uh and
- 11:35
what we wanted to do was really reduce
- 11:37
the friction of back and forth with an
- 11:39
engineering team. So we tried to really
- 11:42
remove all the complicated logic and
- 11:45
make this into a self-s serve UI. So the
- 11:47
screenshot that you actually see is what
- 11:48
actually exists. uh so uh you know like
- 11:51
a product manager or an operator is
- 11:53
going to come to our UI. They're going
- 11:55
to set some of these configs uh on the
- 11:57
platform and then actually run the
- 11:59
calibration loop themselves. So they
- 12:01
don't have to worry about the different
- 12:03
settings that they need to worry about
- 12:05
what are the different uh you know
- 12:07
tweaks that they need to do and they can
- 12:08
actually like you know run a calibration
- 12:10
loop using any model of their choice. I
- 12:12
think in this example I have Gemini they
- 12:14
can use run it using uh you know any of
- 12:16
the claude or the openi models too. The
- 12:19
other important piece was actually uh
- 12:21
making this reviewable. Uh you know
- 12:23
again a lot of this uh is a closed box
- 12:26
where you can't really it's hard to see
- 12:28
what's actually happening. Uh so the
- 12:30
second piece that we built was actually
- 12:32
giving them vis visualization and
- 12:34
visibility into what's actually
- 12:35
happening. So on the left you can see we
- 12:38
and this is like one of the good
- 12:39
examples where we saw like a significant
- 12:41
amount of improvement in the judge
- 12:44
prompt. uh and we actually show the you
- 12:47
know the the previous the original
- 12:49
system prompt and the calibrated prompt
- 12:51
to our partners so that they are also
- 12:53
able to gain that trust uh why as as as
- 12:57
we build this
- 12:59
[clears throat]
- 12:59
>> yeah just want to add to that is this
- 13:02
enables different configurations in
- 13:04
different teams in some teams you have
- 13:06
seen the strategy and operations folks
- 13:08
own the prompt uh you have seen some
- 13:10
teams where the product manager owns the
- 13:11
prompt you have seen some teams where
- 13:13
engineering owns the prompt so this
- 13:14
gives gives the flexibility for teams to
- 13:16
design and evolve because we are all
- 13:18
learning. So the even the org uh design
- 13:20
is improving and we are enabling that.
- 13:23
>> Yeah, that that's a good point. I think
- 13:25
the overall idea was to you know build
- 13:27
something which is as self-s served as
- 13:28
possible so that uh you know people
- 13:31
aren't always necessarily blocked by our
- 13:33
team helping them out. Um and then
- 13:36
finally you know uh the quality loop in
- 13:38
practice. you know as we've been going
- 13:39
through this exercise we've seen a lot
- 13:41
of improvements happening to our product
- 13:43
as well. So you know for example we sort
- 13:46
mentioned we started with the UIs we are
- 13:48
you know now API and workflow first uh
- 13:51
we're trying to reuse a lot of the
- 13:53
existing infrastructure that already
- 13:54
existed at Door Dash uh and that's
- 13:57
helped us uh get a long way. Now some of
- 14:00
uh we we we've seen obviously like you
- 14:02
know really good results. I think a very
- 14:04
good result that we we do like to call
- 14:06
out is we actually did see a lot of
- 14:08
reduction in the spend uh at per
- 14:11
annotation cost as you as you all can
- 14:13
imagine we do have you know thousands of
- 14:16
rows that need to get annotated every
- 14:18
week uh and it can get pretty expensive
- 14:20
at doash scale uh and having this
- 14:23
selfserve uh annotation platform really
- 14:26
helped us reduce increase the velocity
- 14:29
and reduce the cost that we were
- 14:30
actually spending with these annotators.
- 14:33
to to annotate the data for us. Uh
- 14:35
obviously uh this resulted in faster
- 14:38
loops. Uh teams were able to iterate
- 14:40
faster. They were able to uh you know
- 14:43
calibrate their own judges in a
- 14:45
completely self-s served way. Uh and
- 14:47
thus it has resulted us in in in moving
- 14:50
with a very very high velocity.
- 14:54
So uh finally I just wanted to you know
- 14:56
quickly touch on this slide again uh the
- 14:59
eight steps you know continuous loop uh
- 15:01
which is you know you you have your
- 15:03
traces you want to look at your traces
- 15:05
your sessions you want to sample it down
- 15:07
to a size which is which you are
- 15:09
comfortable with uh you want to start
- 15:11
annotating your data sets you really
- 15:13
want to start uh making the data better
- 15:17
with the human knowledge that exists and
- 15:19
the domain knowledge that exists and
- 15:21
then calibrate your workflows was
- 15:23
calibrate your agents, calibrate your
- 15:24
LLM judges with this golden data set and
- 15:28
then repeat this whole cycle uh over a
- 15:30
period of time to you know to ship
- 15:32
reliably and ship with high quality. Um
- 15:36
yeah, we have 4 minutes left. Thank you
- 15:38
once again. I think that was the last
- 15:40
slide. Uh thanks for attending and if
- 15:42
there's any questions, we'd be happy to
- 15:43
hang out after the talk or even happy to
- 15:46
answer them now.
- 15:51
>> [applause]