AI Engineer World's Fair 2026
How do you diffuse AI into the real world? — Varun Shenoy, Long Lake
Read the talk
How Do You Diffuse AI Into the Real World?
Varun Shenoy explains why capable models still fail to change services businesses—and why earned autonomy, operational traces, real outcome-based evaluations, and in-person workflow redesign must advance together.
From a talk by Varun Shenoy
At a glance
Ideas worth remembering
Model capability and operational diffusion are different achievements. Deployment requires changing equipment, workflows, training, and incentives around the model.
Increase autonomy progressively—from copilot through synchronous, asynchronous, and long-running work—rather than beginning with an unproven AI coworker.
Asynchronous services agents need industry-specific ways to represent, fork, execute, and review work; coding’s sandbox-and-pull-request pattern does not transfer automatically.
Operational traces become valuable when they connect agent behavior to real outcomes such as a repaired roof, closed books, or the final data an employee submitted.
Continual learning and adoption form one loop: use supplies improvement data, while improvement must make the agent worth using. Initial enablement is the hard prerequisite.
Services deployment depends on proximity to the work—embedding tools in existing systems, training employees directly, and observing the exceptions that formal process descriptions omit.
The demo is not diffusion
AI can book a flight, complete a support ticket, or produce code ready for review. Those demonstrations now feel plausible. Yet Varun Shenoy asks the audience to enter a 200-person property-management firm, where real employees manage real properties, customers, and money. There, he says, the work may look essentially unchanged. The talk begins with that discrepancy: model capability has arrived faster than operational adoption. 0:43
Shenoy treats that gap as normal for a general-purpose technology. His electricity analogy separates invention from diffusion: electrifying a factory required replacing motors and equipment, reorganizing production, and training workers—not merely connecting an existing process to a new power source. The historical analogy supports his main operating model: deploying AI means changing the surrounding system of work. His claim that diffusion takes a generation, and that AI diffusion will be a defining problem for the next 20 years, is a forecast rather than a measured timetable. 1:34
Long Lake approaches that problem through ownership rather than conventional software sales. Shenoy reports that the company had acquired 35 services businesses across areas including HOA and property management, architecture, and HR services. He describes a roughly 40-person Long Lake team spanning technology, finance, and operations, with more than half focused on products, data, and field deployment. He also cites a recently announced $6.3 billion take-private of American Express Global Business Travel to convey the scale of its operating environment. These figures describe the company at the time of the talk; they do not by themselves establish the effectiveness of its AI deployments. 3:01
Ownership changes the feedback mechanism. A software vendor can attribute weak adoption to a customer’s processes, implementation, or training. Long Lake cannot make that separation inside businesses it owns and operates: when an agent fails to produce the intended result, the failure remains its problem. Shenoy organizes the rest of the talk around three linked questions: how to increase agent autonomy, how to learn from operational data, and how to make improvement compound through enterprise learning loops. 4:17
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Autonomy must be earned one rung at a time
Shenoy describes autonomy as a ladder rather than a binary choice between chatbot and coworker. Each rung changes who initiates work, how long execution lasts, and how much supervision remains:
- Copilot: A user asks a question and receives information, potentially through retrieval and connected systems.
- Synchronous agent: The user initiates a task and interacts with an agent that can call tools and work for roughly one to five minutes.
- Asynchronous agent: Work continues in the background and returns later; an external event or job queue can trigger it without a fresh user request.
- Long-running agent: Execution stretches across hours, days, weeks, or months.
- AI coworker: A proactive partner works alongside the employee and takes broader responsibility for getting work done. 5:21
The temptation is to start with the coworker because it is the easiest product vision to sell. Long Lake’s operating experience leads Shenoy to the opposite sequence: an agent must earn the right to do more. Some tasks still exceed model capability, and employees need time and close field support to understand where the system works. Gradual expansion also gives the deployment team opportunities to observe failures before granting wider autonomy. The talk does not specify formal promotion thresholds between rungs, so “earning” autonomy remains an operating principle rather than a quantified policy. 6:48
Answers a user request and retrieves information.
Each rung extends the agent’s initiative or task horizon. Long Lake’s argument is to move upward only after the system proves useful in the field.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Services work lacks coding’s asynchronous primitive
Coding makes the ladder concrete. A synchronous coding agent can inspect a filesystem, modify code, call tools, and receive immediate human feedback. To make it asynchronous, the developer can place essentially the same agent in an isolated sandbox, let it build and test, and receive a pull request when the task finishes. Shenoy describes this pattern as largely solved relative to the corresponding problem in services. 7:33
Engineers also already accept out-of-order completion. They may launch 10 jobs and remain comfortable when job seven finishes before job three. That habit matches asynchronous agents: each task can occupy a separate sandbox, while the person reviews results as they arrive. Traditional services work usually lacks both the sandbox and the habit. Employees clear an inbox one message at a time, and many business processes assume serial ownership of a case. 8:03
A synchronous services agent is easier to imagine: give it enterprise context, custom tools, integrations, and potentially MCP connections, then let an employee chat with it. The harder frontier is an asynchronous services agent that can safely fork real work. Code has filesystems, sandboxes, builds, tests, and pull requests as natural execution and review boundaries. Architecture and property management need their own equivalents: task representations, isolated workspaces, completion evidence, and interfaces for reviewing results. Shenoy poses this as an open problem, not something Long Lake claims to have fully solved. 8:33
His proposed research direction begins with a model capability that already works well: code generation. Instead of waiting for models to learn every services workflow directly, Long Lake asks whether knowledge work can be represented as code and executed by coding agents. But representation is only one part of the problem. The product must also determine which serial workflows can be parallelized and which interface fits each industry. A launch mechanism that works for code will not automatically fit architecture or property management. 9:24
A bounded request targets files in a repository.
Coding already supplies a forkable environment and a review artifact. Services workflows need industry-specific equivalents before parallel delegation feels natural.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Real work produces traces, ground truth, and regression tests
The most valuable services tasks are often absent from public training data. Shenoy’s examples are deliberately mundane and difficult: closing the books when receipts are missing, scoping construction from a blueprint, and coordinating vendors to repair a broken roof. The procedure may live in a senior employee’s memory, a 20-year-old application, or an informal sequence of judgments. Before an agent can learn from this work, the organization must make the task and its hidden decisions explicit. 10:24
Long Lake’s proposed flywheel begins by having agents collaborate with employees on actual work. That interaction generates traces containing tool calls, corrections, exceptions, small frictions, and failed attempts. The trace is useful because it connects model behavior to an operational outcome. For the roofing workflow, the final question is whether the roof was repaired; for accounting, whether the books closed. Those outcomes provide stronger ground truth than whether a response merely sounded convincing. 10:53
Each observed failure can become an evaluation case. Shenoy says Long Lake hill-climbs against these benchmarks and turns each week’s benchmark improvements into regression tests. The intended ratchet is straightforward: real work exposes a failure, the team changes the agent, and the retained case checks that later changes do not reintroduce the same problem. The talk does not report benchmark sizes, pass rates, or measured business gains, so it explains the feedback architecture rather than demonstrating its quantitative effectiveness. 11:23
The traces support three parallel improvement mechanisms:
- Evaluation signals: Explicit feedback includes ratings and written notes. Implicit feedback includes the difference between AI-generated data and the version an employee ultimately submitted.
- Post-training data: Shenoy says Long Lake has begun internally post-training models on proprietary operational data that is out of distribution for many general-purpose models. He provides no comparative results for those models.
- Agent customization: The surrounding agent must account for differences among companies, individual employees, and clients. In a services business, those variations are often part of the expected service rather than noise to discard. 11:53
This is why polished demonstrations misrepresent the shape of the work. A benchmark task may look like a clear downhill route with success visible from the start. Operational work contains missing receipts, unusual client expectations, software quirks, vendor delays, and exceptions layered on exceptions. Shenoy’s sharpest formulation is that “the exceptions are the job.” In other words, handling the nominal path demonstrates capability; surviving the accumulated edge cases determines whether the system can operate. 13:04
Agents participate in real services work alongside employees.
Operational work creates traces and outcome evidence. Evaluations preserve failures, improvements return to production, and broader use produces more evidence.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Continual learning and adoption are one loop
Enterprise teams often separate continual learning from enablement. Research or platform engineering owns improvements to prompts or weights, while growth, deployment, or customer-experience teams try to persuade employees to use the product. Shenoy argues that this organizational split hides a dependency: the agent only learns from operational use, while employees only keep using it if learning makes it better. 13:43
The resulting loop is a snowball: more usage produces feedback, feedback improves the agent, and a better agent earns more usage. But a loop with no initial motion produces nothing. Giving a capable tool to an entire company does not guarantee adoption, especially in a 100-year-old firm where an employee may have closed the books the same way for 20 years. Technical quality cannot generate training traces if the product never enters the workflow. 14:29
Shenoy calls the response extreme software-service co-design: develop the software together with the people and processes that deliver the service. The analogy is to hardware-software co-design, where neither layer can be optimized independently. Here the service operation is part of the system. Product design, deployment, workflow changes, employee training, feedback collection, and model improvement must inform one another rather than pass work between isolated teams. 15:29
Employees must first incorporate the agent into real work.
Usage is both the source of improvement data and the result of improvement. Initial enablement is therefore a necessary input, not an afterthought.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Co-design means reducing workflow distance—and showing up
Software-service co-design starts by lowering the energy required to try the system. That can mean embedding the product where employees already work: Excel, an ERP, 3D design software, email, or other familiar tools. Asking people to leave their operational system for a separate AI interface adds friction before the agent has demonstrated enough value to justify a process change. 15:59
It also means observing work in person. Shenoy’s examples are intentionally unglamorous: lunch-and-learns, one-on-one or two-on-one training, running a cotton-candy stand at a company conference, and asking employees about day-to-day difficulties while mountain biking. These activities create opportunities to see tacit procedures and small frustrations that a support ticket or scheduled video call may never expose. They also let the deployment team teach the product while collecting immediate feedback. 16:29
The talk ends by making physical proximity part of the technical method. Shenoy argues that a services business cannot be co-designed adequately over Zoom or through support tickets alone. That is a strong operating judgment, not a controlled comparison of remote and on-site deployment. Its practical force comes from the rest of the talk: if valuable workflows live in people’s heads, if exceptions define the real job, and if improvement depends on traces from actual usage, then the deployment team must get close enough to see how work really happens. In his closing phrase, AI diffusion requires teams to “touch some grass.” 16:59
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
- Varun ShenoyReference
Shenoy’s personal site, supplied with the recording’s speaker information.
- Varun Shenoy on XReference
A supplied speaker profile with additional posts about Long Lake’s work in services businesses.
Related talks
- Agents for Everything Else — swyx
Explores how coding-style agents can extend into non-code knowledge work, directly complementing Shenoy’s search for an asynchronous services primitive.
- Improving Agents is a Data Mining Problem
Develops the trace-analysis and production-data side of the continual improvement loop described in this talk.
- AI tools for Forward Deployed Engineering
Examines the hands-on workflow discovery and redesign required when enterprise automation depends on customer-specific operational context.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Hi everyone. I'm Varun. I'm one of the
- 0:14
co-founders at Long Lake and I'm excited
- 0:17
to share a little bit about what we've
- 0:19
been up to for the last 2 years.
- 0:21
It all comes back to a question all of
- 0:24
us have asked time and time again.
- 0:27
The models are getting better,
- 0:29
but the real question is how do you
- 0:31
actually deploy the AI into the real
- 0:34
world?
- 0:35
How do you get the models to complete
- 0:37
economically relevant tasks?
- 0:43
Let me start by saying everyone has seen
- 0:45
the demo. Think of the agent
- 0:47
automatically booking a flight, the
- 0:50
agent automatically completing a ticket
- 0:53
in some kind of customer service portal.
- 0:55
Think of an agent completing a block of
- 0:57
code ready to commit and go.
- 1:00
The reality is we've all seen this and
- 1:02
it feels like magic. 2 years ago any of
- 1:06
this would have been complete science
- 1:08
fiction. The capabilities are real.
- 1:12
Now, walk with me into a 200-person
- 1:15
property management firm.
- 1:16
Real people, real properties,
- 1:19
real dollars, real customers all across
- 1:22
the US.
- 1:24
You would expect AI to show up by now,
- 1:27
but the reality is nothing has changed
- 1:30
at all.
- 1:34
Here's the thing.
- 1:36
This is totally normal and maybe in fact
- 1:38
I'd argue this is what we should expect.
- 1:42
This is true for every general-purpose
- 1:45
technology. You know, take electricity
- 1:47
for example.
- 1:48
Electricity was invented in the 1880s
- 1:51
and it was first demoed at Edison's
- 1:53
Pearl Street Station Dynamo Room over in
- 1:55
Manhattan.
- 1:57
This was the magic demo of its time.
- 2:01
The reality is it took a long time for
- 2:04
electricity to be fully adopted.
- 2:07
Consider a Ford factory.
- 2:09
It's not enough to just have
- 2:11
electricity. You have to rip out the
- 2:13
existing motors and equipment. You have
- 2:15
to bring in the new equipment. You have
- 2:17
to go and train everybody to use that
- 2:19
very same equipment.
- 2:21
Here's a picture of a Ford electrified
- 2:23
moving assembly in 1924.
- 2:26
These things take time.
- 2:28
Diffusion of any technology takes a
- 2:31
generation. And since everyone here in
- 2:34
this room today is talking about AI, I
- 2:36
would argue
- 2:37
AI diffusion is perhaps the single most
- 2:40
important problem for the next 20 years.
- 2:44
The models are going to keep getting
- 2:45
better. The big question is how do we
- 2:48
actually get these models to be in the
- 2:50
real world, complete real tasks, uh and
- 2:52
make people more efficient, happier, and
- 2:54
provide better service.
- 2:57
So taking a quick step step back, who
- 2:59
are we? Uh we are Long Lake. Over the
- 3:01
last 2 years, we've raised over $3
- 3:04
billion from Elad Gil, General Catalyst,
- 3:07
and AlphaWave since our founding.
- 3:10
Here's the strange part. We we don't
- 3:12
sell software. We actually go out and
- 3:14
acquire and partner with real services
- 3:16
businesses in the world. Uh we've
- 3:18
acquired 35 businesses across HOA and
- 3:21
property management, architecture, HR
- 3:23
services, and a lot more.
- 3:26
To give you a little bit more flavor, we
- 3:28
have roughly a 40% team right now split
- 3:30
between technology, finance, and
- 3:32
operations. More than half our team is
- 3:35
part of the technology team focused on
- 3:37
uh building products, data, and
- 3:39
deploying the core products into the
- 3:41
field. Uh we're in a collected group of
- 3:43
folks, a bunch of ex-founders who've
- 3:46
worked in the services before,
- 3:47
ex-military, folks from Palantir, Ramp,
- 3:50
Glean, uh and from the finance side,
- 3:53
Blackstone, H.I.G., et cetera. We are we
- 3:56
are not selling them to these companies
- 3:58
above from the outside. We're actually
- 4:00
deploying into these companies and
- 4:02
figuring out how to get the technology
- 4:03
to work.
- 4:05
And just to show you the scale we're
- 4:06
playing at, we announced recently our
- 4:08
$6.3 billion take private of American
- 4:11
Express Global Business Travel, the
- 4:12
world's largest corporate travel
- 4:14
platform.
- 4:16
We own these businesses.
- 4:19
So, when the AI doesn't work, it's not
- 4:22
their problem. We're not the vendor.
- 4:24
It's our problem.
- 4:27
Concretely, again, we are not the
- 4:29
vendor. We are the operator owners. And
- 4:31
we work very closely with our teams
- 4:33
within the businesses to drive real
- 4:35
outcomes.
- 4:37
Now, I want to step back and get to the
- 4:39
concrete about the how. What are the
- 4:41
lessons we've learned over the last 2
- 4:43
and 1/2 years? And what we've learned
- 4:45
from deploying AI into companies we've
- 4:47
owned.
- 4:51
Three quick lessons. One, how we move
- 4:54
agents from co-pilots to co-workers.
- 4:57
Two, how we leverage real-world data
- 5:00
within these businesses.
- 5:02
Remember, we're seeing all of the work
- 5:04
that's being done in these real services
- 5:06
businesses. There's a lot of interesting
- 5:08
problems and solutions embedded within
- 5:11
that.
- 5:11
And then finally, perhaps the most
- 5:13
interesting and exciting is how do you
- 5:15
actually get all of this technology to
- 5:16
compound over time by learning loops in
- 5:19
the enterprise. We'll get to that at the
- 5:21
end over here.
- 5:23
So, starting off from co-pilots to
- 5:25
co-workers.
- 5:27
There's a spectrum of how much autonomy
- 5:29
you can give an agent.
- 5:31
On the left here, you see a co-pilot.
- 5:32
This is, you know, your simple rag
- 5:34
chatbot from 2 years ago. It's very
- 5:36
quick. You can ask a question. Maybe
- 5:38
it's integrated with some systems. It
- 5:39
can give you information back very, very
- 5:41
quickly.
- 5:44
The second step is a synchronous agent.
- 5:45
Consider something like Claude code,
- 5:47
Codex, Claude co-work. It's real-time.
- 5:50
There's this two-way interaction. It's a
- 5:52
bit more sophisticated than a co-pilot.
- 5:54
You can go let it run off for 1 to 5
- 5:56
minutes. Uh it'll call tools, maybe use
- 5:58
its skills. Uh it's still synchronous.
- 6:00
You still need to step in and ask a
- 6:01
query. So, the next obvious rung of the
- 6:03
ladder is the asynchronous agent.
- 6:06
You can come in here, still ask a query.
- 6:08
The agent will go off into the
- 6:10
background, do some work, and then come
- 6:11
back. Uh and what's really interesting
- 6:14
about asynchronous agents is that the
- 6:15
user does not have to be the one that
- 6:17
triggers them.
- 6:19
You can have external triggers as well.
- 6:21
Maybe someone completes a certain task
- 6:23
and there is an async job queue uh that
- 6:25
allows the async agent to pull off from
- 6:27
and proactively offer advice to the end
- 6:29
user.
- 6:31
Then, I'd argue the next step is a
- 6:33
long-running agent.
- 6:35
How do you get these agents to work for
- 6:37
hours, days, weeks, months, etc.? I
- 6:40
think this is currently a very core
- 6:42
problem that a lot of the labs are
- 6:44
focused on, as are we.
- 6:48
And then finally, at the end, the holy
- 6:50
grail, an AI co-worker.
- 6:53
This is where most people start off.
- 6:55
You want a proactive partner that gets
- 6:57
work done just alongside you.
- 7:00
This is what everyone wants to sell you,
- 7:02
but what we've learned from owning the
- 7:03
outcomes in this business is
- 7:07
you have to earn the right to do more.
- 7:09
It's it's not enough to jump to the
- 7:11
co-worker immediately,
- 7:13
right? For for a bunch of reasons. One,
- 7:15
for certain tasks, the models might not
- 7:17
quite be there yet. And two, you
- 7:19
actually have to work with these
- 7:20
companies in the field, interact and
- 7:23
iterate very, very closely, so that they
- 7:26
understand that this is the beginning of
- 7:28
AI, and you can work up the rungs over
- 7:30
time.
- 7:33
I think a really unique lens to look at
- 7:35
this problem through is the that of the
- 7:37
jagged frontier. We all know that agents
- 7:40
are incredibly good at writing code. So,
- 7:42
what does the, for example, synchronous
- 7:44
agent for code generation look like?
- 7:47
This is super simple. This is just your
- 7:48
coding agent, maybe it's Codex, Cloud
- 7:50
Code, just running on your desktop. It
- 7:52
has access to a file system. You
- 7:54
collaborate within real time. You get
- 7:56
instant feedback and you iterate.
- 7:59
The next step is, you know, if you look
- 8:01
at code code generation, what is the
- 8:02
async agent? This is also fairly
- 8:04
straightforward and largely solved. You
- 8:06
take the exact same coding agent, you
- 8:08
wrap it in a sandbox, and you just let
- 8:10
it go run. It can build, it can test,
- 8:12
and once it's done with its work, it can
- 8:14
provide the code in the form of a PR.
- 8:16
One thing that's really unique about
- 8:18
engineers is folks are incredibly good
- 8:21
at already paralyzing their work.
- 8:23
It's very commonplace to launch 10 jobs
- 8:26
and be comfortable with the fact that
- 8:28
job seven might finish before job three.
- 8:30
So, engineers are incredibly good at
- 8:32
using these async agents.
- 8:35
Now, when we come to services, the
- 8:37
equivalent of a synchronous agent, what
- 8:38
we talked about a little bit earlier,
- 8:40
it's a co-working agent. It's an agent
- 8:42
that has deep context about your
- 8:43
enterprise. It interacts potentially
- 8:45
with MCPs, custom tools, custom
- 8:47
integrations, uh and you can chat with
- 8:49
it synchronously just like any of these
- 8:51
other products.
- 8:53
I think this is a frontier here in the
- 8:54
bottom right.
- 8:56
What does it mean to build an
- 8:58
asynchronous agent for the services?
- 9:00
What does it mean to paralyze work in
- 9:03
industries where work is traditionally
- 9:05
done in a very, very serial manner?
- 9:08
This is where we spend a lot of time and
- 9:10
this is what I wake up every morning
- 9:11
really excited thinking about, you know,
- 9:12
we've we've figured out what the async
- 9:15
and forking mechanism for code is. You
- 9:17
just spin up a bunch of sandboxes and do
- 9:19
work. What does that look like for the
- 9:21
rest of the world?
- 9:24
So, here's a couple questions we think
- 9:26
about pretty seriously. One, you know,
- 9:28
the models are trained on code, they
- 9:30
want to write code, they're incredibly
- 9:31
good at writing code. How do we leverage
- 9:33
these coding agents for actual knowledge
- 9:34
work? You You rather than wait for the
- 9:36
models to catch up on doing services
- 9:38
knowledge work, what if we just use that
- 9:40
code knowledge and represent knowledge
- 9:42
work as code?
- 9:44
Two, as I mentioned, engineers are used
- 9:46
to paralyzing work. How do you paralyze
- 9:48
work that's traditionally serial? You
- 9:49
know, people clean out their inbox one
- 9:51
email by one email, not 10 emails at
- 9:52
once.
- 9:54
And finally, how do you move up the
- 9:55
ladder here both in terms of product and
- 9:57
user enablement?
- 9:59
What are the right form factors? And I'd
- 10:01
argue this varies dramatically from
- 10:04
industry to industry. Just because you
- 10:06
have one way of launching an async agent
- 10:07
for code, doesn't mean that same way is
- 10:09
going to work for architecture or
- 10:10
property management.
- 10:13
The second point I want to cover today
- 10:15
is leveraging real-world data.
- 10:17
We all know this. Frontier models have
- 10:19
learned from everything humanity has
- 10:21
written down,
- 10:22
but the most valuable tasks are not on
- 10:24
the internet.
- 10:26
How do you actually close the books when
- 10:27
you're missing receipts?
- 10:29
>> [snorts]
- 10:29
>> How do you scope a building for
- 10:31
construction in a blueprint, potentially
- 10:33
collaboratively?
- 10:35
How do you coordinate vendors for fixing
- 10:36
a broken roof?
- 10:39
All of this knowledge lives in people's
- 10:40
heads, in 20-year-old software, uh in
- 10:43
the way that one senior person on one of
- 10:45
these teams just knows how to do it. How
- 10:48
do you make this information explicit
- 10:49
and create tasks that you can actually
- 10:51
learn from?
- 10:52
So, we've constructed a little bit of a
- 10:53
flywheel. We get our agents to
- 10:56
collaborate with our employees to do
- 10:58
real work. And this allows us to
- 11:00
generate rich traces of data and
- 11:01
information. Tool calls, the hiccups,
- 11:04
the papercuts, everything that goes
- 11:06
wrong with doing real work.
- 11:08
This in turn allows us to build
- 11:10
real-world evals.
- 11:12
There is a ground truth here. In the
- 11:14
case of the roofing example, the
- 11:16
question is, did the roof get repaired?
- 11:19
Did the books get closed?
- 11:21
And this allows us to hill climb and
- 11:23
build better agents, which leads to more
- 11:25
and more impact. And what's really
- 11:27
exciting is it ratchets up. Every week
- 11:30
our hill climbing benchmarks
- 11:32
become a regression test. So, our agents
- 11:34
get better and better over time.
- 11:37
Just to drive a little bit deeper here
- 11:39
on the traces, there's three upshots of
- 11:42
being able to collect these rich traces.
- 11:44
One, we get to generate amazing evals
- 11:46
that are built and scored automatically.
- 11:49
Uh and we're able to gather both
- 11:51
implicit and explicit feedback. Explicit
- 11:53
feedback in the sense of thumbs ups and
- 11:54
thumbs down, maybe people provide a note
- 11:57
telling us whether this response was
- 11:58
good or not. Uh and also implicit
- 12:00
feedback. Right? Again, we have the
- 12:02
ground truth. Maybe there's some data
- 12:03
that the AI generated and there's a real
- 12:05
diff between the data that the AI
- 12:07
generated and what was ultimately
- 12:09
submitted. That's rich information that
- 12:12
almost no one else has.
- 12:14
Two, we've started post training models
- 12:16
internally on
- 12:18
all of the data that these businesses
- 12:19
operate on and produce, generally
- 12:22
speaking.
- 12:23
This is all data that is completely out
- 12:25
of distribution for most frontier labs.
- 12:27
Think of the task I showed at the
- 12:28
beginning. A lot of the models A lot of
- 12:31
the frontier models today just can't do
- 12:33
these tasks yet and we're trying to post
- 12:35
train our own models internally to be
- 12:37
able to do that on the rich source of
- 12:38
data that we own.
- 12:40
And then finally, the actual agents
- 12:42
themselves.
- 12:43
The real world is incredibly hairy and
- 12:45
messy and you want customization per
- 12:47
company. Every company does things very
- 12:50
differently. Customization per user. The
- 12:52
way each user does their work is very
- 12:54
unique. And customization per client.
- 12:57
The way you work with every client is
- 12:59
different. It's a services business and
- 13:01
you want to uphold those standards.
- 13:04
I love this picture
- 13:06
because it's the whole thing in a single
- 13:08
image. Um the the way we usually talk
- 13:10
about LLM tasks is the top panel. Right?
- 13:13
You just It's It's a slope. You got a
- 13:15
bike. And but there's clear sight to
- 13:18
success.
- 13:19
The reality is most work is not like
- 13:22
that. And And you and I both know that.
- 13:24
Uh there are hills and ravines. Uh
- 13:27
there's death by a thousand paper cuts.
- 13:29
But But that's what real work looks
- 13:31
like. That's the entire job. The
- 13:33
exceptions are the job.
- 13:37
That's That's the demo.
- 13:39
That's the actual job.
- 13:43
Now, on to the final thing I want to
- 13:44
chat with you guys today is learning
- 13:47
loops within the enterprise.
- 13:49
I'd argue there's two hot trends
- 13:51
everyone's talking about in 2026. One,
- 13:54
it's continual learning. How do you make
- 13:56
an agent better over time with feedback?
- 13:58
I think there are plenty of sessions uh
- 14:00
this week on how you can use continual
- 14:02
learning, whether it's in the prompt or
- 14:04
in the weights.
- 14:06
And two, enablement. How do you get in
- 14:08
these enterprises and actually get them
- 14:10
to adopt and use AI?
- 14:12
Traditionally speaking,
- 14:15
these two initiatives are owned by two
- 14:16
separate teams. Right? The continual
- 14:18
learning is owned by your research team,
- 14:20
your platform engineering team.
- 14:21
Enablement's owned by growth or
- 14:23
deployment or customer experience. Uh
- 14:25
usually pretty siloed, not much
- 14:26
interaction between the two.
- 14:29
We think these are part of the exact
- 14:31
same loop.
- 14:32
The agent only improves if people
- 14:34
actually use it.
- 14:37
And people only use the agent if it's
- 14:39
worth adopting.
- 14:41
So, here's a little graphic of a
- 14:42
snowball. More usage drives continual
- 14:45
learning, which drives a better agent,
- 14:46
which drives more usage again.
- 14:49
All this to say, there's still a really
- 14:51
big elephant in the room.
- 14:53
How do you get the initial usage?
- 14:55
I think a lot of people, you know, will
- 14:57
use Claude Code or or give it to their
- 15:00
whole enterprise, expect folks to just
- 15:02
start using it.
- 15:04
Everyone assumes the usage just shows
- 15:07
up.
- 15:08
But as we all know, that's simply not
- 15:10
the case. It never does. Right? Getting
- 15:12
a 100-year-old firm to change its
- 15:15
processes is hard.
- 15:17
You could have the best AI coworker on
- 15:20
the internet or on Earth. And if the
- 15:22
people if the person who's closed the
- 15:24
books for the last 20 years continues to
- 15:26
do things the same way,
- 15:28
nothing changes. Nothing happens.
- 15:32
So, what can you actually do about it?
- 15:33
What you know, this this seems like
- 15:35
incredibly hard. What what's the upshot?
- 15:38
How do you actually get this stuff to
- 15:39
work? Well, I think a lot about Jensen
- 15:41
and how he dominated the market in his
- 15:44
words with extreme hardware software
- 15:46
co-design. Designing the chips and the
- 15:49
software together as one system.
- 15:52
We look at this through the lens of
- 15:54
extreme software service co-design. How
- 15:57
do you co-design our products with the
- 15:59
people and the processes at our
- 16:01
businesses? And I'd argue this is only
- 16:04
possible from being within under the
- 16:06
same roof.
- 16:08
We need to meet the people within these
- 16:10
companies both metaphorically, for
- 16:13
example, bringing products to their
- 16:15
systems so that the energy required for
- 16:17
enablement is kept low, and also
- 16:19
physically. Get on a plane, show up, say
- 16:22
hi, learn what people actually do.
- 16:25
You know, maybe you build a product
- 16:27
that's natively embedded into Excel or
- 16:29
into their ERP system, maybe their 3D
- 16:32
design software, or or maybe even their
- 16:34
Microsoft products like Outlook, Gmail,
- 16:36
etc.
- 16:38
Or you show up in person. You do a lunch
- 16:39
and learn with a bunch of folks at one
- 16:40
of the companies. You go to their
- 16:42
conferences and you create cotton candy
- 16:44
and run a stand for them. You go
- 16:45
mountain biking and ask them about all
- 16:47
the difficulties that they have with
- 16:48
their actual day-to-day jobs.
- 16:50
Or you show up in person one-on-one or
- 16:53
sometimes even two-on-one in this case
- 16:55
and just show them how to use the tools
- 16:57
and learn from the feedback because this
- 17:00
is what the rest of the world really
- 17:01
looks like. It's not like the folks in
- 17:02
this room or in San Francisco. It's a
- 17:04
lot more like this. You cannot co-design
- 17:07
software with the services business over
- 17:08
Zoom
- 17:09
or over a support ticket. You you have
- 17:12
to be there. You have to be in person.
- 17:14
And I'd argue this is the part that
- 17:16
actually makes it work. In order to get
- 17:17
AI diffusion to work, you have to touch
- 17:20
some grass.
- 17:22
Thank you so much. I'll be around for
- 17:23
the rest of day if there's anything I
- 17:24
can help with. My email is up there.
- 17:27
And yeah, thank you.
- 17:43
>> [music]