AI Engineer Code 2025
Claude Agent SDK [Full Workshop] — Thariq Shihipar, Anthropic
Read the talk
Building Agents With the Claude Agent SDK: Bash, Context, Verification, and Practical Tradeoffs
Thariq Shihipar explains how the Claude Agent SDK builds on Claude Code, why files and Bash make agents more adaptable, and how to design systems that gather context, act safely, and verify their work.
From a talk by Thariq Shihipar
At a glance
Ideas worth remembering
Design agents around gathering context, taking action, and verifying work, while treating verification as a continuous property of the entire loop rather than a final checkbox. 21:41
Choose structured tools for controlled atomic actions, Bash for composable filesystem and command-line operations, and code generation for dynamic API composition, while accounting for their different context and latency costs. 25:12
Make unfamiliar problems more accessible by exposing data through interfaces the model can already use, including spreadsheet ranges, SQL, searchable files, and progressively discoverable command-line scripts. 41:41
Protect powerful agents with layered defenses, scoped credentials, sandboxing, deterministic checks, and reversible checkpoints instead of assuming the model alone will enforce application security. 13:03
Control context growth by saving bulky outputs to files, delegating focused work to sub-agents, and reconstructing state from durable artifacts instead of repeatedly loading entire datasets or conversation histories. 29:45
Prototype directly in Claude Code, inspect real execution transcripts, refine instructions and helper scripts, and only then package the working behavior behind a small SDK entry point. 21:41
From fixed workflows to an opinionated agent harness
The progression Shihipar describes begins with individual language-model features, moves through structured workflows, and arrives at agents that build their own context and choose their own trajectories. A workflow can accept tightly defined inputs and produce tightly defined outputs, while an agent can interpret a natural-language request and decide which intermediate actions are necessary. The distinction is not absolute: an issue-triage workflow may still need an agent to clone a repository, start a Docker container, investigate a failure, and return a structured result. 1:35
The Claude Agent SDK is built on top of Claude Code because teams building agents repeatedly needed the same surrounding infrastructure. That infrastructure includes the model, a tool-running loop, agent and tool prompts, filesystem access, skills, sub-agents, web search, compacting, hooks, and memory. Shihipar presents the SDK as a packaged harness that absorbs these recurring implementation concerns so application developers can concentrate on domain-specific search, actions, and verification. 3:56
The architectural commitment is substantial: this style of agent needs Bash and a filesystem, so it runs locally or in a container rather than behaving like an ordinary stateless model request. Shihipar acknowledges the resulting sandbox, hosting, and performance overhead. His argument is that these inconveniences purchase a more capable execution environment, not that every deployment becomes simpler or that every problem requires the SDK. 8:36
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why Bash, files, and generated code expand what an agent can do
The central claim is that Bash is a general-purpose composition layer. Instead of defining a separate model-facing tool for every search, lint, execution, or transformation task, an agent can discover existing commands, invoke package-manager scripts, save intermediate results, generate reusable scripts, and operate software such as FFmpeg or LibreOffice. The filesystem makes those intermediate artifacts inspectable and reusable, turning context engineering into a question of tools, files, scripts, and state rather than prompt text alone. 5:06
A ride-sharing expense example makes the mechanism concrete. An email search alone might return a large collection of messages that the model must interpret directly; with Bash, the agent can save results, search for prices, add them, retain line numbers, and inspect whether each extracted value actually corresponds to a relevant charge. The same pattern extends to joining inbox and contact information or using FFmpeg and JQ to process a recorded meeting. The improvement comes from external computation, composability, and the ability to check intermediate work. 17:08
Shihipar distinguishes three execution options. Structured tools are reliable, controlled, and suitable for atomic or irreversible actions such as writing a file with approval or sending an email, but large tool inventories consume context and compose poorly. Bash lowers upfront context requirements and supports reusable commands, but discovering a command through its help interface introduces latency. Code generation supports dynamic scripts, API composition, data analysis, and flexible logic, but execution may require linting or compilation and is generally slower. Choosing among them is a system-design decision, not a universal rule. 25:12
Reliable, controlled atomic actions
Structured tools, Bash, and generated code expose different control, composition, and execution tradeoffs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Design the loop around context, action, and continuous verification
Shihipar proposes a practical three-part loop: gather context, take action, and verify the work. For a coding agent, gathering context might mean searching a repository; for an email agent, it means locating relevant messages. The model should generally discover useful information through the interfaces it has been given instead of receiving a static bundle that developers assume contains everything it needs. Planning can be inserted between context gathering and action, but it introduces additional latency and is not necessary for every task. 21:41
The quality of the verification surface strongly influences whether a problem is suitable for an agent. Code offers relatively strong checks because it can be linted, compiled, and executed, while research is harder to verify and may depend partly on source citations. Shihipar recommends deterministic checks wherever possible, including rejecting a file write when the agent has not first read that file. Verification should also appear throughout execution, not merely at the end: limits, validation errors, and intermediate checks can give the model actionable feedback and redirect its next attempt. 23:03
Hooks provide another way to introduce deterministic behavior or update context during the agent loop. A hook might validate a spreadsheet after an operation, insert changes a user made while the agent was working, or reject a final response that failed to consult the required data or generate a required script. These mechanisms are particularly useful when the model appears to know an answer already and might otherwise respond from existing knowledge rather than inspect the application’s authoritative inputs. 1:47:05
The framework should remain a guide rather than a rigid mandatory sequence. Shihipar describes expressing the desired behavior in a system prompt while allowing the agent to decide which steps actually apply. A read-only spreadsheet question, for example, does not need the same write-oriented verification as a modification. His recurring operational advice is to read agent transcripts repeatedly, identify why the model chose a particular path, and adjust prompts, interfaces, checks, or available tools accordingly. 21:41
Find relevant files, messages, or data
Gather context, act, verify, and use validation feedback to direct the next attempt.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Treat search interfaces and context budgets as product design problems
A spreadsheet illustrates why context gathering requires more than attaching a generic search tool. Finding revenue for a particular year is a multidimensional retrieval problem: headers alone do not identify the correct intersection of metric and year. Possible interfaces include familiar spreadsheet ranges, SQL queries, XML-oriented access, or command-line processing. Shihipar highlights translating a CSV-like source into an interface the model already understands, such as SQLite and SQL, as an example of making an unfamiliar business problem more compatible with existing model capabilities. 53:43
Search quality can also improve through preprocessing and annotation. An application might transform its source into a queryable representation or have another agent add descriptions and metadata before the main agent searches it. The best interface is domain-dependent, so Shihipar recommends trying multiple approaches against representative tests rather than assuming the first retrieval design is sufficient. Read and write interfaces can often share the same underlying abstractions, whether those abstractions are spreadsheet ranges, SQL, or XML. 57:25
Large datasets expose the limits of simply expanding the prompt. Shihipar explicitly notes that accuracy becomes harder to maintain as spreadsheets or codebases grow and advises against loading an entire spreadsheet into context. Instead, an agent can inspect a small initial view, search for relevant terms, navigate between sheets, and maintain a scratchpad or notes. Long tool outputs can likewise be written to files while the tool returns only the resulting path, allowing later search, processing, and rechecking without flooding the conversation. 29:45
Sub-agents offer another context-management strategy: delegate an extensive search or independent sheet summaries to separate workers and return only the useful result to the main agent. Shihipar also describes clearing coding-session context when the relevant state can be reconstructed from files and a git diff, while acknowledging that designing equivalent resets or summaries for less technical spreadsheet users is harder. For very large codebases, he identifies tradeoffs in bespoke semantic search and points instead to useful project instructions, an appropriate starting directory, hooks, and verification. 1:06:13
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build safety from layered controls, scoped access, and reversibility
Granting an agent Bash and filesystem access creates a genuine security problem, so Shihipar describes a layered defense rather than reliance on a single safeguard. The layers include model alignment, harness-level prompting and permissions, analysis of Bash commands through an AST parser, and sandbox restrictions on network and filesystem operations. Isolating an agent from personal machines or environments containing production secrets adds another boundary. The stated goal is to reduce what a compromised or misdirected agent can actually access or exfiltrate. 12:40
Database access illustrates the tradeoff between strict control and flexible exploration. A narrowly defined tool can expose only approved inputs and outputs when sensitive information must remain hidden, but that structure also limits dynamic querying. Bash or generated code can support iterative SQL development because the model can run a query, inspect an error, and revise it, provided access is guarded appropriately. Shihipar suggests scoped or temporary API keys, backend enforcement, and, where appropriate, proxies that insert credentials without exposing them directly to the agent. 44:13
Safety also depends on whether an action is reversible. Code is comparatively forgiving because version history and checkpoints can restore earlier states, whereas a mistaken interaction in a shopping flow can leave the interface in a more complicated state that requires additional corrective actions. For a spreadsheet or similar product, Shihipar recommends considering checkpoints and restoration mechanisms so users, and potentially agents, can recover from destructive mistakes. He also notes that coordinating parallel Bash-based sub-agents introduces practical concerns such as race conditions. 1:10:36
Independent verification can supplement these controls, but Shihipar ranks deterministic checks ahead of model-based review whenever the rules are clear. When a separate reviewing agent is useful, he recommends giving it a fresh context rather than copying the original agent’s full history, reducing the chance that the verifier inherits the assumptions or errors it is supposed to challenge. This preserves the distinction between a hard enforcement boundary and a probabilistic second opinion. 1:07:52
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Prototype with real data, then preserve what works in the SDK
The workshop’s prototyping approach starts with Claude Code, a real API, a few helper scripts or libraries, and project-level instructions. Shihipar argues that an effective agent should remain relatively small even though finding the right domain abstraction may be difficult. Rather than first constructing an elaborate orchestration layer, developers can observe how the model searches, invokes APIs, writes scripts, and fails, then concentrate their effort on domain-specific retrieval, guardrails, and verification. 1:23:33
The demonstration uses the Poke API and a generated TypeScript library that exposes operations for Pokémon, species, abilities, moves, and related resources. Shihipar contrasts this filesystem-and-code-generation approach with a separate implementation using the regular messages or completion API and individually defined tools. He logs tool calls to inspect execution and uses bun while prototyping because it lets him work with TypeScript without separately managing a TypeScript-to-JavaScript compilation step. 1:25:40
The live example also demonstrates the limitations of an unpolished agent. When asked about generation-two water Pokémon, the model initially appears to rely partly on existing knowledge and does not consistently use the prebuilt API, prompting Shihipar to identify the project instructions as an area for improvement. A later request about building around Venusaur uses a text dataset associated with Smogon; the agent searches references to Venusaur, identifies related Pokémon, teammates, and counters, and generates a script to analyze the material. The lesson is not that the initial prototype is flawless, but that observing its actual behavior reveals where better instructions, preprocessing, or verification are needed. 1:34:25
To move from a successful prototype toward an application, Shihipar describes retaining the useful instructions and helper scripts while adding a relatively small SDK entry point. Deployment can either remain local or run inside a hosted sandbox; a customized user interface can also be served from a development server inside that sandbox and refreshed as the agent edits code. He leaves several questions explicitly open, including cross-agent reuse, per-user container architecture, and the best long-term organization of skills. He also cautions that agents can be expensive and that monetization and expected usage patterns should influence product design from the beginning. 32:54
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:00
[on-hold music] Okay.
- 0:21
Yeah, thanks for joining me. I, uh, I'm still on [REDACTED:location] time, so it feels like I'm doing this at, like, seven AM. [laughs] [laughs] Uh, so yeah, but, um, glad to talk to you about the Claude Agent SDK.
- 0:36
So, um, yeah, I, I think, like, this is gonna be, like, a rough agenda of what we're gonna talk about. We're gonna talk about, like, what is the Claude Agent SDK, why use it.
- 0:46
There are so many other agent frameworks. What is an agent? What is an agent framework? Um, how do you design an agent, uh, using the Agent SDK or, or just in general?
- 0:57
Um, and then I'm going to do some, like, live coding, where Claude is gonna do some live coding on prototyping an agent. Um, and, uh, I've got some starter code, but, uh, yeah, I, I-- The whole goal of this is, like, you know, we got two hours.
- 1:12
We're gonna be super collaborative, ask questions. Um, this is also going to be not, like, a super canned demo in the sense that, like, we're gonna be, like, thinking through things live.
- 1:24
You know, I'm not gonna have all the answers right away. Um, and I think that'll be a good way of, like, building an agent loop, I think is, like, really mu- very much, like, kind of an art or intuition.
- 1:35
So, um, but yeah, before we get started, just curious, a show of hands, like, how many people have heard of the Claude Agent SDK or have... Okay, great. Cool.
- 1:46
And how many have, like, used it or tried it out? Okay, awesome. Okay, so pretty good show of hands. Um, yeah, so I'll, I'll just get started on, like, the, like, you know, overview on agents.
- 1:59
I, I think that, like, this is, I, I, I think something that people have seen before, but I think it's still, it's taking some time to, like, really sink in, uh, how AI features are evolving, you know.
- 2:13
So I think, like, when GPT, you know, three came out, it was really about, like, single LLM feature, right? You're like, "Oh," like, "Hey, can you categorize this?" Like, return a response in one of these categories.
- 2:25
Um, and then we've got more, like, workflow-like things, right? "Hey," like, "Can you, like, take this email and label it?" Or like, "Hey, here's my code base, like, index for your RAG.
- 2:37
Can you give me, like, the next completion or the next, um, the next file to edit?" Right? And so that's what we call, like, a workflow, where you're very, like, structured.
- 2:48
You're like, "Hey," like, "given this code, give me code back out." Right? And now we're getting to agents, right? And, uh, like, the canonical agent to me is Claude Code, right?
- 3:00
Claude Code is a tool where you don't really tell it... We don't restrict what it can do really, right? You're just talking to it in text, and it will take a really wide variety of actions, right?
- 3:12
And so agents, uh, build their own context, like, decide their own trajectories, are working very, very autonomously, right? And so, uh, yeah, and I think, like, as the future goes on, like, agents will get more and more autonomous.
- 3:28
Um, and we, uh, yeah, I think it's like we're kind of at a breaking point where we can start to build these agents. Um, they're not perfect, you know, but it's definitely, like, the right time to get started.
- 3:40
So, um, yeah, Claude Code, I'm sure many of you have, have tried or used. Um, it is, yeah, I think the first true agent, right? Like, the first, uh, time where I saw an AI working for like ten, twenty, thirty minutes, right?
- 3:56
So, um, yeah, it's, it's a coding agent. And, uh, the Claude Agent SDK is actually built on top of Claude Code. And, uh, the reason we did that is because, um, basically we found that when we were building agents at Anthropic, we kept rebuilding the same parts over and over again.
- 4:18
And so to, to give you a sense of, like, what that looks like, of course, they're the models to start, right? Um, and then in the harness, you've got tools, right?
- 4:28
And that's like sort of the first obvious step. Like, let's add some tools to this harness. And later on, we'll give an example of sort of, like, trying to build your own harness from scratch too and, and what that looks like and, and how challenging it can be.
- 4:42
But tools are not just, like, your own custom tools. It might be tools to interact with your file system, like with Claude Code. Um, did the volume just go up or were they not holding it close enough? [laughs]
- 4:53
Okay. Starts with that going to happen. Anyways, um, you got tools, tools you run in a loop, and then you have the prompts, right? Like the core agent prompts, the, um, the, the prompts for the scripts, things like that.
- 5:06
Uh, and then finally you have the file system, right? And, or not finally, but you have the file system. The file system is a way of context engineering that we'll talk more about later, right?
- 5:19
And I, I think, like, I-- one of the key insights we had through Claude Code was thinking a lot more through the, like, context, not just a prompt. It's also the tools, the files, the scripts that it can use.
- 5:30
Um, and then there are skills which we've, like, rolled out recently, and, uh, we can talk more about skills, uh, um, if that's interesting to you guys as well.
- 5:39
Um, and then, yeah, things like, uh, subagents, uh, web search, you know, like, um, like research, compacting, hooks, memory. There are all these, like, other things around a harness as well.
- 5:52
Um, and, uh, it ends up being quite a lot. So the Claude Agent SDK is all of these things packaged up for you to use, right? Um, and yeah, you have your application.
- 6:04
So I, I think, like- Uh, to give you a sense of, uh, yeah. To give you a sense of, like,
- 6:14
maybe why the Claude Agent SDK is, um
- 6:21
... Yeah, like, like, so yeah, people are already building agents on the, uh, SDK. A lot of software agents, uh, you know, software reliability, security, incident triaging, bug finding, um, site and dashboard builders if you're ...
- 6:35
These are extremely popular. If you're using it, you should absolutely use the SDK. Um, and its office agents, if you're doing any sort of office work, tons of examples there.
- 6:45
Um, got some like, you know, legal, finance, healthcare ones. Um, so yeah, there are tons of people building on top of it. Um, I want to ... But yeah, okay.
- 6:55
So why the Claude Agent SDK, right? Like, why did we do it this way? It's why did we build it on top of Claude Code? And we realized basically that as soon as we put Claude Code out, yeah, the engineers started using it, but then the finance people started using it, and the data science people started using
- 7:12
it, and the marketing people started using it, and yeah, I think it just like, it ... We just realized that people were using Cla- Claude Code for non-coding tasks.
- 7:21
And we felt ... And, and as we were building, you know, non-coding agents, we kept coming back to it, right? And so, um, it's a, like ... And we'll go more into why that just works, why we, we can, we could use Claude Code for non-coding task.
- 7:38
Uh, spoiler alert, it's like the Bash tool. Um, but yeah, it's, uh, it, it, it was something that we saw as an emergent pattern that we want to use, and we've built our agents on top of it, right?
- 7:50
And, uh, these are lessons that we've learned from deploying Claude Code that we've sort of baked in. So, uh, tool use errors or compacting or things like that, stuff that is, like, very, can take a lot of scale to find, you know, like, what are the best practices we've sort of baked into the Claude Agent SDK.
- 8:08
Um, as a result, we have a lot of strong opinions on the best way to build agents. Uh, like, I think the Claude Agent SDK is quite opinionated. We'll, I'll talk over some of these opinions and, and why, like, uh, why we chose them, right?
- 8:21
Um, but yeah, one of the big opinions of the Bash tool is the most powerful agent tool. So, okay, um, what, what are, like, what I would describe as the Anthropic way to build agents, right?
- 8:32
And I'm not, I'm not saying that you can only build agents using the API this way, right? But this is, like, um, if you're using our opinionated stack on the Agent SDK, what is it, right?
- 8:42
So roughly Unix primitives, like the Bash and file system, and, you know, we're gonna go over, like, prototyping an agent using Claude Code. And, uh, my goal is really to sort of show you what that looks like in real time, right?
- 8:56
Like, why is Bash useful? Why is the file system useful? Why not just use tools? Um, yeah, agents, uh, I mean, you can also make workflows, and we'll talk about that a bit later.
- 9:07
The agents build their own context. Um, thinking about code generation for non-coding, um, like, we use codegen to generate docs, query the web, like, do data analysis, take, uh, unstructured action.
- 9:20
So, um, there's a lot of, like, uh ... This can be s- pretty counterintuitive to some people, and again, in the, like, prototyping session, we'll, we'll go over how to use code generation for non-coding agents.
- 9:32
Um, and yeah, every agent has a container or is hosted locally because this is Claude Code. Uh, it needs a file system, it needs Bash, it needs to be able to operate on it, and so it's a very, very different architecture.
- 9:45
I'm not planning to talk too much about the architecture today, but we can at the end if that's what people are interested in, in ... Or sorry, by architecture, I mean hosting architecture.
- 9:54
Like, how do you host a agent and, like, uh, what are best practices there? Happy to talk about that at the end. Um, yeah. So
- 10:03
well, let me pause there 'cause I feel like I covered a lot already. Any questions so far on the Agent SDK, agents? Um, yeah, like, what you get from it?
- 10:14
Yeah.
- 10:15
Can you s- can you explain what code generation for non-coding means exactly, Alan?
- 10:19
Yeah. Um, this is, um, like, basically when you ask Claude Code to do a task, right? Like, let's say that you ask it to, uh, find the weather in San Francisco and, like, you know, tell me what I should wear or something, right?
- 10:37
Like, uh, what it might do is it might start writing a script, uh, to fetch a weather API, right? And then start, like, maybe it wants it to be reusable.
- 10:50
Like, maybe you want to do this pretty often, right? So it might fetch the weather API and then get the, like, maybe even get your location dynamically, right? Based on your IP address.
- 11:00
And then it will, like, um, you know, check the weather and then maybe, like, call out to, like, a subagent to give you recommendations. Maybe there's an API for your closet or wardrobe, right?
- 11:13
It's like, so that's an example. I, I think that, like, it's kind of, um, for any single example, we can talk over how you might use code, uh, codegen.
- 11:23
Uh, a lot of it is, like, composing APIs is, like, the high level way to think about it. Yeah.
- 11:29
Uh, yeah. And-
- 11:30
Yeah. Uh, workflow versus agent, uh, like, for repetitive task or, you know, like, a process, a business process that is always the same, do you will still prefer to build an agent versus a fully deterministic workflow?
- 11:43
Yeah. So we do have-
- 11:45
Can you repeat the question, sorry?
- 11:47
Oh, sure. Yeah, yeah. Um, so the quest- the question was about workflows versus agents, and would you still use the Claude Agent SDK for workflows. Is that right? Um, y- yes.
- 11:58
And so, and so, uh, I mean, we ... I just- We just sort of tell you what we do internally [chuckles] basically. And what we do internally is we've done a lot of, like, GitHub automations and Slack automations built on the Claude Agent SDK.
- 12:11
So, uh, you know, we have a bot that triages issues when it comes in. That's a pretty workflow-like thing. But we've still found that, you know, in order to triage issues, we want it to be able to clone the code base, and sometimes spin up a Docker container and test it, and things like that.
- 12:24
And so it's still ends up being like a very... Like, there's a lot of steps in the middle that need to be quite free-flowing, um, and then you, like, give structured output at the end.
- 12:35
So, um, yes. All right, I'll take one more question and then keep going. So yeah, in the blue.
- 12:40
Yeah. Uh, so could you talk about security and guardrails? Like if, if, you know, we're using Claude Agent SDK and, you know, you're leaning towards using Bash as the, you know, all-powerful generic tool-
- 12:51
Yeah
- 12:51
... is the onus on, uh, building the, the agent builder to make sure that, you know, you're preventing against, like, common attack vectors? Or is that something that the model is, is, is doing, um, by itself?
- 13:03
Yeah. So I, I think this is sort of like the Swiss cheese... Oh, yeah. Sorry. So the question was, uh, permissions on the Bash tool, right? Or like how do you think about permissions and guardrails?
- 13:14
The, like, in, like, when you're giving the agent this much power over, you know, your, its environment and the computer, how do you make sure it's aligned, right? And so the way we think about this is, uh, what we call, like, the Swiss cheese defense, right?
- 13:26
So like there is, um, like, on every layer some defenses, and together we hope that it, like, blocks everything, right? So obviously on the model layer, uh, we do a lot of, um, alignment there.
- 13:40
We actually just put out a really good paper on reward hacking. Super recommend you check that out. Um, so like, definitely I think Claude models, like, we try and make them very, very aligned, right?
- 13:51
And, uh, so yeah, there's the model alignment behavior. Then there is, like, the harness itself, right? And so we have a lot of, like, permissioning and prompting. Um, and, uh, like, we do a AST-based parser on the Bash tool, for example.
- 14:07
So we know, um, fairly reliably, like, what the Bash tool is actually doing, and definitely not something you'd want to build yourself. Um, and then finally, the last layer is sandboxing, right?
- 14:19
So like let's say that... And someone has maliciously taken over your agent, what can it actually do? Uh, we've included a sandbox in, like, where you can sandbox network request, um, and sandbox, uh, file sy-system operations outside of the file system.
- 14:35
And so, uh, yeah, ultimately that's what they call, like, the lethal trifecta, right? Is like, um, like the ability to, like, execute code in an environment, change a file system, um, exfiltrate the code, right?
- 14:48
I think I'm getting the lethal trifecta a little bit wrong there [chuckles] but, like, the idea is basically, like, if they can exfiltrate your, like, information back out, right, um, that's like...
- 14:58
Th-they still need to be able to extract information. And so if you sandbox the network, that's a good way of doing it. Um, if you're hosting on a sandbox container like Cloudflare, uh, Modal or, you know, E2B, Daytona, like all of these is, like, sand-sandbox providers, they've also done like some level, level of security there, right?
- 15:14
It's like you're not hosting it on your personal computer, um, or on a computer with like your prod secrets or something. So, uh, yeah, lots of different layers there.
- 15:22
And, and yeah, we can talk more about hosting in depth. Um, so okay. So I'm gonna, uh, talk a little bit about Bash is all you need, you know.
- 15:32
Um, I think this is something that, oh, yeah, um, this is like my shtick, you know. I'm, I'm just gonna like keep talking about this until everyone like, uh, agrees with me.
- 15:44
Um, or like, I, I think this is something that we found at Anthropic. I think it is sort of a, something I discovered once I got here. Um, Bash is what makes Claude Code so good, right?
- 15:53
So I think, like, you guys have probably seen, like, code mode or programmatic tool use, right? Like the, um, different ways of, like, composing MCPs. Uh, Cloudflare's put out some blog posts on that.
- 16:06
We've put out some blog posts. Uh, the way I think about code mode is like, or Bash, is that it was like the first code mode, right? So the Bash tool allows you to, you know, like, store the results of your tool calls to files, uh, store memory, dynamically generate scripts and call them, compose functionality like tail
- 16:24
or grep. Uh, it lets you use existing s-software like FFmpeg or LibreOffice, right? So there's a lot of, like, interesting things and powerful things that the Bash tool can do.
- 16:35
And like, think about, like, again, what made Claude Code so good. If you were designing an agent harness, maybe what you would do is you'd have a search tool and a lint tool and an execute tool, right?
- 16:46
And like, you know, N tools, right? Like every time you thought of like a new use case you're like, "I need to have another tool now," right? Um, instead now Claude just uses grep, right?
- 16:55
Or it, it knows your package manager, so it runs like npm run, like test.ts or index.ts or whatever, right? Like it can lint, right? And it can find out how you lint, right, and can run npm run lint.
- 17:08
If, if you don't have a linter it can be like, "What if I install ESLint for you?" Right? So, um, this is like, you know, like I said, the first programmatic tool calling, first code mode, right?
- 17:19
Like you can do a lot of different actions very, very generically, right? Um, and so to talk about this a little bit in the context of non-coding agents, right?
- 17:31
So let's say that we have an email agent, and the user is like, "Okay, how much did I spend on ride sharing this week?" Um, a, you know, like it's got one tool call or generally it's got the ability to search your inbox, right?
- 17:47
And so it can run a query like, "Hey, search Uber or Lyft," right? And without Bash, it, it searches Uber or Lyft, it gets like 100 emails or something, and now it's just got to like think about it.
- 18:02
You know what I mean? And I, I think, like, a good, like, analogy is sort of like imagine if someone came to you with like- Like a stack of papers and like, "Hey, how much did I spend on ride-sharing this week?
- 18:12
Can you, like, read through my emails?" You know what I mean? Like, that, that would be really hard, right? Like, uh, you need very, very good precision and recall to do it.
- 18:20
Um, or with Bash, right? Like, let's say there's a Gmail search script, right? It takes in a query function. Um, and then you can start to save that query function to a file or pipe it.
- 18:34
You can grep for prices. You know, you can, uh, then add them together. You can check your work, too, right? Like you can say, "Okay, let me grep all my prices, store those as, like, in a file with line numbers, and then let me then be able to check afterwards, like, uh, was this actually a price?
- 18:51
Like, what does each one correlate to?" Right? So there's a lot more, like, dynamic information you can do to check your work with the Bash tool. So this is like, um, just a simple example, but, like, hopefully showing you sort of the power of, like, the composability of Bash, right?
- 19:08
So, uh, I'll pause there. Any questions on Bash is all you need, the Bash tool? Any, anything I can make a little bit clearer? Yeah.
- 19:16
Do you have stats on how many people use YOLO mode versus if it's like
- 19:21
Uh, stats on YOLO mode? We probably do. Um, I mean, internally we, we don't, uh, but that's just... I think we just have a higher security posture. Um,
- 19:32
yeah, [laughs] I'm not sure. Uh, I can probably pull that. Any other questions on Bash?
- 19:38
Okay, cool. Um, yeah, just to give you, like, some more examples. Like, let's say that you had an email API and you wanted to, uh, you know, like, go through, like, fetch my e-- like, tell me who emailed me this week, right?
- 19:54
So you've got two APIs. You've got an inbox API and a contact API. Um, this is, like, a way you can do it via Bash. You can also do it via CodeGen.
- 20:01
This is kind of, like, enough Bash that it's, it is CodeGen, right? Like, um, Bash is
- 20:07
ostensibly a CodeGen tool. Um, and then, yeah, like, let's say that you wanted to-- you had a video meeting agent, right? You want it to say, like, "Find all the moments where the speaker says quarterly results in this earnings call," right?
- 20:19
You can use FFmpeg to, like, slice up this video, right? Um, you can use jq to, like, uh, start analyzing the information afterward. So, um, yeah, lots of, like, def-- like, powerful ways to use, uh, to use Bash.
- 20:34
So, all right. I'm gonna talk a little bit about workflows and agents. They can do both. You could use, uh, build workflows and agents on the Agent SDK. Um, yeah, agents are like Claude Code.
- 20:45
So if, if you are, like, building something where you wanna talk to it in natural language and take action flexibly, right? Then that's where you're building an agent, right?
- 20:55
Like you want... You have an agent that talks to your, like, business data, and you want to get insights or dashboards or answer questions or, uh, write code or something, like, that's an agent, right?
- 21:05
And then a workflow is kind of like, you know, we do a lot of GitHub Actions, for example, right? So you define the inputs and outputs very closely, right?
- 21:12
So you're like, "Okay, can you get a PR and give me a code review?" Um, and yeah, both of these you can use Agent SDK for. Um, when building workflows, you can use structured outputs.
- 21:22
Uh, we just released this. Um, you can, yeah, Google [laughs] Agent SDK structured outputs. Um, but yeah, so you can do both. I'm going to primarily be talking about agents right now.
- 21:34
A lot of the things that you can, like, learn from this are applicable to workflows as well. So, um, yeah, we'll, we'll talk about this. Uh, wait. Uh, show of hands, how many people have, like, designed an agent loop before?
- 21:50
Okay, cool. Okay, great, great. Um, so yeah, I mean, I think the number one thing, the, the meta learning for designing an agent loop to me is just to read the transcripts over and over again.
- 22:03
Like, every time you see, see the agent run, just read it and figure out, like, "Hey, what is it doing? Why is it doing this? Can I, uh, help it out somehow?"
- 22:11
Right? Um, and, uh, we'll do some of that later, right? So we'll, uh, we'll build an agent loop. Um, but here is the, uh, the three parts to an agent loop, right?
- 22:25
So, uh, first, it's gather context, right? Second is taking action, and the third is verifying the work, right? And, uh, this is, like, not the only way to build an agent, but I think a pretty good way to think about it.
- 22:43
Um, gathering context is, uh, like, you know, for Claude Code, it's grepping and finding the files needed, right? Um, you know, for an email agent, it's like finding the relevant emails, right?
- 22:55
Um, and so these are all, like, pretty, um... Yeah, like I, I think thinking about how it finds its context is very important, and I think a lot of people sort of, uh, skip the step or, like, underthink it.
- 23:09
This can be, like, very, very important. Uh, and then taking action, um, how does it, like, do its work? Uh, does it have the right tools to do it, like code generation, uh, Bash?
- 23:19
These are more flexible ways of taking action, right? And then verification is another really important step. And so, uh, the-- basically, what I'd say right now is, like, if you're thinking of buil-building an agent, think about, like, can you verify its work, right?
- 23:35
And if you can verify its work, it's, like, a great, like, candidate for an agent. If you can't verify its work, like it's, like, you know, coding, you can verify by linting, right?
- 23:44
And you can at least make sure it compiles. So that's great. Uh, if you're doing, let's say, deep research, for example, it's actually a lot harder to verify your work.
- 23:52
One way you can do it is by citing sources, right? So that's, like, a step in verification. But obviously, research is less verifiable than code in some ways, right?
- 24:00
Because, like, code has a compile step, right? You can also, like, execute it and see what it does, right? So, um, I think, like, thinking on, you know, like, as we build agents, the ones that are closest to being very general are the ones with the verification step that is very strong, right?
- 24:16
So, um I, I think there was a question here. Yeah.
- 24:18
So, when-- where do you generate a plan of the work you need to run through?
- 24:25
Hmm. Yeah, I mean, you, you might-
- 24:28
Can you repeat the question?
- 24:29
Oh, yeah, sorry. The, the question was, when do you generate a plan, um, before you run through it? So, um,
- 24:37
like in Claude Code, you don't always generate a plan. Um, but if you want to, you'd insert it between the gathering context and taking action step, right? And so, um, plans sort of help the agent think through step by step, but they add some latency, right?
- 24:52
And so there is like some trade-off there. Um, but yeah, the agent SDK helps you, like, do some planning as well. So yeah.
- 24:59
Yeah.
- 25:00
Can you, like, make the agent create that to-do list for like one hundred percent sure that it will create that to-do list and run by it?
- 25:12
Uh, yeah. So the question was, will the agent create the to-do list? Uh, yes. Um, if you're using the agent SDK, we have like some to-do tools that come with it, and so it will, like, maintain and check off to-dos, and you can display that as you go.
- 25:26
So yeah. Um, any other questions about this right now? Okay, cool. Okay, so I'm gonna quickly talk about, like, like how do you do this stuff? You-- Like, what are your tools for doing it, right?
- 25:41
And, uh, there are three things you can do. There, you have tools, Bash, and code generation, right? And I, I think traditionally, I think a lot of people are only thinking about tools.
- 25:52
And, uh, yeah, basically one of the call to actions is just figuring out, like, thinking about it more broadly, right? So tools are extremely structured and very, very reliable, right?
- 26:01
Like, if you want to sort of have as fast an output as possible with minimal errors, uh, minimal retries, uh, tools are great. Uh, cons, they're high context usage.
- 26:12
If anyone's built an agent with like fifty or a hundred tools, right? Like, they take up a lot of context, and the model, it kind of gets a little bit confused, right?
- 26:21
Um, there's no, like, sort of discoverability of the tools, um, and they're not composable, right? And l- l- and I say tools in the sense of, like, if you're using, you know, uh, messages or completion API right now, um, that's how the tools work.
- 26:35
Of course, like, you know, there's like code mode and programmatic tool calling, so you can sort of blend some of these. Um, but then there's Bash. So Bash is very composable, right?
- 26:45
Like, uh, static scripts, low context usage. Uh, it can take a little bit more discovery time. Like, 'cause like let's say that you have, whatever, you have like the Playwright MCP or something like that, um, or sorry, the Playwright CLI, the Playwright, like, Bash tool.
- 27:00
Um, you can do playwright --help to figure out all the things you can do, but the agent needs to do that every time, right? So it needs to, like, discover what it can do, um, which is kind of powerful.
- 27:10
That helps take away some of the high context usage, but adds some latency. Um, there might be slightly lower call rates, you know, just because, like, it has a little bit more time to, um, it, right, it needs to, like, find the tools and, and what it can do.
- 27:26
Um, but this will definitely, like, improve as it goes. And then finally, codegen, highly composable, dynamic scripts. Um, they take the longest to execute, right? So they need linting, possibly compilation.
- 27:39
API design becomes like a very, very interesting step here, right? And I, and I'll talk more about, like, uh, best-- like how to think about API design in an agent.
- 27:49
Um, but yeah, I, I think this is like how you-- like the, the three tools you have. And so yeah, using tools, think-- You still want some tools, but you wanna think about them as atomic actions your agent usually needs to execute in sequence, and you need a lot of control over, right?
- 28:05
So for example, in Claude Code, we don't use Bash to write a file. We have a write file tool, right? Because we want the user to be able to sort of see the output and approve it and, um, we're not really composing write file with other things, right?
- 28:19
It's like a very atomic action. Um, sending an email is another example. Like, any sort of like non-destruct-- like destructible or sort of like, you know, uh, un-reversible change is definitely, like, a, a tool is a good place for that.
- 28:33
Um, then we've got Bash. Uh, so for example, there are like, uh, composable actions, like searching a folder, using GitHub, linting code, and checking for errors or memory. Um, and so yeah, you can write files to memory, and that can be your Bash-- Like Bash can be your memory system, for example, right?
- 28:52
So, um, and then finally, you've got code generation, right? So if you're trying to do this like highly dynamic, very flexible logic, composing APIs, uh, like if you're doing data analysis or deep research or like reusing patterns.
- 29:05
And so, um, yeah, we'll talk more about, uh, code generation in a bit.
- 29:11
Um, any questions so far about like the SDK loop or tools versus Bash versus codegen? Yeah.
- 29:18
Yeah. Uh, I was gonna ask, um, do you have a-- Are you gonna have any ready-made tools for like offloading tool call results?
- 29:27
Offloading tool call results, like into the file system or-
- 29:29
Like if that sequence goes to Bash and then the context explodes.
- 29:33
Mm.
- 29:33
Is it like types of commands that like blew everything up?
- 29:36
Okay.
- 29:37
Or, or otherwise just like long outputs polluting your history?
- 29:40
Sure. Yeah, yeah, yeah.
- 29:41
I'm actually like all the time just offloading them to files.
- 29:45
Yeah, yeah. I, I think that's a good common practice. I think, um, we--
- 29:52
I, I remember seeing some PRs about this very recently on, on, on Claude Code about handling very long outputs, and I,
- 30:02
I, I don't know exactly. Like, I, I think, I think we are moving towards a place where more and more things are being like just stored in the file system, and this is like a good example.
- 30:12
Yeah, like it's storing like long outputs, uh, over time. Um, I think like generally prompting the agent to do this is a good, uh, way to think about it.
- 30:21
Or even if you have-- I think like something I just do always now is like whenever I have a tool call, I, um- I save it, like the results of the tool call to the file system so that you can, like, sort of trust it, and then have the tool call return the path of the result.
- 30:37
Um, just because, like, that helps it, like, sort of recheck, uh, its work. So, um,
- 30:44
yes?
- 30:45
Um, do you find that you need to use, like, the skills, um, kind of structure to help Claude along to use the Bash better or out of the box, you know, that's not necessary?
- 30:59
Yeah. So the question was about skills and, like, do we need skills to use Bash better? Um, yeah, for context, skills ... Maybe I can ...
- 31:10
Skills. Okay, yeah. Skills are basically a way of like, uh, you know, allowing our agent to take longer, complex tasks and, like, sort of load in things via context, right?
- 31:22
So so- like, for example, we have, uh, a bunch of docx skills, and these docx skills tell it how to do code generation to generate these files, right? And so, um, yeah, I, I think overall, skills are, yeah, basically just a collection of files.
- 31:37
They're also sort of like an example of being very, like, file system or Bash tool-pilled, right? Um, because they're really just folders that your agent can, like, cd into and, like, read, right?
- 31:51
Um, and so, yeah, they give ... Like, what we found the skills are really good for is pretty, like, repeatable instructions that need a lot of expertise in them.
- 32:02
Uh, like for example, we released our front-end design skill recently that I really, really like, and, um, it's really just sort of a very detailed and good prompt on how to do front-end design.
- 32:13
Uh, but it comes from, like, our best, you know, like, um, AI front-end engineer. You know what I mean? And he, like, really put a lot of top thought and iteration to it.
- 32:22
So that's one way of using skills. Um, yeah.
- 32:27
Quick question.
- 32:28
Yes.
- 32:28
So I use that front-end skill.
- 32:30
Sure.
- 32:30
And it's pretty cool. Thanks for, uh, publishing it. Uh, I want to understand, uh, there are multiple .md files. Like, CLAUDE.md is also there, and it is also at the user level and supported level, and then there are SKILL.md files.
- 32:45
Like, is there, like, a priority order? Should some stuff be relegated to the CLAUDE.md and some other stuff should only come through SKILL.md?
- 32:53
Hmm. So the question was about SKILL.md versus CLAUDE.md and how to think about, uh, that, right? And, uh, I think like ... I- I will say all of these concepts are so new.
- 33:05
You know what I mean? Like, even Claude Code is, like, released at, like, eight or nine months ago, right? Like, um, and so skills were released like two weeks ago.
- 33:13
Like, I, like, I won't pretend to know all of the best practices for, for everything, right? Um, I think generally
- 33:21
skills are a form of progressive context disclosure, and that's sort of a pattern that we've talked about a bunch, right? Like, with like, uh, Bash, and, you know, like preferring that over like, you know, purely s- like normal tool calls.
- 33:34
It's like, it's a way of like the agent being like, "Okay, I need to do this. Let me find out how to do this, and then let me read in the SKILL.md," right?
- 33:43
So you ask it to make a docx file, and then it, like, cds into the directory, reads how to do it, writes some scripts, and keeps going. So, um, yeah, I, I think, like, there's still some intuition to build around, like, what, what exactly you, like, define as a skill and how you split it out.
- 33:59
Um, but, uh, yeah, I think, uh, yeah, lots of best practices to learn there still. Um, yeah.
- 34:08
Uh, so yesterday, uh, you talked about the future of skills-
- 34:13
Yeah
- 34:13
... and how it evolves over time. Do you see this as ultimately becoming part of the model and you need less of the skills? This is just a way to bridge the gap for now?
- 34:22
Yeah. So the question was, are skills ultimately part of the model? Um, are they a way to bridge the gap? I missed Barry's talk, uh, Barry Mesch's talk yesterday, but, uh, yeah, I think roughly the idea is that the model will get better and better at doing a wide variety of tasks, and skills are the best way
- 34:39
to give it out-of-distribution tasks, right? Um, but I, I would broadly say that, like, it's really, really hard, especially, like, you know, if you're, like, m- uh, not at a lab to, like, tell where the models are going exactly.
- 34:56
Um, my general rule of thumb is, like, I try and, like, rethink or rewrite my, like, agent code, like, every six months, uh, just 'cause I'm like, uh, things have probably changed enough that I've, like, baked in some assumptions here.
- 35:08
And so, like, I think the ... Like, our Agent SDK is built to, as much as possible, sort of advance with capabilities, right? Like, the Bash tool will get better and better.
- 35:18
Uh, we're building it on top of Claude Code, so as Claude Code evolves, you'll get those wins out, out of the gate. Um, but at the same time, like, you know, things are so different now, like, than they were a year ago in, in terms of, like, AI engineering, right?
- 35:34
And I think, like, a general best practice to me is sort of like, "Hey, we can write code 10 times faster. We should throw out code 10 times faster as well."
- 35:43
Um, and I think thinking about, like, not so, like, hedging your bets on, like, where is the future right now, but, like, what can we do today that really works, right?
- 35:53
And like, like, let's get market share today and not be afraid to throw out code later. Um, if you're a startup, this is arguably your largest advantage that you have over competitors.
- 36:04
They're like, you know, larger companies have, like, six-month incubation cycles, and so they're always, like, stuck in the past of, like, the agent capabilities, right? And so your advantage is that you can, like, be like, "Hey, the agent-- the capabilities are here right now.
- 36:18
Let me build something that uses this right now." Right? So, um, yeah. Uh,
- 36:25
any, any other questions on ... For-- We're talking about skills in Bash. Okay. It seems like there are a lot of skill questions. So, um- Yeah. Uh, I, I think at the back, someone.
- 36:38
You might have to shout
- 36:40
Yeah. So why would you use a skill versus an API? They look very similar to, like you could, that Python program there could be a package, right?
- 36:48
Yeah. So the question was why use a skill versus an API. Um, good question. I, I think that like, um, when you ... Like these are all forms of progressive disclosure basically to the agent to figure out what it needs to do.
- 37:02
Um, and I'll go over like, uh, examples of like you just have an API, right, in, in our like in, in our prototyping session. Um, it's totally like use case dependent, right?
- 37:14
Like just, I, I think like I don't have a ... Like I don't think there's a general rule. I think it's like read the transcript and see what your agent wants.
- 37:21
If your agent always wants, like thinks about the API better as like a API.ts file or something, or API.py file, do that. You know, that's great. Like I think skills are like a, like sort of an introduction into like thinking about the file system as a way of storing context, right?
- 37:37
And they're a great abstraction. Um, but there are many ways to use the file system. Um,
- 37:44
and I, I should say that like something about skills is that like you need the Bash tool, you need a virtual file system, things like that. So the Agent SDK is like basically the only way to really use skills to like their full extent right now.
- 37:55
So, um, yeah. Yeah, back there.
- 37:59
Can we expect a marketplace for skills from Anthropic?
- 38:02
Yeah, the question was can we expect a marketplace for skills? So, um, yeah, Claude Code has a plugin marketplace that you can also use with the Agent SDK. Um, we're evolving that over time.
- 38:13
You know, like it was like a very much a v0. Um, and by marketplace, I'm not sure if people will be charging for this exactly. It's more just like a discovery system, I think.
- 38:23
Um, but yeah, that exists right now. You can do /plugins in Claude Code. Um, and, and you can find some, so yeah. Yeah.
- 38:31
What's your current thinking about when you're gonna reach for like the SDK, you know, to solve a problem?
- 38:36
When ... Yeah, so the question is when do I use the SDK to solve a problem? Uh, if I'm building an agent, basically. I, I think that like, um, my overall belief is that
- 38:50
like for any agent, the Bash tool gives you so much power and flexibility, and using the file system gives you so much power and flexibility that you can always eke out performance gains over it, right?
- 39:01
And so, um, yeah, in the prototyping part of this talk, we're going to like look at an example with only tools and an example without, with, you know, Bash and the file system, and compare those two.
- 39:13
Um, and yeah, that's what I mean by be- being Bash-pilled. I'm like I, I just like start from the Agent SDK, you know? And I think a lot of people at Anthropic have started, like doing that as well.
- 39:23
So, um, of course, I, I do wanna say that there are lots of times where the Agent SDK is kind of annoying 'cause you've got like this network sandbox container, and you're like, "I hate ...
- 39:32
Like I don't wanna do this," you know what I mean? Like, "I want to run on my browser locally," right? Um, I totally get that. I think it's, it, there is like a real performance trade-off.
- 39:41
Um, the way I think about it is sort of like React versus like jQuery. You know, like I, like I ... When I was coming up, I was like very into web dev, and like, you know, I was using jQuery and Backbone, and then React came out, and it was by Facebook.
- 39:55
And they're like, "You have to ... W- Here's JSX. Like we just made this up, and, and now there's a bundler," right? And I'm like, "Ah, it's so annoying."
- 40:01
Um, but like it generally makes the model, or it makes, it made web apps more powerful, right? And I think we're sort of like ... The Agent SDKs are like the React of agent frameworks to me because it's like we build our own stuff on top of it, so you know it's real.
- 40:18
And all the annoying parts of it are just like things where we're annoyed about it too, but we're like, it just, it just works. Like you have, like you gotta do this, you know?
- 40:26
Um, so yeah. Uh, yeah, okay. More, more skill questions, I guess? Yeah, right here.
- 40:33
Uh, one of the s- quick Bash question.
- 40:34
Oh, sure. Bash question. Great. I love Bash.
- 40:36
If you got custom internal like Bash tools-
- 40:38
Yeah
- 40:39
... how do you let the agent discover that, or do those have to become tool to tools?
- 40:44
Okay, the question is if you have custom agent Bash tools, how do you let the agent discover that? By custom Bash tools do you mean like Bash scripts or-
- 40:51
Like we build internal things. We have, we have Bash scripts, yeah-
- 40:52
Yeah
- 40:53
... that we built.
- 40:54
Um, yeah. So I, I think, uh, where is it? You just put it in the file system, and you tell it like, "Hey," like, "here is a script." Uh, you can call it ...
- 41:03
You know, I- I'm generally thinking in the context of the Claude Agent SDK where it has the file system and the Bash tools are tied together. This is kind of an anti-pattern I see sometimes where people are like, "Oh," like, "we're gonna host the Bash tool in this like virtualized place, and it's not going to interact with
- 41:20
other parts of like the agent loop," you know? And that sort of, you know, makes it hard 'cause if, if you've got a tool result that's saving a file, then your Bash tool can't like, uh, read it, you know what I mean?
- 41:32
Unless it's all in one, one container. So do, does that answer your question? Like-
- 41:37
Um, yeah, kind of. I mean, like, so you're saying you would just put it in like system prompt or something?
- 41:41
Yeah, just put it in system prompt and be like, "Hey, you have access to this." Uh, I would like sort of design all my CLI scripts to have like a --help or something so that the model can call that, and then it can like progressively disclose like every, like subcommand inside of the script.
- 41:55
Yeah. Uh, yeah, back there.
- 41:57
Yeah. So, uh, like my question is around when to reach for the Agent SDK.
- 42:01
Yeah.
- 42:02
So have you designed, or rather would you recommend someone use the Agent SDK to build like a generic chat agent as compared to like, oh, you know, I'm building an agent where you have some input, and the agent goes and does some stuff, and finally I care about the output, as compared to let's say someone ...
- 42:19
Like are you using or do you foresee using the agent to build, like the Agent SDK to build like Claude the, the app rather than Claude Code?
- 42:28
Uh, yeah. So the question is when do we reach for the Agent SDK? Uh, does, um- Like, uh, like would we use the Agent SDK to build Claude.ai, which is a more traditional chatbot, uh, than Claude Code?
- 42:45
Um, I, uh, one, I think Claude Code is like a very... Like, like the interface is not a traditional chatbot interface, but like the inputs and outputs are, are, right?
- 42:54
Like you input code in, you, you get like... Or you input text in, you get c- text out, and you can take actions along the way. Um, you might have seen that like when we rolled out doc creation for Claude.ai, um, now it has the ability to spin up a file system and like create spreadsheets and PowerPoint
- 43:15
files and things like that by generating code. And so that is like, you know, we're in the midst of sort of like, um, like merging our agent loops and stuff like that.
- 43:24
But, but broadly, like, uh, like yeah, Claude.ai will... Like, it's getting more and more. Like you see it with skills and the memory tool and stuff, more and more file-system-pilled, right?
- 43:34
So, uh, we do think it's like a broad thing that you can use just, just generally, and happy to talk through examples and, um... Yeah, one more question, and then we'll, we'll keep going.
- 43:44
Yeah.
- 43:44
Um, still trying to understand the rule of thumb on when to build a tool or use a tool, when to wrap something with a script or just let the agent go wild on the bash.
- 43:56
Because I'll, I'll give you an example. Let's say I need to access a database
- 44:02
from time to time. I can use an MCP, I can wrap it in a script, and I can just let the agent call an endpoint from ba- directly from bash, right?
- 44:13
Yeah. Great question. Great question. So it- still trying to grok like when to use tools versus bash versus codegen, and he gave an example like, okay, I have a database.
- 44:22
Um, I want the agent to be able to access it in some way. What should I do? Should I create a tool that queries the database in some way?
- 44:29
Um, should I use the bash? Should I use codegen, right? These are all-- These are three ways of doing it. Um, I think that they are... Like, you could use any of them, and I, I think like part of it is like...
- 44:40
I, I think unfortunately there's no like single best practice, right? This is like kind of a system design problem. But let's say that you want to access your bash, your database via a tool, you would do that if your database was very, very structured, and you had to be very careful about like, I don't know, you're accessing
- 44:57
like user sensitive information or something like that, and you're like, "Hey, I, I can only take in this input and I need to like give this output, and I c- have to mask everything else about the database from the agent."
- 45:11
Right? Obviously, that like sort of limits what the agent can do, right? Like it can't write a very dynamic query, right? Um, if you're writing a full-on SQL query, I would definitely use Bash or codegen, uh, just because when the model is writing a SQL query, it can make mistakes, and the way it fixes it mi- its
- 45:30
mi- is its mistakes is by like linting or like running the file, looking at the output, seeing if there are errors, and then iterating on it, right? Um, and so I generally like, if I'm building an agent today, I'm giving it as much access to my database as possible, and then I'm like putting in guardrails, right?
- 45:50
Like I'm probably limiting its like write access in different ways. But what I-- probably what I would do is like I would give it write access and put in specific rules, and then give it feedback if it tries to do something it can't do.
- 46:06
You know what I mean? And, and so I know this is like kind of a hard problem, but I think this is the like set of problems for us to solve, right?
- 46:13
Like, we've built a Bash tool parser, um, s- b- and that's a super annoying problem. Uh, but we need to solve that in order to like let the agent work generally, right?
- 46:23
And same thing with like database. Like, like yes, it's quite hard to understand what is a query doing, but if you can solve that, you can let your agent work more generally over time.
- 46:32
So, um, yeah. I, I think thinking about it, uh, like flexibly as much as possible and keeping tools to be like very, very like sort of atomic actions, right?
- 46:43
That you need a lot of guarantees around. Um-
- 46:46
A follow-up-
- 46:47
Yeah, of course
- 46:47
... on the same thing.
- 46:48
Yeah.
- 46:48
Like, uh, how do you ensure that role-based access controls are taken care of effectively?
- 46:55
How do you ens-, uh, ens... So the question is how do you ensure that the role-based acces- uh, access controls are taken care of? Usually, that's in like how you provision your API key or your back-end service or something like that, right?
- 47:05
Like, um, I think that like probably what I do is like I create like temporary API keys. Sometimes people create proxies in between to insert the API keys, um, if you're concerned about exfiltration of that.
- 47:18
Um, but yeah, I would create s- like API keys for your agents that are scoped in certain ways, and so then on the back end, you can sort of check it's like, you know, what it's trying to do and like, uh, if it's a, an agent, you can like give it different feedback.
- 47:31
So, yeah. All right, you have one question.
- 47:34
Um, anything you could tell us, uh, more about the, the memory tool, the internal memory tool?
- 47:41
Um, I have-- I, I'm not trying to like keep a secret. I, I don't know exactly, like I haven't read the code. But I, I think it generally works on, on the file system.
- 47:51
And so, uh-
- 47:52
Is it exposed to, uh, to the, uh, Agent SDK or has it already been built?
- 47:57
Um, I would say that like... We, we've had this question a bunch. I would just use the file system on, in the Claude Agent SDK. I would just create like a memories folder or something and tell it to write memories there.
- 48:07
Um, it's like I, I don't know the exact implementation of the memory tool, but it does use the file system in, in, in that way. So yeah.
- 48:16
Um, all right. Last, yeah, last question on this. Yeah.
- 48:19
How you are manage-- For the Bash and the code, how you are managing the, like, like, uh, reusability? Suppose the same agent is allowed to hundreds of users and, uh, same code every time it is generating and every time it is executing.
- 48:33
So how can we use the reusability?
- 48:36
Yeah, that's a really good question. So, uh- Yeah. Let's say you have two agents interacting with two different people. The question is, like, how do you think about reusability between agents or how do agents communicate, right?
- 48:50
Um, I think, uh, this is a thing to be discovered, I think. Like, but I think there's a lot of best practices and system design to be done on, like, um-- Because traditionally with web apps, you're serving one app to, like, a million people, right?
- 49:06
And with agents, like with Claude Code, we serve like, you know, a one-to-one, like, container. When you use Claude Code on the web, it, it's like it's your container, right?
- 49:16
And so there's not a lot of, like, communication between containers. It's a very, very different paradigm. I'm not gonna say that, like, I know exactly the best system design to do that, right?
- 49:26
And like, I think there's lots of best practices on like, okay, these agents are reusing work, um, how can we give them like, like cut-- like general scripts that combine together the work that they've done?
- 49:37
How can we make them share it? Um, I would generally think, uh, this is sort of like a tangent, but on, like, agent communication frameworks, I would say that, like, we probably don't need like a whole-- We don't-- I, I think this is more of a personal opinion.
- 49:52
I think, like, we probably don't need to reinvent, uh, like, a new communication system. There are-- Like, the agents are good at using the things that we have, like HTTP requests and Bash tools and API keys and, uh, named pipes and all of these things.
- 50:07
And so like, probably, like, the agents are just making HTTP requests back and forth from each other, you know, using HTTP server. Um, there's a bunch of interesting work there.
- 50:16
I've seen people make like a virtual forum for their agents to communicate, and they, like, post topics, and we-- like, reply and stuff like that. Um, kind of cool.
- 50:28
I think there's a lot of things to explore and, and discover there. Yeah. Okay. Um, gonna keep going a little bit. How are we doing for time? Okay. It's about an hour left, I think.
- 50:39
Okay. Um, cool. So an example of designing an agent, uh, this is like-- Yeah, let's-- This is not the prototyping session, but I think this is like-- will be a good sort of like, like way into it.
- 50:54
Let's say we're making a spreadsheet agent. Uh, what is the best way to search a spreadsheet? What's the best way to execute code in-- Like, or what's the best way to take action in a spreadsheet?
- 51:04
What is the best way to lint a spreadsheet, right? These are all, like, really interesting things to do. Uh, I'm going to do, like, a Figma, and we can go over it.
- 51:11
Um, if someone could grab a water as well, that'd be great. I, [chuckles] like, could really use water, man. Yeah. Okay. Um, thanks. Uh,
- 51:22
okay, so we're going to, um... Yeah. Let, let's, let's talk through it. Uh, or why don't you spend, like, a couple minutes yourselves thinking about this question? You have a spreadsheet agent.
- 51:34
You want it to be able to search. You want it to be able to, like, gather context, take action, verify its work. How would you think about it, right?
- 51:41
So, like, just spend some time thinking through that. Take some notes or something.
- 51:47
Great. [laughs] [coughing]
- 52:57
Okay. Has everyone g- had a little bit of time to think about this? Does anyone want more time, or wanna just dive into it? Okay. Uh, what's the best way for an agent to search a spreadsheet?
- 53:09
Realizing I have to type with one hand now. Um,
- 53:15
I should figure this out 'cause I'm gonna need to type later. Okay. Um, the-- Okay, searching a spreadsheet. Uh, any, any ideas? How do you search a spreadsheet? Like, what would you do?
- 53:24
CSV. Python.
- 53:27
Okay, you've got a CSV. Okay, now, like, your agent wants to, like, search the CSV. What, what does it do?
- 53:34
Grep.
- 53:34
A grep search. Okay. Uh, what does the grep look like?
- 53:38
Needs to look at all the headers.
- 53:39
Looks at the headers. Okay.
- 53:40
Headers of all sheets.
- 53:43
Okay. Great. Yeah. And let's say I'm looking for the revenue in twenty twenty-four or something. Um, now I've got my headers, like, uh, I'm j- I'm just gonna pull up a spreadsheet, right?
- 53:57
Um, let's say that the revenue is in-- there's a revenue column, and then there's like a,
- 54:04
uh... So yeah, let's see. Okay, so yeah, let's say it's something like this, right?
- 54:24
Like, um, how do I get revenue in twenty twenty-six, right? So this is sort of like a tabular problem, right? Like, there is revenue here, and there's also twenty twenty-six here, right?
- 54:35
So it's like a multidimensional step, right? We could look at the headers that will then give us- Uh, like if you just pull this, you'll get one hundred, two hundred, three hundred, right?
- 54:46
So we need a little bit more. Any, uh, any other ideas?
- 54:52
Yeah, we can-
- 54:52
Yeah.
- 54:52
There's a Bash tool for it, the awk, A-W-K, I think.
- 54:56
awk? Okay.
- 54:57
Yeah.
- 54:58
Yeah, yeah, yeah. And what would it awk for?
- 55:01
Well, it depends on what you, what you're looking for. [laughs]
- 55:03
Yeah. Yeah, yeah. That-- Well, that's the question, right? Like what, what is the user looking for, right? They're probably looking for something like this, like revenue in twenty twenty-six, right?
- 55:11
Um-
- 55:11
Maybe use the APIs to use the Google tools to s- add all the numbers together or VLOOK-VLOOKUP something like this. Like-
- 55:19
Yeah. So the idea is like use the APIs, like use the Google APIs to like look it up. Um, that's great. Uh, but yeah, let's say we're working locally, we need to sort of design these APIs.
- 55:28
Yeah.
- 55:29
SQLite or DuckDB can query the CSV directly, and it works pretty well.
- 55:34
Oh, interesting. Okay. Yeah, I didn't know that. That's great. So yeah, you, you use SQLite to query a CSV. Um, that's a great, like sort of creative way of thinking about API interfaces, right?
- 55:44
Like, um, if you can translate something into a interface that the agent knows very well, that's great, right? And so like, if you have a data source, if you can convert it into a SQL query, then your agent really knows how to search SQL, right?
- 56:00
So thinking about this transformation step is really, really interesting. It's a great way of like, designing like an agentic search interface. So, um, yeah. Over there.
- 56:08
Sorry, real quick. When we're talking about tools, 'cause you can use TSV for some of the stuff as well.
- 56:11
Yeah.
- 56:12
Um, is there any kind of ranking within the tool? Is Claude smart enough to start ranking the right tool for the right job? 'Cause that's kind of what we're talking about here, is right tool for the right job.
- 56:20
Yeah. Is Claude smart enough to right-- rank the w-right tool, tool for the right job? Uh, yeah, if you prompt it, you know, like, or like I, I think that's one of those things where like, I don't know, let's find out.
- 56:28
Like let's read the transcript. Uh, if it's not, like how can you help it?
- 56:33
Like close match or anything?
- 56:34
Yeah, just sort of like, I, I think all of these things are like an intuition, you know? And it's like, like kind of like riding a horse. Not that I've ever rode a horse, but I don't know, just like- [laughs] I imagine it's like riding a horse. [laughs]
- 56:47
Um, yeah. Like you, you, you like, you know, you're sort of giving these signals to the horse. You're calming it down. You're trying to understand what it-- how, how do you push it faster?
- 56:58
You know what I mean? And sort of like it's a very organic, like thing, right? Um, like I think we like to say that models are grown and not designed, right?
- 57:07
So we're like sort of understanding their capabilities. Yeah. Uh, yeah, what-- anybody else? Yeah.
- 57:13
Quick question. So is there a way to add, like metadata to the spreadsheet? Can you give descriptions, like a different document?
- 57:18
Hmm. Yeah, that's-
- 57:19
So for example, like KPIs are kind of an idea that I can use to in fact build intelligence to ask questions about the spreadsheet.
- 57:25
Yeah. So that's another great pattern is like, okay, can you add metadata to a spreadsheet? So th-these are some questions that you might wanna think about before, like when you're thinking about search, is like what pre-processing can you do to make the search better, right?
- 57:39
And so one example is that you translate it into like a SQL format or something, where you use something that can query it, right? That's like a translation step.
- 57:47
Another step is like maybe you have a tool or, um, like a, a pre-processing step where a-another agent annotates the, the spreadsheet and, and like adds information so that the agent can then like search across that information better, right?
- 58:02
So, um, yeah. One more.
- 58:04
Um, I was just curious-
- 58:05
Yeah
- 58:06
... like what-- I mean, all those tools sound great, but-
- 58:09
Yeah
- 58:09
... why can't the agent just, you know, do what was suggested, read the header, and then just get the date and run? Like I feel like that should be pretty trivial for, um, for, for retask.
- 58:21
Yeah, probably I should have like prepared this in code. [laughs] [laughs] Um, but yeah, w- I, I built a ton of spreadsheet agents before. Basically it's-
- 58:28
Does that work?
- 58:29
It, it's kind of hard to do.
- 58:30
Okay.
- 58:30
Yeah, yeah. So, um, basically what I, what I would s-think about is like, so we, we've got like... Okay, I--
- 58:38
Sean, do you have suggestions on how I can-
- 58:39
Voice to text
- 58:39
... how I can code at the ti-same time? Go ahead.
- 58:42
Install voice to text on your computer.
- 58:45
Oh, I see. Yeah, yeah, yeah. [laughs] Mark, do you work at Wispr Flow or something, or? [laughs]
- 58:50
Stick the mic in your shirt.
- 58:51
There's a microphone button. [laughs] There's a microphone button on the back.
- 58:56
Stick the mic in your shirt. [laughs]
- 58:59
Oh, I, I just don't trust that stuff, man. Okay. Um, [laughs]
- 59:05
maybe I, maybe I shouldn't be working in an AI lab. [laughs] Um, okay. So, uh, let's see.
- 59:12
Make someone hold it for you.
- 59:13
Yeah. [laughs]
- 59:14
Make someone hold it for you.
- 59:15
Hold on, hold on. Okay. Um, like let's search. So
- 59:22
one way to do it is like y-you see in spreadsheets, right? Like you can say here, you can design formulas, right? So like B3 to B5.
- 59:39
Right. So this is a syntax, for example, that the agent's pretty familiar with, like B3 to B5, right? And so you can design an agentic search interface, which is like this, right?
- 59:48
Like B3:B5 or something, right? So like your agentic search interface can take in a range, right? It can take, take in a range string, right? And these are things that like the agent knows pretty well, right?
- 1:00:01
Like you can, um, do SQL queries, right? The agent knows SQL queries pretty well, right? Um, and, uh, like these, you can also, uh, do XML, right? Sorry, the font is so small.
- 1:00:17
Um- [laughs] Okay. Uh, yeah, you can also do XML. I, I, it-- I'm not sure if you guys know, but like, uh, XLSX files are XML in the back end, right?
- 1:00:29
And XML is very structured. Uh, you can do like an XML search query. Uh, and there are different libraries that can do that. So that's one example, right? Is like, how do you search and gather context?
- 1:00:39
And I hope this sort of like illustrates to you that like gathering context is really, really creative, right? Like, and, and like there's so many iterations, and if you just If you've only tried one iteration, it's probably not enough, right?
- 1:00:51
Like, think about, like, as many different ways as you can. Like, try these out, right? Like, try SQL or try, try the search, try, try the grep and awk, and, like, all of these things, and, um, have a few tests that you're trying across different things and, and see what the agent likes and what it, what it
- 1:01:06
doesn't like. Um, it's gonna be different for each case.
- 1:01:08
Thariq?
- 1:01:09
Yeah.
- 1:01:10
If you ... When you say agent, you're referring to Claude, the, the model? Or ... 'Cause we're building an agent here.
- 1:01:18
Yeah.
- 1:01:18
And you're relying on already pre-existing knowledge of how to handle XML. Who's, who's doing that, the model?
- 1:01:26
Yeah, 'cause the question is, like, who, w- uh, where does the knowledge come from? Is it the model? Is it, like, what is ... What do I mean by the agent?
- 1:01:32
Yeah, generally, what, I think what you're looking for is, like, you have a problem, you want to make it as in distribution as possible for the agent, right? And so the agent knows a lot about a lot of different things.
- 1:01:44
It knows a lot about, for example, finance, right? So if you ask it to make a DCF model, it knows what DCF is, right? And you can ... If, if you want to give it more information, you can make a skill, right?
- 1:01:55
But, so it's, it knows what DCF is, it knows what SQL is. Can it combine those things together, right? And so, like, uh, ideally, you want to, like ...
- 1:02:05
Your, your problem is gonna be out of distribution in some way, right? Like, like, there's some, like, information that's not on the internet or something that you have, uh, or is something, is somewhat unique to you, and you want to try and, like, massage it to be as in distribution as possible.
- 1:02:19
Um, and, uh, yeah, it's, it's very, very creative, I think. Like, uh, you know, it's not like a ... It's not a science to me. [laughs] It's very much like an art.
- 1:02:31
So, um, yeah, okay, so we, we've tried gathering context, then taking action. Um, we can probably do a lot of the same stuff here that we've done before, right?
- 1:02:43
Like, we can do, like, insert 2D array, right? Um, if, if we've got, like, a SQL interface, right, we can, um, we can do a SQL query. We can edit XML.
- 1:02:58
Um, these are, like, often very similar, right? Like, taking action and gathering context, that, that you probably want a similar API back and forth. And then the last thing is verifying work, right?
- 1:03:08
Like, how do you think about, how do you think about that? Um,
- 1:03:12
check for null pointers, right? Is one of the ways to do it. Um, any other ideas on, on verification or ... Yeah.
- 1:03:23
Sorry, I'm, I'm a bit confused-
- 1:03:25
Oh, yeah
- 1:03:25
... if I may say.
- 1:03:26
Yeah, yeah.
- 1:03:27
Because, like, when, when you're using other SDKs to build the agent-
- 1:03:31
Yeah
- 1:03:31
... I don't need to tell it how it should gather the context.
- 1:03:34
Sure.
- 1:03:35
I just give it the context-
- 1:03:36
Yeah
- 1:03:36
... and explain this is what ... Like, basically, I explain in plain English-
- 1:03:40
Yeah
- 1:03:40
... what it's meant to do.
- 1:03:41
Yeah.
- 1:03:42
And what I tend to do, and you tell me if I'm wrong, I actually end up creating a separate agent for QA-
- 1:03:49
Oh, interesting
- 1:03:51
... to-
- 1:03:51
Yeah
- 1:03:51
... to verify, because I don't trust the agent to verify itself.
- 1:03:55
Hmm.
- 1:03:56
But I'm just, I'm, I'm just a bit, uh, a bit confused about the level of detail I need to provide the agent in that example.
- 1:04:03
Yeah. Okay, so the question is about, um, giving context to the agent versus having it gather its own context. Uh, you mentioned that you sometimes use a QA agent.
- 1:04:14
Uh, can I ask, like, what, like, domain you, you're building your agent in? Or ...
- 1:04:19
In, uh, cybersecurity.
- 1:04:21
Okay. Sure. Yeah, yeah. Um, I think that ... I, I think I'd need to, like, look into more specifics, but the Claude Agent SDK is great for cybersecurity, and, like, I would generally push people on, like, let the agent gather context as much as possible.
- 1:04:39
You know, like, let it find its own work as much as possible. Um, you're trying to give it the tools to find its own work. The way I think about this is kind of like, let's say that someone locked you in a room, and they were, they were, like, giving you tasks, you know?
- 1:04:53
Like, so that's what your, what your job was. Like, a MrBeast sort of, like, scenario, right? Like, you get $500,000 if you stay in this room for six months.
- 1:05:01
Um, then, like, like, someone's giving you a message. What tools would you want to be able to do it, right? Like, would you just want, like, a list of papers?
- 1:05:11
Or, like, would you want a calculator or, like, a computer, right? Probably, I would want a computer, right? I'd want Google. I'd want, like, all of these things, right?
- 1:05:20
And so, like, I wouldn't want the person to send me, like, a stack of papers, be like, "Hey, this is probably all the information you need." I'd rather just be like, "Hey, just give me a computer, give me the problem, let me search it and figure it out," right?
- 1:05:31
Mm-hmm.
- 1:05:31
And so that's how I think about agents as well. Like, they need, like, like, you know, they're stuck in a room.
- 1:05:37
So I need to give them tools. So if you can go back to the slide you have, to the graph you had.
- 1:05:44
To the graph. Like, like this-
- 1:05:46
Like, uh-
- 1:05:46
... you mean, or?
- 1:05:46
Yeah, this one.
- 1:05:47
Yeah.
- 1:05:47
So basically, that gathering context is basically these are the tools I'm offering it.
- 1:05:52
Yeah, exactly. Yeah. You, you're ... I'm giving it, like, maybe an API for code generation. Maybe I'm giving the SQL tool. Maybe I'm giving a Bash. These are all, like, examples, right?
- 1:06:02
So, yeah. You have one more question?
- 1:06:04
Question. So, uh, for all the agents that you're, uh, having to serve the same-
- 1:06:08
Yeah
- 1:06:08
... context, do they, they share the same context window? And what's the size of it?
- 1:06:13
Interesting, yeah. So do agents share the context window? I think, I think this is, like, an interesting question just overall about how you manage context. Uh, I think, and I haven't talked about this too much yet, but subagents are, like, a very, very important way of managing context.
- 1:06:28
Um, I think that this is like we're using more and more subagents inside of Claude Code, and I would think about, like, doing subagents very generally. So, like, what we might do for the spreadsheet agent is maybe we have a search subagent, right?
- 1:06:44
So, like, subagents are great for when you need to do a lot of work and return an answer to the main agent. So for search, let's say the question is like, "How do I find my revenue in 2026?"
- 1:06:55
Maybe you need to do a bunch of results. Maybe you need to, like- Uh, search the internet, maybe you need to search a spreadsheet, things like that. And there's a bunch of things that don't need to go into the context of the main agent.
- 1:07:05
The main agent just needs to see the final result, right? And so that's a great subagent task. Um, I don't have a dedicated subagent slide here, but, like, yeah, they're very, very useful, and I, I think a great way to think about things.
- 1:07:19
Um, yeah.
- 1:07:20
And just to, just to build on that question-
- 1:07:22
Yeah
- 1:07:22
... actually. For verification, for example, you could imagine doing that through a skill or a subagent. You might even want to have an adversarial... Like, the security example is a great one.
- 1:07:32
You want to have it really go to town on it and not really have any sympathetic relationship with the work already done. Uh, it's a very... I, I get it's a spectrum, but do you like...
- 1:07:41
Are you saying yes, you'd use a subagent here, you'd use a skill? How would you think about this?
- 1:07:45
Yeah, definitely. So question on like, uh, do subagents or how-
- 1:07:50
Not sure it'll work, just to make sure.
- 1:07:51
Oh, sure. Okay, yeah, yeah. Thank you. Appreciate it. Um, okay, yeah. Uh, can you use subagents for verification? Uh, yes. I, I think this is a pattern. I think like ideally, the, the best form of verification is rule-based, right?
- 1:08:07
You're like, uh, if there are like a null pointer or something, uh, that's like easy verification. It, does it lint or compile? Like, like, as many rules as you can, try and insert them, and again, be creative, right?
- 1:08:19
Like for example, uh, in Claude Code, if the agent tries to write to a file that we ha- know it hasn't read yet, like we haven't seen the, we haven't seen it enter the read cache, we throw it an error.
- 1:08:31
We, we tell it like, "Hey, uh, you haven't read this file yet. Try reading it first." Right? And that's an example of sort of like a deterministic tool that we insert into the verification step.
- 1:08:42
And so as much as possible, like any time you are thinking about, you know, verification, first step is like, what can you do deterministically? What, like what, like, you know, outputs can you do?
- 1:08:52
And again, like when you're choosing which a- like types of agents to make, the agents that have more deterministic rules are better. You know, like they just like, like it, it just makes a lot of sense, right?
- 1:09:03
So, um, of course, as the models get better and better at reasoning, then you can have these subagents that check the work of the main agent. The main thing there is to like avoid, uh, context pollution.
- 1:09:15
So you probably wouldn't want to like fork the context. You'd probably want to start a new context session and just be like, "Hey, yeah, adversarially check, um, the work of..."
- 1:09:25
Like this, this output was made by a junior analyst at McKinsey or something. They graduated from, uh, like not a great school, like their GPA... Like, you know, like, like just like feed it a bunch of stuff and then tell it to critique it, right?
- 1:09:39
Like that's like one of the tools of the subagent, right? And so, um, yeah, the more you like, uh, yeah, as the models get better and better, that sort of verification will become better as well.
- 1:09:50
Um, but doing it deterministically is like a great start. Yeah. Joshua?
- 1:09:56
Um, just a question about the verified work. So-
- 1:09:59
Yeah.
- 1:10:00
Um, so let's say we found null pointers. It's probably easy to just say, "Okay, fix it." But like, you know, let's say we deploy to production and the client is using it, that's not us, and they somehow get into a spot where the whole spreadsheet's deleted.
- 1:10:17
And so like, like on what level do we need to bake in like ability to like undo tools and stuff? 'Cause like, um, let's say the QA agent determines that their spreadsheet is empty.
- 1:10:31
Yeah.
- 1:10:31
Not necessarily is able to undo before. So like, you know, like what, what's your advice there?
- 1:10:36
Yeah. So the question is like how do you think about state and like undoing and redoing-
- 1:10:42
Mm-hmm
- 1:10:42
... being able to, um, fix errors basically, right? I think this is like, uh, a really good question and honestly another sort of like, um, like when you think about like what are agents good at, right, like or what problem domains are agents good at, how reversible is the work is like a really good intuition, right?
- 1:11:04
So code is quite reversible. You can just like go back, you can undo the Git history. We, we come with like, you know, these atomic operations right out of the gate, right?
- 1:11:13
Like I use Git constantly through Claude Code. I, I don't type Git commands anymore, right? So, um, that's like a really good example. A really bad example is computer use, you know.
- 1:11:22
Because computer use has, is not reversible in state, right? Like let's say you go to like doordash.com and you add, like the user wants you to order a Coke and you add, order a Pepsi.
- 1:11:35
Now, like you can't just go back and click on the Coke. You have to like go to the cart and you have to remove the Pepsi, right? And so your mistake has like compounded this like, you know, this state, and the state machine has gotten more complex, right?
- 1:11:49
And, and so like whenever you're dealing with like very, very complex state machines that you can't undo or redo of, it does become harder, right? And I think one of the questions for you as an engineer is like can you turn this into a reversible state machine, kind of like you said.
- 1:12:03
Can you store state between checkpoints such that the user can be like, "Oh, my spreadsheet is messed up right now, just go back to the previous, uh, checkpoint." Right?
- 1:12:12
Uh, potentially even can the model go back to previous checkpoints. Um, I, I think someone had this like time travel tool, um, that they were giving one of the coding agents, which was kind of cool, where you're like, it's like you can time travel back to a point before this happened.
- 1:12:27
Do you know what I mean? Uh, it's kind of fun. I, I think like all of these tools, uh, some of them don't work that well yet, but you know, we'll, we'll get there.
- 1:12:35
Um, but yeah, thinking about state and verification is, is very useful, right? So, um, yeah. Got a question at the back?
- 1:12:44
Yeah. Um, I'm, I'm kinda curious about scale. Um, so what if the spreadsheet is like millions of rows and million, and, and thou- hundreds of thousands of columns, right?
- 1:12:56
Uh, or just like any sort of database. Like in that type of situation, how would you go about searching? There's obviously a context limit you'd have to account for.
- 1:13:06
Yeah, this is great. Um, I probably should've done the spreadsheet example as my coding example. For, for a preview, my coding, like, agent is a Pokémon agent. Um, probably spreadsheet would've been better.
- 1:13:18
Okay. Uh, the question was, what if the spreadsheet is very big? If you have a million rows, uh, how do you think about-
- 1:13:26
Yeah, a hundred columns. I mean, like, a hundred thousand
- 1:13:27
... a hundred, yeah, a hundred thousand columns or a hundred columns or whatever. Like, how do you think about it, right? Like, y- your database is also very big.
- 1:13:33
Like, how do you, how do you do that? Um, I think for all of these things, uh, one, of course, as the data becomes larger and larger, it's just a harder problem.
- 1:13:42
Like, you know, it, it just absolutely is. Your accuracy will go down, right? Like, Claude Code is worse in larger code bases than it is in smaller code bases, right?
- 1:13:50
As, as the models get better, it will get better at all of that. Um, for all of these, I would think about, like, how would I do this? If I had a spreadsheet that was, like, a million columns and a million rows, what would I do?
- 1:14:02
I, I mean, I would need to start searching for it, right? I would need to be like, like, if I'm searching for revenue, I'd be, like, searching Control+F revenue, and then I'd go check each of these, like, results, and I'd be like, "Is this right?"
- 1:14:14
And then, like, I'd see, like, hey, is, is there a number here? And then I'd probably keep a scratch pad, like a new sheet where I'm like, "Hey," like, "Equals, revenue equals this," you know?
- 1:14:25
And, and, and store this reference and, and keep going. So I, I think that's a good way of thinking about it is, like, the model shouldn't ... You should never, like, read the entire spreadsheet into context because it would, it would take too much, right?
- 1:14:36
Like, um, you want to give it, like, the starting amount of context, and that's also how you work, right? Like, let's say that you open up the spreadsheet. What you see is rows is this, right?
- 1:14:46
You see, like, the first 10 rows and the first, like, you know, tw- 30 columns or something, right? That's what you see. You don't load all of it into context right away.
- 1:14:56
You probably have an intuition for, like, hey, I should load more of this into context, right? And, and like, oh, I should navigate to this other sheet, right, and this other sheet has more data, right?
- 1:15:06
Um, but you need to, like, sort of, you gather context yourself, right? And so the agent can operate in the same way. It can, like, navigate to see these sheets, read them, like, try and, like, keep a scratch pad, keep some notes, and can keep going.
- 1:15:20
So that's how I would think about it. Uh, yeah, at the back.
- 1:15:24
Yeah, so my question is about managing context pollution. It actually, I guess, relates to the previous question.
- 1:15:29
Yeah.
- 1:15:29
Um, do you have a rule of thumb for, you know, what fraction of the context window to use before you start hitting diminishing returns or just it becomes less effective?
- 1:15:39
Hmm. Yeah, the question is... Yeah, context management, do you have a rule of thumb for, like, uh, how much of the context window to use before it becomes less effective?
- 1:15:48
This is actually, I'd say, a pretty interesting problem right now. Um,
- 1:15:54
I think a lot of times when I talk to people who are using Claude Code, they're like, "I'm on my fifth compact." I'm like, "What?" Like, like, I've, I, like, almost have never done a compact before.
- 1:16:05
You know what I mean? Like, I have to, like, test the UX myself by, like, like, forcing myself to get compacted, um, just because, like, I, I tend to, like, clear the context window very often, right, when I'm using Claude Code myself.
- 1:16:18
Just because, like, um, at least in, in code, the state is in the, the files of the code base, right? So let's say that I've made some changes. Uh, Claude Code can just look at my Git diff and be like, "Oh, hey, these are the changes you made."
- 1:16:32
It doesn't need to know, like, my entire chat history with it, you know, in order to continue a new task, right? And so in Claude Code, I clear the context very, very often, and I'm like, "Hey, look at my outstanding Git changes.
- 1:16:44
I'm working on this. Can you help me extend it in this way?" Right? That's, like, a way of thinking about it. And, um, when you're building your own agent, like let's say we're building a spreadsheet agent, it gets a little bit more complex 'cause your users are less technical, right?
- 1:16:58
And they don't know what a context window is, right? Um, that is, like, I'd say, a, a hard problem. I think there's, like, some UX design there of, like, can you reset the conversation state, right?
- 1:17:09
Like, can you ... Maybe every time the user asks a new question, can you do your own compact or something, and can you, like, uh, su- summarize the context?
- 1:17:18
Um, does it ... Like, in a spreadsheet, a lot of the state is in the spreadsheet itself, so it probably doesn't need, you know, to know the entire context.
- 1:17:27
Um, can you store user preferences, um, as it goes so that you remember some of this stuff? You know, like, there's a lot of, like ... Again, like, it's an art.
- 1:17:35
There are, like, so many different angles and ways in which you can do this, right? Um, but yeah, you are trying to, like, sort of minimize context usage. Um, you probably don't need sort of million context or something, you know what I mean?
- 1:17:47
Like, you just need good context management, like UX design. Yeah. Um, yeah.
- 1:17:52
Um, just, I just wanted to ask, the subagents were made to protect the context of the core agent, right?
- 1:17:59
That's right, yeah. Subagents were made to-
- 1:18:00
With the spreadsheet, would we be able to use multiple subagents and try to make the process so we chunk up the spreadsheet in the case where it's super large, so then the agent's gonna kinda run through each portion, like, in parallel of each other?
- 1:18:11
Yeah. Yeah. I mean, um, yeah, so, like, one of the things I love about Claude Code is that we are, like, the best experience for using subagents, like, especially subagents with Bash.
- 1:18:22
It is very, very good. I didn't really quite realize, uh, all the pain. Um, I think if anyone's going to QCon, I believe Adam Wolf is giving a talk on QCon about how we did the Bash tool.
- 1:18:33
Adam's a legend, and the Bash tool, he's done such a good job. Um, when you're running parallel subagents at the same time, Bash becomes, like, very complex, and there are lots of, like, like, race conditions and stuff like that.
- 1:18:45
And, and so there's a lot of work that we've solved there, right? So this is, like, one of the things I love about Claude Code is that you can just be like, "Hey, like, spin up three subagents to do this task," and it will do that.
- 1:18:56
And in the Agent SDK as well, you, you can just ask it to do that. So number one- Yeah. Subagents are a great primitive in the Agent SDK, and I haven't seen anyone do it as well.
- 1:19:05
So that's like a big reason to use it. Um, yes, generally you want it, you want these subagents to preserve context. Let's say you have-- if you have a spreadsheet, you could potentially have multiple read subagents going on at the same time, right?
- 1:19:16
So maybe the main agent is like, "Hey, can this agent read and summarize sheet one? Can this agent read and summarize sheet two? Can this agent summarize sheet three?"
- 1:19:25
And then they return their results, and then the agent maybe spins off more subagents again, right? So this is like another knob you have. Um, and I, I think what I want to say is like
- 1:19:38
i-i-- there's like-- we've talked so many-- so much about like all these different creative ways that you could like do things. This is like the level at which you should think about-- should have to think about your problem.
- 1:19:49
You should not really, in my opinion, think about like, uh, like how, like how do I spin off a process to make a subagent or like, you know, like the system engineering between like, uh, behi- and like what is a compact or something, right?
- 1:20:02
So like we take care of all of this for you in the harness so that you can think about like, "Hey, what subagents do I need to spin off," right?
- 1:20:09
And like, how do I create a, a agentic search interface, and how do I like verify its work? These are the really core and hard problems that you have to solve, and any time you spend not solving these problems and, and solving like lower level problems, uh, you're probably not delivering value to your users, you know?
- 1:20:25
And, and so, um, yeah, I, I think subagents, big fan of the Agent SDK subagents. Yeah. Uh, yeah. Great question.
- 1:20:34
So, uh, like we have this, uh, uh, action and the verification task.
- 1:20:39
Yeah.
- 1:20:39
So where exactly we need to put the verification? In this example, like let's say after generation of the S-SQL query-
- 1:20:45
Yeah
- 1:20:46
... I can verify it is the right query generated or not. That is the one path. Second path is like, uh, generation the query directly executing, and once I will re- get the output, then I will, uh, do the verification.
- 1:20:59
So, uh, and how do-- how agent can tune dynamically, like which one is the right path?
- 1:21:04
Yeah. So the question is like, where do you do verification? Uh, is it only at the end? Do you do it in the middle? Like things like that. I would say like everywhere you can.
- 1:21:12
Just like constantly verific-verification, right? Like, uh, like I said, we do some verification in the read step of the, of Claude Code, right? So that's like a great example.
- 1:21:21
Um, you can do it at the end. You should absolutely do it at the end. But at any other point, if you have rules or heuristics especially, uh, like if, for example, you're like, "Hey, one of my rules is that you shouldn't do like the, the total number of columns you should search is, should be under ten
- 1:21:37
thousand or under a thousand or something," that's like a, a nice way of doing it, right? Like similarly here, like maybe you shouldn't be inserting like a huge like row like of, of values.
- 1:21:47
Like give feedback to the model and be like, "Hey, chunk this up," right? You throw an error and give it feedback, and the great thing about the model is like it listens to feedback.
- 1:21:54
It will read the error outputs, right? And then it'll just keep going. So yeah, verification is definitely like... I, I know I have it in this like as a sort of a loop, but, um, it's definitely more like verification can happen anywhere and, and should happen any-anywhere.
- 1:22:11
Like, like put it in as many places as you can. So, um, all right, I do need to start doing some of the prototyping, but I'll, I'll take one more question.
- 1:22:19
So right, right here. Yeah.
- 1:22:20
How do we say-- how do we form the steps? Like how do we say the agent that go search first and then-
- 1:22:26
Yeah
- 1:22:26
... do this step and then do that step? How does the loop actually start from the start point to the end? How do we actually find-
- 1:22:33
You just tell it. So like, uh-
- 1:22:34
Like, like in the, in the, is it in the system prompt or?
- 1:22:37
Yeah, in the system prompt. Yeah. So like with Claude Code, we just give it the Bash tool and we're like, "Hey, like gather context, read your files, uh, do stuff, like run your linting."
- 1:22:46
You know what I mean? Um, and so yeah, again, with the agent, you don't need to enforce this, right? You don't need to tell it, "Hey, like you need to do this," because like sometimes it might not be necessary, right?
- 1:22:55
Like let's say that someone is asking a read-only question for your spreadsheet. You don't need to like verify that, uh, like your, that there are no compile errors, right?
- 1:23:09
Because there's, you haven't done any write errors, write, write operations, right? So, um, let the agent be intelligent and, and like in the same way that you would like that same freedom when you're doing your work, right?
- 1:23:19
Uh, you're trapped in this box or whatever, like same way, right? Uh, so okay, cool. I, I, I do want to try and see if I can do some prototyping now that we have this, um, uh, the, the holder as well.
- 1:23:33
Um, okay. Yeah, execute, lint. We've done a bunch of Q&A. Okay. Prototyping. Okay. Let's say that you have an agent, right? Like you want, you want to build an agent.
- 1:23:44
You come out of this talk and you're like, "Great, I have a bunch of ideas. How, how do I do this?" Um, I think what I say overall is like building an agent should be simple.
- 1:23:54
Your agent at the end should be simple, but simple is not the same as easy, right? So like it should be very simple to get started, and it is.
- 1:24:02
Just go to Claude Code. Give Claude Code some scripts and libraries and, uh, a cl-custom CLAUDE.md and ask it to do it, right? That's what we're going to do, right?
- 1:24:12
Um, that's like it should be so easy to be like, "Hey, this is my API. This is like an API key. Uh, can you like go search like, you know, I don't know, like my customer support tickets or something and organize them by priority or something like that, right?"
- 1:24:29
And then look at what Claude Code does and, and, and iterate on it, right? And this is like a great way of like just skipping to like the hard domain specific problems that you have, right?
- 1:24:40
So you have a lot of like domain problems, like how do you organize your data, your agentic search? How do you like put guardrails on your database? These are all questions that you can just start solving right away with Claude Code, right?
- 1:24:51
And so try and like build something that feels pretty good with Claude Code. And I think generally what I've seen is that you can do this and get really good results just out of the bat using Claude Code locally, right?
- 1:25:02
And, and you should have high conviction by the end of it, right? And so, um- Yeah. I think, like, [laughs]
- 1:25:11
I forgot this. For more info, watch my AI Engineer talk. Uh, this is, like, a deck for internal that we were using. Um, okay, so, ah, yeah, I'm gonna be inserting this.
- 1:25:23
So y- yeah, you're getting what we're, we're-- what we show customers, right? So, um, okay, uh, yeah. So yeah, use, use Claude Code. Uh, again, simple, but simple is not easy, right?
- 1:25:36
So, like, the amount of code in your agent should not be, like, super large. Doesn't need to be huge, doesn't need to be extremely complex, but it does need to be elegant.
- 1:25:46
It needs to be, like, what the model wants. You want to have this interesting insight. Let's turn the, the model into a SQL query, or, like, just turn this question into a SQL query and then go from there, right?
- 1:25:55
So, um, think about it that way, and Claude Code is, like, a great way of doing that. So, okay, uh, let's make a Pokémon agent, right? This is what we're gonna do.
- 1:26:04
Uh, Pokémon is a game with a lot of information. There are thousands of Pokémon, each with a ton of moves. Um, uh, we want to be pretty general, and so there is actually, like, a PokéAPI.
- 1:26:16
Um, and the reason I chose Pokémon is just 'cause, like, I know that you guys have your own APIs as well, right? And they're all, like, very unique, right?
- 1:26:23
And, uh, so I wanted to choose something with a kind of complex API that I haven't tried before. Um, so the PokéAPI has, like, you know, you can search up Pokémon, like Ditto.
- 1:26:34
Uh, you can search up, like, items and things like that. Um, and so it's got this, like, yeah, this custom API about, uh, uh, everything in the games, right?
- 1:26:44
So, um, and yeah, like, one of the quest things your agent might want, your user might want to do is make a Pokémon team, right? I love Pokémon. I know very little about making an interesting Pokémon team for competitive play.
- 1:26:59
Uh, could my agent help me with that? That'd be, that'd be cool, right? So, um, the-- my goal is to make an agent that can chat about Pokémon, and then we will, like, you know, see what we can do, right?
- 1:27:10
And, and, and how far we get. So, um, I've done, like, some of this work already, and I will, like, open up and show you. So, um, the first step and the prompt here is, like, the first step is I'm, I'm gonna do mostly code generation for this, right?
- 1:27:28
And so, um, let me-
- 1:27:32
Is that gonna be on GitHub somewhere?
- 1:27:34
Uh, actually it is. Uh, yeah, it's on my personal GitHub. Oh yeah, I was going to commit all of this as well.
- 1:27:43
Can I officially-
- 1:27:43
Yeah.
- 1:27:45
Awesome.
- 1:27:45
Um, yeah, yeah. So, uh, I think my personal GitHub is... Let's see. All right.
- 1:27:51
Is it secure GitHub or does it have malware in it? [laughs]
- 1:27:55
You, you guys are AI engineers. You know, like, if you get- That's- -an error, that's, that's your fault. [laughs]
- 1:28:02
Um, yeah. So, um, yeah, you can, you can cl- clone this if you'd like. Um, it need to first ask to change this. So, okay. So, um, yeah. Can, can you guys see this?
- 1:28:14
Should I put it in dark mode instead, or is this fine, like, um- Dark mode. Dark mode? Okay. [laughs]
- 1:28:21
Tabs. They even have that. [laughs] Okay, is this better? Yeah. No? Yeah. You want a different dark mode? [laughs]
- 1:28:38
Dark hard. Okay. I don't think this is what you guys are gonna get, guys. Um, okay. Okay. Let's see. I-- How does this work? Can you guys still hear me or- Yeah.
- 1:28:50
Okay. Um, okay, so here is an example of, like, I've taken... The, the prompt I gave it was, "Hey, I-- Go search PokéAPI for its API and create a TypeScript library," right?
- 1:29:04
And so this is all vibe-coded. Um, and so you can see here that it's created this, like, interface for Pokémon, right? And so it's created, like, this Pokémon API.
- 1:29:14
I can get by name, I can list Pokémon. I can get all Pokémon. I can get species and abilities and stuff like that. And so, like, this is just a prompt that I gave it, right, and it generated this, like, TypeScript API.
- 1:29:27
It also did it for moves. Um, and then it's created this, um,
- 1:29:33
like, uh, it's created this, like, API that I can use. Import PokéAPI, right, from the PokéAPI SDK. And, uh, yeah, you can see, like, sort of how it's, like, set, set this up.
- 1:29:45
And, uh, now in contrast, right? And, and so this is the CLAUDE.md, right? This is the TypeScript SDK for the PokéAPI. Um, this is, like, the, the modules in the PokéAPI.
- 1:29:57
Here are some of the key features. Um, uh, I'm asking it to write scripts in the examples directory, and then it will execute those scripts to help me with my queries, right?
- 1:30:09
Um, and I give it some example scripts. It doesn't always need all this information, right? Like, uh, but yeah, fetching Pokémon, listing the resources, getting data, things like that.
- 1:30:18
So this is, like, my agent really. It's, like, a prompt I gave it to generate a TypeScript library and then this CLAUDE.md, and I, I can chat with it in Claude Code.
- 1:30:28
I'll also show you a version of it that is just tools, right? So here I'm using the messages completion API, right? And I've given it a bunch of tools from the API.
- 1:30:39
So, like, get Pokémon, get Pokémon species, uh, get Pokémon ability, get Pokémon type, get move. So you've defined all of these tools, and you can see that, like, you know, I also just gave it a prompt and told it to make the tools.
- 1:30:53
Um, it doesn't want to make 100 tools, right? Like, there's a ton of Smogon-- or sorry, um, PokéAPI data. Um, but, like, it, it, you know, th- there's only so many parameters it can do.
- 1:31:06
So it's got this, like, tool call and-
- 1:31:11
I, I made like a little chat interface with it, right? So let me now go here and say like, uh, this is my tool calling. Um-
- 1:31:23
You pushed the latest one.
- 1:31:25
Did I? Yes. Great. So yeah, here we've got this chat.ts, right? Um, I, I use Bun when I'm prototyping stuff just 'cause, like, I don't wanna compile from TypeScript to JavaScript.
- 1:31:43
Um, and, uh, again, Bun has, like, linting built into it. Uh, it, it's a way of, like, simplifying for the agent, so the agent doesn't need to remember to compile.
- 1:31:52
But TypeScript is better for generation 'cause it has types, right? So I'm gonna start this, like, Bun chat, and then I'm gonna try, like, okay, what are the generation
- 1:32:02
two water Pokémon? Um, and you'll see that it's, it's starting to, like, search, and I'm logging all the tool calls here. This is very, very important, right? Because, like, it needs to, like, do the tool calls.
- 1:32:16
And so you can see that what it's doing is, like, it's searching a bunch of Pokémon. Um, and then it told me, "Okay, here are the water Pokémon for gen two," right?
- 1:32:25
It's got Totodile, Croconaw, Feraligatr. You can see sort of, like, how it stop... Like, between each step it's thinking through, um, the previous steps, right? Now, like, let's say that I want to do with Claude Code, I think I might need to,
- 1:32:43
uh-
- 1:32:45
Thariq
- 1:32:45
I need to delete this example.
- 1:32:47
Thariq?
- 1:32:47
Um, oh, yeah.
- 1:32:49
Small question. How do you log the, the tool calls?
- 1:32:53
Oh.
- 1:32:53
Just like a, just, just an argument you could-
- 1:32:55
Oh, yeah. This is, um, this is, like, in the normal API, right? So I just, like, uh, in the model, every time it logs in, I just call this.
- 1:33:06
This is in the, like, normal Anthropic API. Um, in the SDK, I, I'll get back to get to the SDK. Um, it's just like you just log every assistant message.
- 1:33:16
So, um, just doing it in console logs. Does that make sense? Or no? Okay. Yeah. So, so the chat interface you were showing- Yeah ... is that just using the regular API or- Yeah, that's using the regular APIs.
- 1:33:28
So not the agent SDK. Not the agent SDK, yeah. Yeah. And so what I'm gonna do here is, um, here, I'm gonna delete this script because I don't want it to cheat.
- 1:33:39
Um, but okay, so here you, you know that, um, I've-- I'm just opening Claude Code. I've created a bunch of files here. I'm gonna say, like, "Can you tell me all the generation two water Pokémon?"
- 1:33:52
Um, and then we'll see what it can do, right? So, um, [coughs]
- 1:33:57
I forget if I need to prompt it to write a script or something. I think I'll be fine. We'll, we'll see what happens.
- 1:34:00
Do you mind going to the core SDK file and just showing, you talked about getting context and then action, and then verification. Can you show that in the code and how we're configuring the tool description?
- 1:34:13
Yeah. So, uh, we haven't done the SDK part yet. So, so far I've just put m- put some APIs in Claude Code. Yeah. Yeah.
- 1:34:23
Sorry, I thought that I missed that. That's why I'm like-
- 1:34:25
No, no, no. Yeah. Yeah. Of course.
- 1:34:26
Sorry.
- 1:34:26
Okay. Um, but yeah, so okay, you can see here, um, it, it's given me a lot more, right? And, um...
- 1:34:40
Yeah, it's given me a lot more. So it, it, it's, it's saying there's 20 water Pokémon, right? And I think this is roughly right. I've like, um...
- 1:34:49
Uh, what did it do? I think it just knows. [laughs]
- 1:34:56
Yeah.
- 1:34:56
That's funny. Live demo this. Um, all right. Um, anyways, uh,
- 1:35:07
yeah, the Pokémon is slightly in distribution, which is, which is, I, I guess good. [laughs]
- 1:35:12
Um, but yeah, so like, what, what it will do is, like, it will try and, like, write, like, a script and, uh, because you don't want it to think as much, right?
- 1:35:22
So here it's like, okay, what I'm going to do is, um, let's see. Gen two water type Pokémon. Yeah. Present.
- 1:35:34
Okay. So yeah, you can see here it, it knows like, okay, the start of the generations, it fetches these, uh, for API. Um, I guess it's decided not to use, like, my pre-built API here.
- 1:35:46
Um, and then, uh, yeah, and, and then runs it, right? So, um, I think I need to, like, improve the CLAUDE.md for this. But anyways, you can see that, like, it, it's able to, like, check 200 plus Pokémon and then check for their type and, and, you know, get their, get their information, right?
- 1:36:05
So this is like, uh, just a quick example on, like, how to do code gen and how to use Claude Code to do it, right? So, um, we'll run this script and then like, uh, um, like keep going, right?
- 1:36:19
So, uh, it will give me the output and, um, yeah, basically what I want to show, let's see, we have roughly 15 minutes left. Um-
- 1:36:32
Does that one play Pokémon?
- 1:36:34
Does that one play Pokémon? Yeah. Yeah. Actually, this is one of the demos I was thinking of doing, um, Claude Code plays Pokémon. So like, let's say you want to do like an agentic version of Claude plays Pokémon.
- 1:36:44
How would you do it? Um, what you would do, I think, is like it would give it access to the internal memory of the, uh, the ROM, right? And so let's say that it wanted to find its party, it could search that in memory, and Pokémon Red is like a very well in distribution, uh, reverse engineered, uh,
- 1:37:05
game, right? And so it could search in memory to be like, "Hey, these are the Pokémon." Um, these are like, this is how I figure out where the map is, this is how I navigate it, right?
- 1:37:15
So this is like maybe-- I actually leave it to the reader if you want to try it out, it's like, um, there is like a Node.js GBA emulator. Um, y- I think I have to legally say you have to go buy Pokémon Red and try it.
- 1:37:27
Um, but yeah, I, I think like, uh- Well, uh, yeah, good example. A- anyways, here. So it's, it's fetched all of them, and it ha- it's listed all their types, and, um, yeah, you can see how it's, like, used code generation to do this, right?
- 1:37:40
So, um, a quick example of using Claude Code to prototype this. Um, now there can be, like, more interesting, like, data here. So, um, I do want to leave time for examples, so I, I think I'll just sort of, like, for questions.
- 1:37:55
So I'll just sort of go through, like, an example. Let's say you're making competitive Pokémon. Competitive Pokémon has a lot of different variables and data. So this is, like, a s- a
- 1:38:07
text file from this online, like, a library, basically, which stores, like, all of the Pokémon and their, like, moves and who they work well with and don't work well with and, you know, like, who they're countered by and all of these things, right?
- 1:38:24
So there's a ton of data here, right? And it's all in text file, um, which is actually pretty good for Claude Code, right? Because I can say, like, "Okay, um, hey, I'm gonna give it a little bit more data."
- 1:38:35
Normally, I'd put this in the, um, check the data folder. Tell me,
- 1:38:41
I, I want to make a team around Venusaur. Can you give me some suggestions based on the Smogon data? Um,
- 1:38:53
and Smogon is, like, this online API. And so I'm n- I'm not entirely sure what it'll do here yet. [laughs] I haven't done this query before, uh, but we'll see.
- 1:39:00
I think it'll be, it'll be fun. Um,
- 1:39:05
where we're at now. This... Oh, I see it. Okay. Um, yeah, but what I wanted to do is sort of grep through this, this data, right? And, and sort of figure out from itself, from first principles, not having seen this data before, how can I, like, answer my query, right?
- 1:39:26
So, um, while it do- it does that, I'll, I'll take any questions. Yeah.
- 1:39:32
Um, first of all, great workshop. Uh-
- 1:39:34
Thanks
- 1:39:34
... so this is, like, purely on top of Claude Code.
- 1:39:37
Mm-hmm.
- 1:39:37
And so my question is, if we were to deploy this customer-facing-
- 1:39:43
Yeah
- 1:39:43
... are we supposed to have Claude Code running in, like, a, like, a s- swarm? Or are we somehow able to take the Claude Code part out, just the bot and the Agent SDK?
- 1:39:54
Hmm. Yeah, so let, let me show you, like, very quickly, like, what the... What, what it looks like to use the Agent SDK here. Um, so I've already done this file system, right?
- 1:40:06
And again, I want you to think about the file system as a way of doing context engineering, right? Like, this is, like, a lot of the inputs into the agent.
- 1:40:13
So my actual agent file is, like, 60 lines, right? Um, and it's mostly just, like, random, like, boilerplate, right? Like, I guess, yeah, it's decided to stop it from, uh, writing scripts outside of the custom scripts directory.
- 1:40:28
Again, fully documented. So, um, yeah, you can see, like, it just runs this query, takes in the working directory, um, and, uh, like, like, runs it in a loop, right?
- 1:40:40
And so probably I'd want to, like, turn into, like, some allowed tools here and stuff, but it, it's very simple. And, and so, um, if I were to, like, productionize this, the first step I'd do is, like, okay, I- I've tested it on Cla- on Claude, Claude Code.
- 1:40:54
It seems to do pretty well. I write this file, then I put it... There are two ways to do it. So one is, I do think that, like, local apps might be coming back with AI because I think that, like, there's such an overhead to running it.
- 1:41:10
Like, for example, Claude Code is a front-end app, right? Like, it works on your computer. So maybe the way I ship this as a Pokémon app is like, hey, I have, like, an app that you install, and it works locally on your computer, and it's writing scripts.
- 1:41:21
I think that's one way of doing it, right? Um, the other way is, yeah, you have-- you host it in a sandbox. Um, and again, there's a bunch of different sandbox providers that make it really easy.
- 1:41:32
Like, Cloudflare has a good example, um, of using the Agent SDK, and it's just, like, sandbox.start, you know? And then, like, bun agent.ts, and that's kind of all it takes, right?
- 1:41:44
Like, it- it's like, like, they've abstracted away a lot of it. Um, so you run, like, the sandbox, um, and then you communicate with it. And, um, yeah, I think there is, like, some very interesting stuff that I'm not sure I had time to get to, but, um, like, I, I think some interesting questions are, like, um,
- 1:42:06
yeah, like, how do you do this sort of, like, service? Now we're spinning up a sub- like, a sandbox per user. Um, there's a lot of, like, I'd say best practices to solve here.
- 1:42:15
One thing I just wanna call out for you guys to think about, um, if you're making a, an agent with a UI, like, let's say that you have, uh, yeah, my Pokémon agent, and I want it to have a UI that is adaptable to the user, right?
- 1:42:29
Like, maybe some users are doing team building, some users are helping it with their game, some users just want pictures of Pokémon. How would, how would I have an agent that adapts in, in real time for my user, right?
- 1:42:41
Um, the way I would do it is in my sandbox, I would have a dev server, right? And the dev server would expose a port. Um, it would run on Bun or Node or something.
- 1:42:50
It would, like, expose a port. The agent could edit code, and it would live refresh, and, and your s- user would be interacting with that website. This is how a lot of, like, site builders, like Lovable and stuff, work, right?
- 1:43:02
They, they use sandboxes, and they essen- host essentially a dev server. And so thinking about this for your, your user, if, if you want a customized interface, this is a great way to do it.
- 1:43:13
Um, okay, let's see, let's see what it did.
- 1:43:18
Um, okay, cool. Okay, so, um, it's, like, written this, like, script. It's generate- like, showed me some base stats and suggested a, like, um- Uh, a move set and some teammates, and you can see sort of like...
- 1:43:41
Let's see. What did it do? Um, Control E.
- 1:43:47
Um, yeah. Okay, so you can see here what it started doing is, like, it started searching for Venusaur, right? And it started finding, uh, those types, the, the, the, like, those Pokémon.
- 1:43:59
And when it does that, it also gets other Pokémon that mention Venusaur. So it gets, like, its teammates and its counters and stuff, right? And it sort of, over this time, found interesting Pokémon, right, that, like, it might work with, right?
- 1:44:13
So it's done a bunch of these searches, and it's got this profile. It's found its most common teammates and, and written a script to, to analyze it, right? And so this is all based on a t-text file.
- 1:44:23
Of course, I could have pre-processed a text file a little bit more. Um, but yeah, the, it's, like, done this sort of, like, interesting,
- 1:44:32
um, an-analysis for me, right? And again, I'll, I'll push out more code to the GitHub repo, and, um, also tweet about this. I'm on Twitter. I'm, uh, [REDACTED:username]. Uh, I tweet a lot, so, uh, definitely, like, mostly about Agent SDK stuff.
- 1:44:46
Um, but yeah, we have about eight minutes left, so I want to spend the rest of the time taking questions about kind of anything. You know, and I'm sorry we didn't get to do more prototyping, um, but, uh, yeah.
- 1:44:57
Over there.
- 1:44:58
Yeah. I was gonna say with the Claude Play, can you, uh, sort of plug this in with that just to see if the agent will, uh, be more selective with the teaming to, uh, try to counter it?
- 1:45:06
Yeah. Put it in, in Claude Plays Pokémon. Yeah. I do want to make Claude Code Plays Pokémon. I think that would be fun. Yeah. I, I think Claude Plays Pokémon, I think we try and keep it, like, a pure reasoning task as much as possible.
- 1:45:16
Yeah. Um, other questions. Yeah.
- 1:45:19
I was curious about how people are monetizing Claude SDK. Like, um, in your product that you, you
- 1:45:26
know, kind of like go and take a Python spot. You kind of, like, lose the opportunity to get all the margins if you make your own input agent.
- 1:45:30
Yeah.
- 1:45:30
I'm curious, like, how have you been-
- 1:45:32
Mm
- 1:45:32
... shipping your own Claude SDK keys so that they kind of take the usage-based as it were versus price margin?
- 1:45:38
Yeah. I, I do think overall, especially right now, agents are kind of pricey. You know what I mean? Because, like, um, the models are, have just started to get agentic.
- 1:45:49
We really focus on, like, having the most intelligent models, you know? And, like, you generally... This is just, like, an overall, like, SaaS business software thing. You'd rather charge fewer people more money that really have, like, a hard problem, you know?
- 1:46:03
And so I think this is still good. Like, you probably should find, um, you know, these hard use cases. But I would say, like, number one, make sure you're solving a problem that people want to pay for, right, is, is like the number one step, right?
- 1:46:15
And then number two, um, yeah, I think you could do subscription or token based. I, I, I think this kind of comes down to, like, how much you expect people to use your product, uh, versus, like, how much you expect them to, like, use it occasionally.
- 1:46:29
Like, Claude Code, obviously, people use a lot, and in order to, like... We do a mix of, like, if we give you some rate limits, and if you exceed it, we do, uh, usage-based pricing.
- 1:46:39
Um, I think that, like, yeah, it's very, like, dependent on your own user base and kind of, like, what they will do. But I will say monetization is something you should s-think about upfront and design your, you know, agent around because it's hard to walk back these promises, so.
- 1:46:57
Um, yeah, back there.
- 1:46:59
Um, I haven't heard you talk at all about hooks and would be curious to hear your take on how hooks might fit into-
- 1:47:05
Uh, yeah. There's so much to talk about. Um, hooks are great. We, we, we do ship with hooks. Um, hooks are a way of doing deterministic verification in particular or inserting context.
- 1:47:17
So, um, you know, we fire these hooks as events, and you can register them in the a- in the Agent SDK. There's, like, a guide on how to do that.
- 1:47:24
Um, examples of things you might use hooks for is, like, for example, um, yeah, you can run it to verify the, like, a spreadsheet each time. Uh, you can also, like, like let's say I'm working with an agent and, uh, I'm-- the agent is doing some spreadsheet operations, and the user has also changed the spreadsheet.
- 1:47:41
This is an interesting, like, place to use a hook 'cause you can be like, "Hey, has..." After every tool call, insert changes that the user has made. Uh, and you, and so you're giving it kind of live context changes, um, in an interesting way.
- 1:47:55
So, um, yeah, I think, uh, uh, yeah, there, there's more stuff on, like, the docs about hooks. Um, I am happy to, like, talk about it afterwards as well.
- 1:48:07
Yeah. More questions? Yeah.
- 1:48:09
So one problem, like, workflow I, I do is-
- 1:48:12
Yeah
- 1:48:12
... let's say I have some sample data. I go through the sample data in Claude Code.
- 1:48:17
Yeah.
- 1:48:17
Then I realize, okay, it's working.
- 1:48:19
Yeah.
- 1:48:19
And I want to take the same conversation that I've already done, because I'm going through a similar data-
- 1:48:24
Yeah
- 1:48:24
... and convert that into a new script.
- 1:48:26
Okay.
- 1:48:27
Uh, which is that I've followed a few steps. Now it's actually working.
- 1:48:30
Mm-hmm.
- 1:48:30
So I don't want to rewrite all of the code to write the SDK, like, again. It's like, because it works.
- 1:48:38
Yeah. Sure. Is it... Yeah. So, like, let's say you've done this prototyping, you've found something that works. What I would do is, like, I'd summarize a CLAUDE.md. Like, obviously, like, when I tried doing this one time, it, like, didn't use my API directly, and it wrote JavaScript.
- 1:48:51
I should've been more specific in my CLAUDE.md to be like, "Hey, you should use this." Um, I... Yeah. I, I think, like, so that's one thing. Um, the second thing is, uh, yeah, do summarize into a CLAUDE.md, have the helper scripts that you need, and then, like, write something like this agent.ts script to, like, to run the
- 1:49:12
Agent SDK. Uh, yeah. More questions? Yeah, in the gray.
- 1:49:15
Uh, yeah. I did, uh, just type, put it up for mine, and it's lying about using the output scripts answer. It tries a couple times. Like, my SDK isn't very good because I coded it.
- 1:49:25
Sure, sure.
- 1:49:25
But then it tries twice, and then it's like, "Oh, here's your comparison table." But it just is-
- 1:49:30
Ah.
- 1:49:30
It's in distribution. Do you have any advice for that kind of problem or-
- 1:49:32
Yeah. This is, this is a good question. And, and, you know, like, I'm-- I think it, there is some messiness, right? Like, I, I think one of the things, if an agent knows so- an answer, um, and you want to, like, sort of like, fight it kind of to be like, "Okay, like, you know, it's generation nine
- 1:49:47
now, and, like, Venusaur stats have changed, and there's this, like, this new, like, character." Like, um- This is hard. I actually think, uh, one of the ways of doing that is hooks.
- 1:49:57
So you can say, for example, like, hey, uh, don't-- um, if, if you've, like, returned a response without writing a script, you know, you can check that. You can be like, give feedback to be like, "Please make sure you write a script.
- 1:50:10
Please make sure you've read this data," right? And, and you can use hooks to, like, give that feedback in, in the same way that in Claude Code, uh, we have these, like, rules, like make sure you read a file before you write to it, right?
- 1:50:20
So add some determinism. Uh, it can definitely be, like I said, it's an art, you know, sometimes, you know, yeah, maybe like wr- like riding a horse, I guess, probably.
- 1:50:29
Um. [laughs] Yeah, in the gray.
- 1:50:32
How are you guys dealing with, like, large code bases? I'm working, like, a fifty million plus line code base, and so-
- 1:50:38
Yeah
- 1:50:38
... grep tool doesn't work really.
- 1:50:40
Mm.
- 1:50:40
Um, so I'm having to build, like, my own, like, semantic indexing type thing to kind of help with that, right?
- 1:50:46
Sure. Sure.
- 1:50:46
Is there any kind of, like, at Anthropic maybe thinking about how that can be more native to the product? Like, you know, in a couple months, is the thing I'm writing just gonna go away?
- 1:50:54
Or like, how, how are you guys thinking about that?
- 1:50:57
Okay, your last question, in a couple months, is your thing gonna go away? [laughs] Generally, yes. Yeah.
- 1:51:01
Yeah. [laughs] [laughs]
- 1:51:01
Like, anytime you ask about AI, yeah. Uh, I think that, um,
- 1:51:07
semantic search, uh, uh, this is a Claude Code question more than an Agent SDK question, but happy to answer it. Like-
- 1:51:12
Okay
- 1:51:12
... um, we, you know, there are trade-offs of semantic search. It's more brittle. Um, I think you have to, like, index and s- and, and search. And we've-- it's not necess- the model is not trained on semantic search, and so I think that's sort of like a problem.
- 1:51:26
Like, you know, grep it's trained on because it's like, it's easy to do that. But like semantic search, you're implementing your bespoke query. Um, for like very large code bases, you know, we have lots of customers that work in large code bases.
- 1:51:38
I think what I've seen is sort of like they just do, like, good CLAUDE.md files. You start in, you know, try and make sure you start in the directory you want.
- 1:51:49
Have, like, good, like, verification steps and hooks and links and things like that. And so, um, you know, that's what we do. We don't have, you know, a custom...
- 1:51:57
We, we dog food Claude Code, right? So, um, yeah. Okay, yeah, one last question.
- 1:52:02
We have to close unfortunately, actually.
- 1:52:04
Okay.
- 1:52:04
So let's give it up for Thariq, everyone. [applause] [upbeat music]