AI Engineer Summit 2025
Specialized RAG Agents: Lessons learned from deploying complex AI systems in production
Read the talk
Production RAG Agents: Turning Enterprise Context Into Real Business Value
Douwe Kiela explains why enterprise AI succeeds when teams build specialized retrieval systems around proprietary knowledge, design for production early, integrate with existing workflows, and make failures observable.
From a talk by Douwe Kiela
At a glance
Ideas worth remembering
The context paradox explains why impressive general model capabilities do not automatically translate into enterprise ROI: differentiated outcomes require organization-specific context and expertise. 1:16
Treat the complete RAG system, not the language model in isolation, as the unit responsible for solving the business problem. 4:24
Specialize around proprietary institutional knowledge and build systems capable of handling noisy enterprise data at scale. 5:31
Design for production-scale documents, users, security, and compliance early, while iterating quickly with feedback from actual users. 7:45
Drive adoption by reducing routine engineering overhead, embedding AI into existing workflows, and helping users discover an immediate, meaningful benefit. 9:51
Manage unavoidable errors through observability, evidence-backed attribution, audit trails, and claim checking, while targeting problems capable of producing meaningful business value. 13:20
The context paradox behind enterprise AI
Douwe Kiela, CEO at Contextual AI, frames enterprise AI as a mismatch between enormous expectations and inconsistent realized value. Organizations invest heavily, yet leaders still struggle to show a clear return. He connects this tension to a robotics paradox: tasks that appear intellectually difficult can be easier for machines than seemingly ordinary activities that require situational understanding. 0:17
In enterprise AI, the analogous problem is context. Language models can perform impressive coding and mathematical tasks, but they still struggle to place information within the right organizational situation. Human specialists do this almost automatically by drawing on accumulated expertise and intuition; enterprise systems must reproduce something closer to that contextual judgment before they can solve consequential business problems. 2:18
This creates a tradeoff between convenience and differentiation. General-purpose assistants can make employees more efficient, but business transformation depends on handling the specific context embedded inside an organization. The more differentiated the desired outcome, the more effectively the system must incorporate enterprise knowledge rather than relying on generic model capability alone. 2:18
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build specialized systems around enterprise knowledge
Kiela argues that a language model may represent only 20% of a much larger system. In enterprise deployments, that broader system often takes the form of a RAG pipeline that connects generative AI with organizational data. His practical comparison is that a merely adequate model surrounded by an excellent retrieval pipeline can outperform a stronger model surrounded by a poor one. The relevant unit of engineering is therefore the complete system that solves the business problem. 4:24
The next design choice is specialization over AGI. General-purpose assistants struggle to match the expertise already present inside a company, especially when the problem is difficult, domain-specific, and sufficiently well understood. A specialized system can be organized around that existing institutional knowledge rather than expecting generalized intelligence alone to reproduce expert performance. 5:31
Proprietary enterprise data supplies the basis for this differentiation. Kiela treats organizational data as a durable expression of what makes a company distinctive, while cautioning against assuming that extensive manual cleaning must precede useful AI deployment. The harder but more valuable capability is enabling AI to work with noisy data at scale; success there turns existing information into a competitive advantage instead of requiring the organization to sanitize everything first. 6:38
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Design for production early, then improve through real use
A working pilot can create a misleading impression of readiness. A small RAG demo assembled from an existing framework and a modest document collection may impress its first users, yet production introduces much larger corpora, many more users, numerous distinct use cases, and organizational expectations that the demonstration never had to satisfy. Kiela emphasizes that moving to tens of thousands, hundreds of thousands, or millions of documents is substantially harder than building the initial proof of concept. 7:45
Production also introduces security and compliance requirements alongside scaling challenges. Kiela’s recommendation is to design for production from the beginning rather than optimizing exclusively for a pilot and attempting to retrofit operational requirements afterward. This is not a call to delay release until every detail is polished; it means accounting early for the conditions under which the system must eventually operate. 7:45
Within that production-oriented architecture, speed matters more than perfection. Teams should put a barely functional system in front of actual users early, rather than relying exclusively on friendly testers, and then improve it through their feedback. This iterative approach makes deployment a process of incremental improvement toward usefulness instead of a single attempt to produce a flawless system before confronting real working conditions. 8:46
Start with a barely functional deployment.
Early deployment creates feedback that supports iterative improvement.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Optimize engineering effort and user adoption together
Fast iteration requires deciding what engineers should not spend their time doing. Kiela points to chunking strategies and prompt adjustments as examples of implementation details that vary across use cases and frameworks but can divert attention from the more important question of delivering differentiated business value. Where capable RAG-agent platforms can abstract those details effectively, engineering effort should shift toward the problems that actually distinguish the organization from competitors. 9:51
A deployment is not successful merely because it is technically running. Kiela describes cases where production AI systems attract almost no usage, either because organizational review processes leave them barely useful or because employees do not know how to apply them. Workflow integration addresses this adoption gap: the closer a system fits existing enterprise workflows, the more likely it is to become part of actual day-to-day work. 11:01
Onboarding should also minimize the time required for users to experience a concrete benefit. Kiela describes Contextual AI running in production globally with Qualcomm and thousands of customer engineers; in one example, an engineer discovered a seven-year-old document that had been hidden from view and finally obtained answers to longstanding questions. The lesson is not simply that retrieval can locate old files, but that a personally meaningful discovery can create the early confidence and internal advocacy needed for broader adoption. 12:08
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make failures legible and pursue meaningful outcomes
Kiela treats accuracy as necessary but insufficient. Even if a system reaches 90% or 95% accuracy, an enterprise must still decide what happens in the remaining cases, and perfect accuracy may be unattainable. Once a minimum quality threshold is met, the practical issue becomes how the organization understands, investigates, and manages the errors that remain. 13:20
His answer is observability, including careful evaluation, audit trails, and attribution to supporting documents. In regulated settings especially, an organization needs to understand why an answer was generated and what evidence supported it. Kiela also recommends checking generated claims and applying post-processing so that attributions are substantiated rather than merely attached superficially. 13:20
Finally, ambition should be measured by the potential business value of the problem being solved. Kiela contrasts consequential enterprise applications with assistants limited to basic questions about benefits providers or vacation allowances: those narrow conveniences may be easy to deploy without producing meaningful ROI. His closing argument brings the lessons together: build complete systems instead of chasing models, specialize around enterprise expertise, make failures inspectable, and choose problems whose successful resolution would materially matter. 14:22
Document used to support an answer.
Document attribution and claim checks help organizations investigate generated answers.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:00
[on-hold music] My name is Douwe Kiela.
- 0:17
I'm the CEO at Contextual AI, and I'm here to talk to you about RAG in production, RAG agents specifically. Um, and I'll, I'll share some of the lessons that I've learned.
- 0:27
So my background is in AI research, uh, but after that, I became the CEO of a AI company focused on enterprise. Uh, so I thought I would share some of my learnings with you, uh, in the hope that that's useful.
- 0:41
So if you, uh, look at enterprise AI, uh, if you work in this space, you probably notice that there's a huge opportunity, uh, ahead of us, right? Everybody wants to grab that opportunity.
- 0:52
There's, there's these huge numbers flying around. Four point four trillion dollars is the, is the estimated added value to the global economy, according to McKinsey. So we have this giant opportunity.
- 1:03
But at the same time, if you actually look at what's happening in enterprises, you see a lot of frustration. It's probably even true for some of the people in the audience right here.
- 1:12
If you're a VP of AI, then you're probably under some pressure right now. It's like, "Where's the ROI? We're investing all this money in AI, but where is it actually leading us to?
- 1:23
Are we getting something out of this?" So, uh, Forbes has this interesting study where they showed that only one in four businesses actually get value from AI. So why is that happening, right?
- 1:35
It feels a bit like a paradox. Uh, so to, to explain it, we can look, uh, at a paradox that might be familiar to you. It's, it's something called Moravec's paradox.
- 1:45
It's from robotics. And in robotics, they were very surprised when they found out that it's actually much easier to beat humans at chess than to have a robot that can vacuum clean your house or have a self-driving car.
- 1:59
Um, so the, the, the paradox here is really that things that seem hard are actually much easier for computers than you would expect, and things that seem easy actually turn out to be much harder, right?
- 2:11
So there's something very similar happening right now in, in enterprise AI specifically, and this is around context. So on the one hand, we have these amazing language models, right?
- 2:21
You've, you've-- That's why we're all here, basically, because we see this revolution happening right in front of our eyes. So language models can generate code much better than most humans.
- 2:31
They can solve mathematical problems much better than, than most of us here can do, uh, and we're pretty smart. Um, so it's really amazing what they can do. But one of the things that they really still struggle with, and that's one of the things that as humans we are very good at, sort of without effort, is putting
- 2:49
things in the right context, right? So as humans, we build on our expertise, we build on our intuition that we've developed over the years, especially if we're a specialist.
- 2:58
This is something that is very easy for us to do, is to put something in the right context and, um, and in the right situation so that you can make sense of the information or the, the problem that you're solving.
- 3:11
So I would argue that this is really the key observation, this, this context paradox, um, for unlocking ROI with AI. And the reason for that is that where we are right now here is, is in the bottom left, right?
- 3:24
So we're, we're mostly focused on convenience. We have general-purpose assistants. They're very useful. Mostly, if you're lazy, they help you sort of solve your problems faster. But where you really want to get to is differentiated value.
- 3:37
If you're an enterprise, i-it's nice that you can make, uh, things more convenient. You probably can make people more efficient and more productive. That's great. But where you want to get to is this business transformation ideal, right?
- 3:49
That's what all the CEOs are probably telling you as a VP of AI. Like, "I want to change my entire business. How am I going to do that?" So getting to that differentiated value, that's where you want to get to.
- 4:00
But the problem is that the higher you go on that axis, the further you go on the context axis. So the, the better you need to be at handling the context, uh, uh, that exists within your enterprise.
- 4:14
Um, so what should we do about that? Um, so that observation is really why I started the company that I'm currently the, the CEO of, Contextual AI. Um, and we started this two years ago to try to help bridge this gap, um, and we've learned some lessons along the way that I thought I would share with you,
- 4:32
uh, in the hope that they're also useful for you.
- 4:35
So the first observation i-is really that language models are awesome, but often they're only twenty percent of a much bigger system. Um, so if you have an enterprise AI deployment, usually that means it's a RAG system.
- 4:50
Uh, so I, I think everybody here probably has heard of RAG. Uh, RAG is something that I, uh, originally pioneered with my team at Facebook AI Research when I was there.
- 5:00
Uh, so RAG is, is really kind of the standard way that you get gen AI to work on your data. So what happens very often these days is a new language model comes out, everybody goes, "Whoa, lu- new language model.
- 5:12
It's great." Everybody starts to think just about the language model, but very few people actually think about the system around the language model, and that system needs to solve the problem, right?
- 5:21
So you can have a relatively mediocre language model, but an amazing RAG pipeline around it, and that's going to be much better than an amazing language model with a terrible RAG pipeline around it.
- 5:32
So the basic observation here, or the lesson, is that you should be thinking about systems, not about models. The model is only a small part of the system, and the system is the thing that solves the problem.
- 5:43
The next observation is that if you're in an enterprise, expertise is really your fuel, right? So, uh, one of the, the things that you want to be able to do as an enterprise is unlock all of that expertise.
- 5:54
So you have all of this i-institutional knowledge in your company. How do you get it out? Um, so one way to try to do that is, is using these generalist, kind of general-purpose assistants, but it's very hard to get them to, to, uh, to match the expertise of people in your company.
- 6:11
So ideally, what you want to do is to specialize so that you can capture that expertise much better. So, uh, at my company, we call this specialization over AGI.
- 6:21
AGI is great. There are lots of use cases for it. If you really want to solve a very difficult problem that is very domain-specific where you understand the use case, you want to specialize for it, and you'll get much further.
- 6:32
So that's, I guess, uh, pretty counterintuitive if you look at the sort of broader, uh, interests, right? Most people are much more excited about AGI, but solving real problems is much easier with specialization.
- 6:45
The next lesson is, uh, at an enterprise scale is your moat. So if you think about what a company really is, is a company maybe its people? Probably a little bit, right?
- 6:55
But over time, what the company really is or what makes a company a company is its data, because even people are transient, right? So the data that a company owns, that is the company in the long term.
- 7:07
So now as an enterprise, you need to think how you can unlock all of that potential, right? And so, uh, one of the, the big issues that we see a lot is that enterprises think that, uh, you need to scrub the data and clean it and invest a lot of time in, in, uh, making your data accessible
- 7:24
with AI. But what you really want to do is make sure that AI can work on your noisy data at scale. And doing that is incredibly difficult, but if you succeed in doing that, that's how you get to differentiated value.
- 7:36
That's how you get that moat, because the data makes, makes your company your company, and so that data is really your moat.
- 7:44
Uh, one observation, uh, and this is really a, a hard truth that, that we've learned, and I, I think that many of you might have learned already or that you're about to find out, uh, if you're earlier in, in your journey, is that pilots are very easy.
- 7:58
Uh, building a demo is not very difficult these days, right? If you want to build a RAG system, you take one of the frameworks, you put in some documents, you have a working solution.
- 8:07
It's great. You give it to your ten users, they all tell you it's fantastic. And then you show it to the CEO, and he says, "Okay, we're going to fire half the customer support team, and we're going to replace them with AI, and we're gonna do that in three months."
- 8:20
And now you're on the hook for productionizing something that is actually much, much harder. And so getting this to work at tens of thousands or hundreds of thousands or millions of documents, you can't do that with any existing tools, uh, that are out there on the open source market.
- 8:35
It's very, very difficult to do that. Making this scale to thousands of users is very hard. Um, making it work for lots of different use cases. If you're an enterprise, maybe you have twenty thousand different use cases that you want to cover.
- 8:46
So how do you scale if that's the problem that you're solving? And then there's, of course, enterprise requirements around security and, and compliance. So bridging that gap is much harder than you think.
- 8:56
And, and the, the right way to deal with that is to really focus on production from day one. So don't design for the pilot, design for production, uh, and that can save you a lot of time.
- 9:06
And that brings me to the next observation, is that speed is really much more important than perfection. What we see, um, in, in terms of production roll-outs of, uh, RAG agents, it's all about speed.
- 9:20
Um, and, and what that means is, uh, you need to give it to your users relatively early, real users, not, not sort of, uh, testers who are, are kind of friendly.
- 9:30
You wanna give it to real users to get their feedback. You wanna do that early. It doesn't have to be perfect. It just needs to be barely functional. And if you do that, then you can hill climb to actually get to this level where it's good enough.
- 9:42
If you don't do that, and you wait too long, and then you try to design something that is perfect, it's gonna be very hard to, to bridge that gap from pilot to production.
- 9:51
So iteration is really the key to a lot of, uh, successful, uh, production AI deployments in, in enterprises.
- 9:59
Next observation is, is related to this too, which is that, uh, if you want your engineers to be fast, and if you want to follow that, that speed maxim I just talked about, then you don't want them woring-- working on boring stuff.
- 10:12
Uh, sounds kind of obvious, but it turns out that engineers are working on a lot of very boring stuff. Um, and, and so one of, one of the, the things that they have to worry about, for example, is what is the optimal chunking strategy for my RAG system?
- 10:25
And it's different for every use case, and it's different for every framework. And then they have to think about what the right prompt is or really basic things that ideally they don't have to think about too much because you really want your engineers to think about, how am I going to deliver business value, right?
- 10:41
How, how do I make sure I have this differentiated value and that I'm actually better than my competitors? Um, so make sure that your engineers spend time on the things that matter and not on the chunking strategy or, or things that, that can be abstracted away, uh, very well these days by, by state-of-the-art platforms for, for RAG
- 10:59
agents. Next one is, is about making AI easy to consume. So what, what I mean by that is we actually see, uh, this happen quite often where companies have gen AI running in production.
- 11:13
And then the next question I often ask them is, "Okay, how many people are actually using it?" And, and surprisingly often, the answer is zero. Almost nobody's actually using it.
- 11:23
They did all this work, but they had to make sure it, it came through, uh, sort of model risk a-and, and, uh, teams like that. So it was really like kneecapped almost, and now it's barely useful.
- 11:35
Uh, so that's one scenario, or, or very often people just don't actually know how to use the technology. So it, it really is a journey that you are on, uh, and the easier you can make your solutions to consume, the better it is.
- 11:48
And what that, what that means for most enterprises is not just thinking about your enterprise data and how you make AI work on it, but also how you integrate it into their workflows.
- 11:58
So the closer you can integrate it into a workflow that already exists in your enterprise, the more successful you're going to be with real production usage.
- 12:07
Uh, next one is, is related, uh, to the, to the previous one as well, uh, where it's really about getting usage. It's about sort of being sticky and, and so this sounds maybe a, a little obvious-
- 12:20
But the, the quicker you can wow users or get this sort of spark where they, they suddenly get it, like this for, for me as a CEO of a AI company, that's really the special moment when people suddenly go like, "Wow, I didn't know that it could do this."
- 12:34
Um, so you can try to design your experiences for onboarding users around this observation too, right? So where, where they get to the wow as quickly as possible. So for us, we have this really nice example with someone at Qualcomm.
- 12:47
So we're, we're running in production globally with Qualcomm, with thousands of customer engineers, and one of them became so happy when they found this document. It was seven years old.
- 12:56
It was hidden away somewhere. They didn't know it existed. They had all these questions, and they just never knew what the answers were, and suddenly because they asked our system, they got these answers, and like their, their world was never the same again after that.
- 13:09
Um, so these are the, the small, the small wins sort of, uh, that, that really matter for, for, uh, evangelizing, uh, production in AI.
- 13:20
Uh, so that brings me to the, the penultimate learning, which is that it's not even really about accuracy anymore. So accuracy is almost table stakes, right? Uh, so I, I think as AI practitioners, we probably know that getting a hundred percent accuracy is very hard, if not impossible.
- 13:36
Getting ninety-five percent accuracy, maybe you can get there, or ninety percent. But what enterprises are, are thinking about much more these days is, what about the missing five percent?
- 13:45
Or what about the missing ten percent? How do I deal with the things that might go wrong, right? Um, so there's a minimum requirement for accuracy, but beyond that, it's really about inaccuracy, and the way to deal with that is through observability.
- 13:58
So you wanna be very careful with how you evaluate these systems. You wanna be very careful with making sure you have proper audit trails, especially if you work in a regulated industry.
- 14:06
This is incredibly important, right? Making sure that you have an audit trail that says, "This is why I generated this answer. It's because I found it here in this document."
- 14:15
Basic things like that. So attribution essentially in a RAG system actually becomes very, very important for dealing with the inaccuracies. And similarly, what you can do is you can check the claims that your system generates.
- 14:27
So do a lot of post-processing to ensure that you have proper attributions, uh, that, that you can really back up, uh, as evidence. [coughs]
- 14:39
Struggling with the clicker a bit. So, uh, fi-final one, uh, that I wanna end on, and this sounds maybe a little bit cheesy, but it, it really is, is true, is be ambitious.
- 14:50
We, we actually see a lot of projects fail, not because people are aiming too high, but because people are aiming too low, where, where folks are going like, "I have gen AI running in production," and then, "What does it do?"
- 15:02
"It answers basic questions about who your 401(k) provider is, or how many days of vacation I get." Like, that's not really where the ROI of AI is, right? So you want to aim for really ambitious things where if you solve them, you actually have, have ROI, and you don't just have a gimmick that people don't really, uh,
- 15:19
use anyway. So try to be, be ambitious because we really live in special times. We have the, the astronaut here on the slide. Um, so, so I, I think it was a pretty special time to be alive during the, the Moon landing and when all of that happened, right?
- 15:34
We're in a, in a similar moment right now, where AI is, is really going to change everything, is going to change our entire society in the next couple of years.
- 15:43
Um, and so you have an opportunity being in the role that you're in to, to really, uh, uh, affect that change in society yourself. Uh, so, so be am- ambitious when you do that.
- 15:54
Uh, and don't aim for the, the low-hanging easy fruit. Uh, aim for, for the sky.
- 16:00
So, uh, that's really, uh, what, what my lessons were for you here. Uh, this context paradox is not going away, um, but by understanding these lessons that I, I shared with you, hopefully you can turn some of these challenges, uh, that we see everywhere in enterprise AI into opportunities for yourself.
- 16:17
Um, so, so it's really build better systems, think about systems, not models, focus on your expertise and specialize for it. Don't settle for general solutions. Specialize for, for the expertise that you have in your company, and be ambitious, and then you'll be very successful.
- 16:34
Thank you. [outro music]