How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI
Read the talk
How long can your skills be before your agent forgets what you told it?
Laurie Voss reruns IFScale and finds a roughly tenfold increase in simultaneous keyword constraints over a year. Longer skills become plausible, while four different failure patterns explain why outputs still need checking.
From a talk by Laurie Voss
At a glance
Ideas worth remembering
IFScale measures exact-word inclusion in a business report. Its results demonstrate constraint tracking, rather than reliable reasoning over equally large real skills files.
The original ceiling near two hundred to three hundred requirements moved into the thousands in the newer tested versions. GPT-5.5 reached ninety-nine percent accuracy at five thousand rules.
Omissions, API refusals, exhausted reasoning budgets and late abandonment require different detection. Fluent prose can still be an incomplete deliverable.
Reconsider fragmentation imposed by an old capacity limit, while measuring cost, latency and compliance. Wording and order can still change whether the instructions are followed.
Two hundred instructions disappear quickly
A skills file can accumulate two hundred instructions surprisingly quickly. Conditional behavior, required sections, forbidden phrases, tone rules and edge cases each add another obligation. If an agent quietly loses track of them, the file’s apparent sophistication exceeds what the model actually does. Laurie Voss, head of developer relations at Arize AI and co-founder of npm, begins with that practical concern: how many instructions can fit before compliance starts to fall?
The investigation starts with an aside in a talk by Dexter Horthy at AI Engineer Miami: an agent could follow about two hundred instructions before forgetting some, based on a 2025 result. That would make a substantial skills file a risky engineering choice. A polished response offers little reassurance. It might follow the full specification, or merely resemble the answer the developer expected.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
IFScale turns instruction following into a count
IFScale’s task is simple: write a business report containing a list of exact words. The opening examples are customer and revenue. Each required word becomes one separately checkable instruction. After generation, the benchmark counts how many appear.
Two quantities describe the experiment:
- Density, N: The number of required words supplied together in one request.
- Accuracy: The percentage of those requirements satisfied by the generated report.
The report gives the model somewhere to use the words. Its business insight is not what the score measures.
Follow the revenue requirement through generation. It stays on the input list whether or not the answer sounds convincing. If the report omits that exact word, the requirement fails. Adding more words increases the obligations the model must carry into its answer; counting their appearances makes the loss visible. This answers a question that casual reading struggles with: did each named constraint survive?
The connection to skills is a proxy. Requiring revenue resembles requiring a pricing section because both specify an identifiable obligation. Real instructions can demand reasoning, conditional behavior or resolving conflicting requirements. Voss therefore treats keyword capacity as an optimistic ceiling for more complicated work. Passing this task demonstrates constraint tracking; it does not prove that an equally large real skills file will work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Replicate the old ceiling, then outgrow the benchmark
Before testing newer models, Voss repeats the original experiment. Only three of the paper’s ten models remained accessible through APIs when he ran the replication: GPT-4.1, Claude Sonnet 4 and Gemini 2.5 Pro. Availability determined the comparison set. By the talk, another had been retired—hence his warning, “don’t get attached to your models.”
The replicated curves matched the original paper within what Voss describes as its noise boundary. Accuracy began deteriorating around two hundred to three hundred rules. At five hundred, the models were missing roughly thirty to fifty percent of the required words. The old ceiling had a measurable basis.
The same prompts and words then went to GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro and DeepSeek V4 Pro. All four scored one hundred percent on the original test. Opus 4.8 appeared a week after testing, so these are results for the named tested versions. The benchmark’s maximum of five hundred required words no longer reached their failure region.
The report example now changes in one crucial way: the list grows. Five hundred requirements become one thousand, then two thousand, with testing extended to a ten-thousand-word vocabulary. Each expansion adds words that must survive into the answer. The experiment can finally expose where accuracy bends instead of stopping at a perfect score.
How far did the ceiling move? The results discussion at 7:12 compares deterioration around two hundred to three hundred requirements with newer curves extending nearer two thousand, and the strongest reaching five thousand. Voss describes this as close to a tenfold gain over about twelve months. The horizontal axis is logarithmic: equal distances represent multiplicative changes in rule count. A steep-looking decline can span many additional constraints.
This is a gain in a specific capability: keeping many named requirements in play during generation. It can be much larger than the improvement a developer feels in ordinary conversation. Architecture built around a small instruction budget may now preserve fragmentation that its chosen model no longer needs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The business report fails in four different ways
Extending the list changes more than the score. Older models mostly returned reports with missing words. The newer tested versions fail at different points between receiving the request and finishing the answer. Their behavior matters because a single keyword-accuracy curve can hide very different causes.
-
DeepSeek V4 Pro: missing requirements. Forgetting starts around seven hundred and fifty rules. By two thousand, nearly half are dropped. A report still arrives, and keyword counting measures the damage directly. Voss prefers this failure’s predictability.
-
Claude Opus 4.7: API-level refusal. Combinations in the random vocabulary trigger refusals. Voss attributes them to safety classification, citing
anthraxandcyanideappearing together. A large random list can resemble a dangerous request despite its benchmark purpose. -
Gemini 3.1 Pro: reasoning consumes the answer budget. Performance remains strong to about five thousand instructions. Beyond that, Voss describes the model spending its token budget checking requirements, leaving little or no useful report. The request incurs work without delivering the answer.
-
GPT-5.5: a plausible report followed by abandonment. The strongest tested result reaches ninety-nine percent keyword accuracy at five thousand rules. When pushed far enough, the model begins writing, objects to the task and stops. Failure becomes apparent at the ending.
Claude required an input change. Voss passed the vocabulary through OpenAI’s safety filter and removed words that looked problematic; the filtered task then performed well. Claude’s successful completion therefore depended on vocabulary filtering. A refusal caused by the list’s content is a different limit from losing track of its length.
GPT’s partial report makes the opening anxiety observable. The model accepts the assignment and writes substantial prose, then rejects the premise of producing a coherent business report containing thousands of unrelated words. The unfinished answer lacks many required keywords. Its objection may be reasonable; the deliverable is still incomplete.
Where can the report disappear or become incomplete? The diagram separates refusal before writing, budget exhaustion during reasoning and defects in returned text. It makes the detection problem visible: an API refusal announces itself, while omissions and late abandonment require inspecting the answer. Receiving fluent prose does not establish completion.
A growing list of exact-word constraints accompanies the request.
Refusal happens before a report; reasoning can exhaust the budget; returned text can omit requirements or end in abandonment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Longer skills turn capacity into a cost decision
A two-hundred-instruction budget encourages compression and fragmentation: keep the main file short, send the agent to sub-skills and coordinate additional files or specialized agents. Voss calls the result a “Byzantine labyrinth.” Greater capacity makes it plausible to put a hundred or three hundred straightforward requirements together without that machinery.
A style guide illustrates the potential change. Brand rules and legal disclaimers can amount to thousands of named constraints. Keeping more of them in one prompt reduces the need to distribute instructions solely to accommodate the old ceiling. The practical question then becomes whether each addition earns its cost: very large prompts take more time and money to process. Shortening or splitting instructions can still serve those purposes.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Tracking requirements does not guarantee clear reasoning
There is no universal breaking point. Voss gives a range from roughly seven hundred and fifty to more than nine thousand requirements across the tested models. Selection matters even for this narrow task. A separate question remains: can a model reason clearly over a giant prompt, especially when its instructions conflict?
Voss cites subsequent long-input research across eighteen models reporting accuracy declines of thirty to fifty percent before the context-window limit. He describes a counterintuitive result in which coherent, structured text suffered more than shuffled material, while leaving its cause unexplained. Fitting text into the window, retaining named constraints and reasoning over that material are distinct capabilities.
The unfinished report creates a practical warning. An API refusal announces failure. A confident beginning followed by abandonment makes failure harder to notice. Monitoring must inspect what the application actually returned, rather than treat acceptance of a large prompt or polished opening paragraphs as success.
Voss reports approximately two thousand three hundred calls across seven models costing $29. A simple experiment made a consequential engineering assumption affordable to retest. For tricky production tasks, he recommends monitoring outputs with another LLM. That is his proposed semantic evaluation method; IFScale’s exact-word requirements show why mechanical obligations can instead be checked by counting.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Capacity increased; verification still has work to do
The later reliability discussion introduces Revisiting the Reliability of Language Models in Instruction-following, a paper testing forty-six models. Voss describes sensitivity to rewording and instruction order: the same intended requirements can produce substantially different compliance when their presentation changes. High capacity on one formulation does not guarantee stable behavior across equivalent prompts. A universally correct ordering remains an open research question.
The research direction also moves toward real, messy constraints. Voss names FireBench, CCR Bench and GuideBench as efforts to measure how well models follow many such requirements together. They address the gap between a report containing random words and an application balancing meaningful obligations.
The ending separates two engineering jobs. Compression fits requirements into usable model capacity. Verification checks whether the answer satisfies them. Voss’s declaration that the compression problem is gone expresses the scale of improvement in the tested task. The new hard part is knowing whether the model did what was requested.
Revisit prompt-size assumptions made six months earlier, then test returned outputs against the requirements that matter. A better prompt requests complete work. An eval determines whether complete work arrived. For readers who want to investigate the experiment itself, Voss closes with an invitation to obtain its code and data at the repository shown in the recording around 21:44.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Further reading
Develops the ending’s verification problem into practical evaluation design, including code graders, model judges, human calibration and misleading benchmark failures.
Shows how to capture application inputs, outputs and execution spans so incomplete or expensive agent responses can be investigated.
Related talks
- Ship Real Agents: Hands-On Evals for Agentic Applications
Voss’s hands-on continuation teaches tracing, deterministic checks, custom LLM judges, calibration and experiments for improving an agent.
- Don't Build Agents, Build Skills Instead
Explains reusable skill packages and selective loading, providing the architectural context for this recording’s instruction-capacity question.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
All right. Hello everybody.
- 0:15
Thank you for coming to this
- 0:17
delightfully nerdy talk. Uh this talk
- 0:20
has a really long title. Uh so let me
- 0:22
give you the short version up front. You
- 0:24
write skills files uh and stuff them
- 0:26
full of instructions. At some point the
- 0:28
model stops keeping track of all of
- 0:30
them. The question is where is that
- 0:32
point? At what point have you put too
- 0:34
many instructions in your skills files?
- 0:36
Uh and the answer has changed a lot in
- 0:39
the last year. I'm Lori. I'm head of
- 0:42
developer relations at Arise AI. Uh in a
- 0:44
former life, I co-founded npm Inc. So
- 0:46
some of you may know me from the days of
- 0:47
JavaScript. These days I spend a lot of
- 0:49
time thinking about AI and how to test
- 0:51
it.
- 0:53
Uh a few months ago I was at AI engineer
- 0:55
in Miami which was a good conference. Uh
- 0:57
and I was watching a talk by Dexter
- 0:58
Horthy. Uh it was a good talk. It was
- 1:01
not about this topic at all. Uh but
- 1:03
while he was giving that talk he
- 1:04
mentioned as an aside uh that an agent
- 1:07
can follow up to about 200 instructions
- 1:10
uh before it starts forgetting those
- 1:12
instructions. Uh and he then he moved on
- 1:15
in his talk and it was entirely an
- 1:16
aside. Uh and he he mentioned that that
- 1:19
figure is from 2025 so things might be
- 1:21
better now. Um, and I stopped listening
- 1:24
for a second because I was like, 200
- 1:26
instructions. Uh, is not very many
- 1:30
instructions at all. Right? A decent
- 1:31
skills file blows past 200 instructions
- 1:33
almost immediately. Um, if the user says
- 1:36
X, do Y, always include a section on Z,
- 1:39
never use the phrase W, every one of
- 1:41
those is a separate instruction. Uh, and
- 1:43
if the model quietly stops tracking them
- 1:44
after 200, that's a really hard ceiling
- 1:47
on the complexity of what you can build.
- 1:49
Uh, so I wanted to know where he got
- 1:51
that number first. Uh, and I wanted to
- 1:53
know if it was true. Uh, so you know the
- 1:56
feeling that I'm talking about. You
- 1:57
write this big beautiful skills file,
- 1:58
pages of rules, edge cases, tone,
- 2:00
formatting, you hand it to the agent, it
- 2:02
does the thing. Uh, and you look at the
- 2:04
output and go, did it actually pay
- 2:06
attention? Did it actually follow all of
- 2:08
these rules or did it just sort of, you
- 2:10
know, do what it felt like and sort of
- 2:12
give me a close simulacum of what I was
- 2:14
expecting? Um, you can't really tell or
- 2:18
can you? More on that later. Um, and so
- 2:21
you live with this lowgrade anxiety
- 2:23
every time you hit run. Um, and that
- 2:25
feeling is what this research is about
- 2:27
and what we're trying to find out if we
- 2:29
can avoid.
- 2:30
So here's my promise for your next 18
- 2:32
minutes. Uh, I'm going to show you where
- 2:34
that 200 number came from, whether it's
- 2:36
still true, and what the real number is
- 2:38
today. Uh, because it moved by an order
- 2:40
of magnitude. Uh, and then we're going
- 2:42
to talk about what that means for you to
- 2:45
take away. uh how long your skills and
- 2:47
prompts can actually be and what that
- 2:50
should what changes you should make to
- 2:52
your workflow as a result.
- 2:54
So the 200 number isn't folklore. Uh it
- 2:57
comes from a real benchmark called
- 2:58
IFScale uh from a paper uh by this guy
- 3:02
whose name I'm going to mess up
- 3:03
Jeroslowitch uh and co-authors last
- 3:06
year. And the test is beautifully
- 3:08
simple. Uh here's how if scale works. If
- 3:12
you ask the model to write a business
- 3:13
report uh and you give it a list of
- 3:15
specific words that it has to include
- 3:17
exactly in the report, include the exact
- 3:20
word customer, include the exact word
- 3:21
revenue, and so on for as many words as
- 3:23
you want. Each of those is an
- 3:25
instruction that it has to follow. Uh
- 3:27
and then you count how many of those
- 3:29
exact words showed up. Um
- 3:32
so because the test is so simple, you
- 3:34
only have to keep two numbers in your
- 3:36
head. One is density, which we call n.
- 3:38
That is how many rules we're talking
- 3:40
about at once. And the second is
- 3:41
accuracy, which is the percentage of
- 3:43
those rules uh that it was able to
- 3:45
actually follow. Uh now you might say uh
- 3:49
that including random words in a report
- 3:50
is not the same as following real
- 3:52
instructions and fair enough and we're
- 3:54
going to talk about that. Um but the
- 3:56
keywords are a proxy. Uh include the
- 3:59
word revenue is the same shape of task
- 4:01
as include a section on pricing, right?
- 4:03
Or never use this phrase. It is a
- 4:04
discrete named constraint that you've
- 4:06
told the agent that it has to follow.
- 4:08
Um, if a model can't track 200 words in
- 4:11
one prompt, it's definitely going to
- 4:12
struggle with 200 more complicated
- 4:14
instructions. Uh, so uh, if anything,
- 4:19
it's going to do worse. So, this number
- 4:20
is a ceiling. Uh, this number is as high
- 4:23
as you can go. If you give it more
- 4:24
complicated instructions, the number is
- 4:26
probably going to get lower. And 200 is
- 4:28
a really low ceiling. Um so before
- 4:31
chasing new models you have to do good
- 4:33
science which means that you have to
- 4:34
replicate the uh old result and make
- 4:36
sure uh that the 200 ceiling is real. So
- 4:40
I reran the original benchmark. Um the
- 4:43
original paper tested a whole batch of
- 4:44
models uh and models uh live and die
- 4:47
really fast. So uh by the time I got
- 4:49
around to doing this testing only three
- 4:51
of the models in the original set of 10
- 4:53
models that they used were still
- 4:54
available via any kind of API. Uh so
- 4:57
they were GPT 4.1, Claude Sonnet 4, and
- 4:59
Gemini 2.5 Pro. Those were models that
- 5:02
were available 12 months ago that are
- 5:03
still available now. Um and that is why
- 5:06
we tested those three because they were
- 5:07
what was left. Um and since I first
- 5:10
published this research a couple of
- 5:12
weeks ago, uh one of those three models
- 5:14
has been retired. So this was the last
- 5:15
possible time that I could have run this
- 5:17
test. Um so of that lineup, we're
- 5:20
already down to two. So don't get
- 5:21
attached to your models. Um here is the
- 5:23
results that we got replicating the
- 5:25
original if scale finding. Uh that is
- 5:27
accuracy on the vertical axis. So it
- 5:30
starts at 100% and begins to fall off.
- 5:32
Uh and then the number of rules uh going
- 5:34
up along the bottom on log scale. So
- 5:36
every time it gets halfway across it has
- 5:38
doubled uh the number of rules that it's
- 5:40
dealing with. Um so by 500 rules you're
- 5:44
losing 30 40 50% of them. Uh our curves
- 5:47
matched the results in the original
- 5:49
paper within the noise boundary. So the
- 5:50
finding was real. uh a year ago
- 5:53
somewhere around 200 to 300 rules
- 5:55
frontier models started falling apart.
- 5:57
That is a really low ceiling. Uh so that
- 6:01
is our baseline and now comes the fun
- 6:03
part where we took the exact same test
- 6:04
and pointed it at the current frontier
- 6:07
or rather what the current frontier was
- 6:09
when I ran this test. So I ran GPT 5.5,
- 6:12
Claude Opus 4.7 because 4.8 came out a
- 6:15
week after I ran this test. Uh Gemini
- 6:17
3.1 Pro and Deepseek V4 Pro. So, I gave
- 6:20
them the same prompt, the same words,
- 6:22
the same everything. And I immediately
- 6:24
ran into a problem, which is that they
- 6:26
aced it. They all scored 100%
- 6:29
immediately on this test. Absolutely no
- 6:31
bugs. Uh,
- 6:34
so we'd built a test to find the ceiling
- 6:36
and the models had walked straight
- 6:37
through the ceiling without noticing
- 6:38
that the ceiling was there. Um, and that
- 6:40
was a problem because the benchmark was
- 6:42
written to top out at 500 words. So, I
- 6:44
had to change the benchmark in order to
- 6:45
be able to find the new ceiling. So, I
- 6:47
moved the goalposts. I gave it more
- 6:49
words to include. I doubled uh it from
- 6:51
500 to a,000. I doubled it again from
- 6:53
a,000 to 2,000. And I kept doing that
- 6:55
until I hit a 10,000word vocabulary. And
- 6:58
that is where I began to find the
- 6:59
ceiling of what models can do these
- 7:02
days. Um, so let me put up the this is
- 7:05
the money slide. This is the results.
- 7:07
Remember log scale on the on the uh
- 7:11
x-axis there. So it's going from 500 to
- 7:13
1,000 to 5,000 to 10,000. Uh so it looks
- 7:16
like that scale is falling off of a
- 7:18
cliff and it's actually happening over
- 7:19
like a thousand numbers. Um
- 7:22
but uh look how far to the right these
- 7:25
new curves get before they bend. A year
- 7:26
ago they were falling over at 200 to 300
- 7:29
instructions and now depending on the
- 7:30
model the boundary is closer to 2,000.
- 7:32
And for the best of them it is up to
- 7:34
5,000 instructions before they begin to
- 7:37
fall off a cliff. So in about 12 months
- 7:39
frontier models got close to 10 times
- 7:41
better at following instructions
- 7:43
simultaneously. That is the headline
- 7:45
fighting and there is a lot of nuance
- 7:47
that we need to get into. Um the
- 7:50
capacity to track 2,000 named
- 7:51
constraints in a single prompt is there.
- 7:54
Um and that's really interesting because
- 7:56
I think uh I don't know if everybody
- 7:58
else feels this way but like it feel it
- 8:01
felt to me like the the jump from you
- 8:04
know GPT 5.1 to GPT 5.5 was kind of
- 8:06
incremental, right? It didn't feel like
- 8:08
we'd got 10 times better. But this is a
- 8:11
test that really matters to uh a very
- 8:14
practical thing like how long can my
- 8:16
skills file be? Uh and in the course of
- 8:18
a year we got 10 times better. Uh and
- 8:22
that thing that gets me is that this
- 8:23
benchmark is barely a a year old. A year
- 8:25
later 500 is a rounding error. Uh and
- 8:28
this keeps moving under my feet. I
- 8:29
tested 4.7 uh opus 4.7. Opus 4.8 is even
- 8:33
better. Um so this chart is a little out
- 8:36
of date already which is kind of the
- 8:37
whole point. If you set your engineering
- 8:39
assumptions about how skills files
- 8:41
should work, about how prompt how long
- 8:43
your prompt can be, and you did that
- 8:45
more than about six months ago, you are
- 8:47
incorrect now, and you should probably
- 8:49
be re-engineering how you do stuff. Uh,
- 8:52
but there is more to this story uh
- 8:54
because the way that the a models failed
- 8:58
uh changed dramatically. Uh, and the way
- 9:01
that they failed is very important. This
- 9:03
part was a completely unexpected finding
- 9:06
when I started running the experiment.
- 9:08
Uh, and it totally messed up my test to
- 9:09
start with because uh, the old failure
- 9:12
mode was boring. They would just forget
- 9:14
instructions and I could measure how
- 9:15
many instructions they had remembered or
- 9:17
forgotten. Uh, but the new ones fall
- 9:19
apart in their own weird extremely
- 9:21
onbrand way. Uh, so let me introduce you
- 9:24
to how these four models fail. Uh,
- 9:27
Deepseek 4 is a traditional model. It
- 9:29
just forgets things. It doesn't have any
- 9:31
drama. um it starts forgetting
- 9:34
instructions around 750 rules and by
- 9:36
2000 it's dropping nearly half of them.
- 9:38
Uh so it just forgets which frankly is
- 9:40
the failure mode that I trust most
- 9:42
because it's predictable. It's very easy
- 9:43
to measure. Uh and the other models were
- 9:46
not nearly as cooperative. Uh Opus 4.7
- 9:51
uh would decide repeatedly that the test
- 9:53
was dangerous. Uh and what it would do
- 9:55
is it would refuse at the API level to
- 9:58
complete the test. I didn't know that
- 10:00
there was an API response that you could
- 10:01
get from Claude where it was like, "No,
- 10:03
I could do this, but I'm not going to."
- 10:06
Uh, but that's absolutely an API level
- 10:09
response that Claude supports because
- 10:10
they care so much about safety. Uh, and
- 10:12
I started getting those all of the time.
- 10:15
Uh, and the reason that was happening is
- 10:16
because Claude has a very sensitive
- 10:18
safety classifier. Uh, and if you put in
- 10:20
certain combinations of words like say
- 10:22
anthrax and cyanide, it decides that the
- 10:24
whole request is dangerous and it bails
- 10:26
out. Uh, and if you remember what my
- 10:28
test does, my test is throwing uh 5 to
- 10:31
10,000 random words into uh into an
- 10:34
instruction file. And so my my randomly
- 10:37
selected words contained all sorts of
- 10:38
things that looked dangerous in
- 10:40
combination to the safety filter. And so
- 10:41
it kept bailing saying that I was asking
- 10:43
it to, you know, make a bomb or
- 10:44
something. Um,
- 10:47
so, uh, we had to for to get Claude to
- 10:52
cooperate, I had to take all of my words
- 10:54
and run them through OpenAI safety
- 10:55
filter and filter out all of the naughty
- 10:57
looking words so that it could get to
- 10:58
anywhere. Once I given it that, Claude
- 11:01
did really well. Uh so but the failure
- 11:03
mode is that Claude is more likely to
- 11:05
decide what you're doing is dangerous
- 11:07
very early on uh at you know even two or
- 11:10
300 instructions if what you're doing uh
- 11:13
is you know contains anything to do with
- 11:15
medical advice because medical things
- 11:16
often are dual purpose. They can be
- 11:17
dangerous. They can be safe. Um so uh
- 11:21
the third failure mode was Gemini 3.1
- 11:23
Pro. Gemini is rock solid all the way
- 11:26
out uh to 5,000 instructions. It does
- 11:28
extremely well. um genuinely one of the
- 11:31
best on the chart. Uh and then past that
- 11:33
it gets weird. Um it doesn't forget the
- 11:36
instructions, it gets overwhelmed by the
- 11:39
instructions. What it tries to do is it
- 11:42
uh it uses thinking tokens to make sure
- 11:44
that it is following all of the
- 11:45
instructions at once. And when the
- 11:47
number of instructions gets really high,
- 11:48
it uses all of its thinking tokens. It
- 11:50
uses its entire token budget thinking.
- 11:53
And then it doesn't give any output.
- 11:55
It's it gets to like nine, you know, if
- 11:57
you've given it 10,000 tokens worth,
- 11:59
it'll get 9,500 tokens worth of thinking
- 12:02
and then give you a 500word response
- 12:04
which doesn't contain any of the tokens.
- 12:06
Uh so it thinks itself into a corner and
- 12:09
runs out of room to actually answer,
- 12:10
which is very expensive, uh and totally
- 12:13
unhelpful, which is kind of on brand,
- 12:15
isn't it? Um [snorts]
- 12:18
uh
- 12:19
which, you know, I would never say that
- 12:20
out loud. Uh, and finally comes the
- 12:24
winner, which is uh, GPT 5.5. GPT 5.5 is
- 12:27
the best of the lot. 99% accuracy all
- 12:29
the way out to 5,000 rules. Um, but if
- 12:32
you push it far enough, it is by far the
- 12:34
weirdest of the bunch. Uh, because it
- 12:36
doesn't refuse outright. It doesn't
- 12:38
silently forget. Instead, what it does
- 12:39
is it gets frustrated and tells you that
- 12:42
the test is stupid.
- 12:44
Uh, it starts the report. It gets a few
- 12:47
like that's the thing. It doesn't start
- 12:49
out just saying no. It starts the
- 12:51
report, it starts writing the report,
- 12:52
and like 500 words into the report, it's
- 12:54
like, "No, this is dumb. I'm not going
- 12:56
to do this." And then it politely tells
- 12:57
you, "This is dumb. I'm not going to do
- 12:59
this anymore." Uh, that is the actual
- 13:01
response that it gave me, but that that
- 13:02
was like 5,000 words into the into this
- 13:04
business report that I told it to
- 13:06
generate. Um, so it's not wrong, right?
- 13:10
I was asking for a coherent business
- 13:12
report that on no particular subject
- 13:14
that contains 5,000 random words. You're
- 13:17
right, Gemini DPT. this is a a stupid
- 13:20
thing to ask for. Um,
- 13:23
which is a deeply unreasonable request
- 13:25
and GPT called this out on it. Um, but
- 13:27
it still counts as a failure in the test
- 13:29
because the half-finish report that it
- 13:30
gives you is missing most of the
- 13:32
keywords and it is also the hardest one
- 13:34
to detect because claude bails
- 13:36
immediately. Claude says, "No, I'm not
- 13:38
going to do this." Uh, Deepseek does its
- 13:40
best. Uh, but GPT does what looks like a
- 13:44
good job unless you read all the way to
- 13:46
the end of the report where it says,
- 13:47
"No, actually I'm going to bail because
- 13:48
this is stupid." Um,
- 13:51
so if you step back and look at the four
- 13:53
together, Deep Sea quietly forgets,
- 13:54
Claude gets scared and refuses, Gemini
- 13:56
overthinks itself into silence, and GPT
- 13:59
5.5 finishes half of the job and tells
- 14:01
you that the rest of it is beneath it.
- 14:03
Um, and the point was the point isn't
- 14:06
which one of these is funniest, although
- 14:08
it is genuinely a little funny. Uh the
- 14:10
point is that did it follow my
- 14:12
instructions no longer has one failure
- 14:14
mode. It has four different ways that it
- 14:16
can fail and you can't recognize that
- 14:18
failure unless you know which model
- 14:20
you're dealing with and what its m what
- 14:22
its pattern of failure is going to be.
- 14:24
Uh so the models get 10 got 10x better.
- 14:27
They fail in funny ways. Why should you
- 14:29
care when you uh get back to your desk?
- 14:32
Because three things have changed to
- 14:33
your workflow. The first is that a year
- 14:36
ago, the smart move was to keep every
- 14:38
skills file very very short. Uh under
- 14:41
200 instructions, then point off to
- 14:42
subsklls and a whole like you know
- 14:45
byzantine labyrinth of uh additional
- 14:48
skills files and subfiles and things
- 14:50
like that. Uh and you mo you were
- 14:52
compressing your your instructions to
- 14:54
fit into a very small available space
- 14:56
and you don't need to do that anymore.
- 14:58
Your skills files can be very long. Um,
- 15:01
number two is that if your use case
- 15:03
needs a 100 specific rules or 300, you
- 15:05
can just put them all in the prompt. Uh,
- 15:09
you don't have to lie awake wondering
- 15:10
whether which ones the model silently
- 15:12
ignored. Um, and if you've been thinking
- 15:15
uh about uh your own lived experience of
- 15:18
using models, uh you probably recognize
- 15:21
this. you've discovered that you've got
- 15:22
less worried about how long your your
- 15:24
prompt is going to get uh because the
- 15:26
models have genuinely got 10 times
- 15:28
better at following your prompts. Um
- 15:32
2,000 named constraints is an entire
- 15:34
style guide, right? Like it's it's every
- 15:36
brand rule, every legal disclaimer. Uh a
- 15:38
year ago, you'd have had to shard that
- 15:40
across a dozen specialized agents and
- 15:42
hope that your specialized agents are
- 15:43
hand are are handing off to each each
- 15:45
other cleanly. But now you can ignore
- 15:48
that. Um but the third thing is the big
- 15:50
one. The question used to be can the
- 15:52
model even do this? And the answer is
- 15:54
now firmly yes. Well reasonably firmly.
- 15:57
Uh is it worth the cost is the new
- 16:00
question because you can include 10,000
- 16:03
words of of sorry 10,000 different
- 16:05
instructions into your prompt. But that
- 16:06
is going to be an enormous prompt. It's
- 16:08
going to be a very expensive prompt.
- 16:09
It's going to be a very slow prompt. So
- 16:11
what used to be a hard wall that you
- 16:12
would run against has now become a soft
- 16:14
trade-off of is it worth me adding all
- 16:16
of these extra instructions if it's
- 16:18
going to give me more cost and more
- 16:19
latency.
- 16:21
Uh and now some caveats uh to head off
- 16:25
the Q&A. Um first and important first
- 16:28
and most important I mentioned this
- 16:29
earlier this is a proxy task including
- 16:31
random words uh in a in a fake business
- 16:34
report um is evidence that long skills
- 16:37
file works. It is not the same as proof
- 16:39
that a long skills file works. Um, also
- 16:43
the models hit the wall at wildly
- 16:44
different points anywhere from 750 to
- 16:46
9,000 plus. So you have to pick your
- 16:48
model very carefully. Uh, what our test
- 16:53
doesn't do is measure whether the model
- 16:55
reasoned clearly over a giant prompt. So
- 16:59
uh, the good news is since I did my
- 17:00
research several weeks ago, uh, a whole
- 17:02
bunch of people have piled in on this.
- 17:04
Um and now there's good research uh
- 17:07
actual scientists have got involved and
- 17:09
done uh Chroma's has done context rot
- 17:12
work uh across 18 models showing that
- 17:15
accuracy on long inputs can fall 30 to
- 17:18
50% well before you hit the context
- 17:20
window limit. Uh and the weird part of
- 17:23
their finding was that uh coherent well
- 17:26
ststructured text is more likely to hit
- 17:28
that failure mode uh than if you just
- 17:30
put your instructions into a random
- 17:31
order and shuffle them in. Uh, I don't
- 17:35
know why that's the case. I'd have to
- 17:36
read their report. Um, so the model can
- 17:40
track 2,000, 5,000, possibly 10,000
- 17:42
instructions, but it's not necessarily
- 17:44
going to uh reason clearly over them.
- 17:47
It's not necessarily if those if those
- 17:49
instructions conflict, if there is
- 17:50
tension between them, it's not
- 17:52
necessarily going to get that right. Um,
- 17:55
and then there's the other one I
- 17:56
mentioned briefly. Uh, collude's
- 17:59
refusals are annoying, but they are
- 18:00
loud. You get an error, you know it
- 18:02
failed. Uh GPT's polite half-finish
- 18:04
report is much more dangerous because it
- 18:06
looks like a real answer. Uh you have to
- 18:08
read the whole thing to notice that it
- 18:09
gave up quietly halfway, which means
- 18:11
that you can't trust the output. It
- 18:14
means you have to read the output every
- 18:16
single time to make sure whether or not
- 18:17
it's working. Uh so the model will
- 18:20
accept your 2,00 rules and it will hand
- 18:22
you back something that looks at least
- 18:23
to begin with confident and polished but
- 18:25
could be bailing out halfway through.
- 18:28
Um,
- 18:30
so, uh, as an aside, people always ask
- 18:33
me, "How much did all this cost me?" It
- 18:34
cost me $29 to run all of these queries.
- 18:37
2,37
- 18:39
2,300 calls across seven models, uh,
- 18:42
came to $29. Uh, it turns out novel
- 18:44
research doesn't cost very much. Um,
- 18:48
and this is the part of the talk where I
- 18:50
was saying that you have to check this
- 18:51
stuff in production because you can't
- 18:53
trust that your model isn't going to
- 18:55
silently fail. Uh, so you knew I was
- 18:58
going to mention evals eventually
- 18:59
because I work at Arise and this is
- 19:00
where I do that. Um, but there are
- 19:02
plenty of plugs for Arise. So I'm just
- 19:04
going to say one true thing which is
- 19:06
that if you are building a real AI
- 19:07
application and you are giving it
- 19:09
genuinely tricky tasks, you are going to
- 19:11
run into one or more of these failure
- 19:12
modes with a frontier model. Uh, and
- 19:15
unless it's Claude telling you just to
- 19:16
off at the API level, the only way
- 19:19
to know that something went wrong is
- 19:21
monitoring your outputs with another
- 19:22
LLM. That is an eval. And that is what
- 19:24
Arise does. And I'll leave it at that.
- 19:27
Uh, I already mentioned that there's
- 19:29
been new research since we did our own.
- 19:31
Here's another important one. A paper
- 19:32
landed testing 46 models called
- 19:34
revisiting the reliability of language
- 19:36
models in instruction falling, which you
- 19:38
can bet made my ears perk up after I did
- 19:40
that research myself. Uh, and they found
- 19:42
something uncomfortable, which is that a
- 19:44
model can ace a benchmark like ours and
- 19:46
still be wildly unreliable. because if
- 19:48
you reword the same instruction in a
- 19:51
slightly different way, it can make a
- 19:52
radical difference to how well uh it
- 19:55
follows those instructions. So the model
- 19:57
can follow 2,000 instructions and it can
- 19:59
do it really well. But if you put the
- 20:01
same instructions, the same 2,000
- 20:03
instructions in a different order, it
- 20:05
can suddenly make the model much worse
- 20:07
at following those instructions. And how
- 20:09
ex how exactly to do that? what is the
- 20:12
correct order of instructions to give
- 20:14
your model such that it follows them
- 20:15
perfectly as opposed to getting confused
- 20:17
is still research that is being done. So
- 20:20
capacity went up but reliability is
- 20:23
still a problem. Um and then this is
- 20:26
just a little brag because uh I was
- 20:29
happy about it like I'm not a scientist.
- 20:30
I did some research and then a whole
- 20:32
bunch of other actual scientists piled
- 20:34
in uh and did real science on the same
- 20:36
question. There's now a whole bunch of
- 20:37
benchmarks that have shown up uh to
- 20:39
measure this same question. Firebench,
- 20:41
CCR bench, Guidebench uh are all trying
- 20:43
to measure the same thing. How well
- 20:45
models follow a lot of real messy
- 20:47
constraints at once. Uh and now the
- 20:49
whole field is looking at it. So if you
- 20:51
want better science than my, you know,
- 20:53
10,000 random words, uh the real science
- 20:56
exists now. Uh so that gets me to where
- 21:00
I will leave you. A year ago, the hard
- 21:02
part of writing a skill was fitting
- 21:03
everything in without the model losing
- 21:04
the plot. That was a compression
- 21:06
problem, and the compression problem is
- 21:08
gone. uh the model will hold your 2,000
- 21:10
instructions just fine. The new hard
- 21:12
part is knowing whether it actually did
- 21:14
what you said and that is a verification
- 21:16
problem. Uh a verification problem
- 21:18
doesn't get solved by writing a better
- 21:20
prompt. It gets solved by checking the
- 21:21
output every time uh the same way that
- 21:24
you would test any other code, which is
- 21:25
to say an eval. The ceiling moved by 10x
- 21:28
in one year. Uh so go back and check the
- 21:31
assumptions that you made six months ago
- 21:33
about how big your prompts should be,
- 21:35
how big your uh instructions can get. uh
- 21:38
because they might already be wrong.
- 21:39
Boom. Be wrong. So that is the talk. If
- 21:42
you want uh all of the code and all of
- 21:44
the data, uh it is at this GitHub URL.
- 21:47
Uh and this other QR code is uh
- 21:50
something marketing made me insert. We
- 21:52
are having a World Cup watch party
- 21:54
tonight at 5:00 p.m. Uh you can come to
- 21:56
our party. That link is to the Luma that
- 21:58
will get you into the get into get you
- 22:00
into the party. Uh I hope this talk has
- 22:03
given you some novel information or at
- 22:05
least a couple of laughs. And thank you
- 22:06
so much for your time and attention.