Training Taste — Thais Castello Branco, Taste Labs
Read the talk
Training Taste
Thais Castello Branco’s approach to better AI design starts by decomposing “taste” into measurable constraints, contextual judgments, and inference-time systems that can resist repetitive, ill-fitting output.
From a talk by Thais Castello Branco
At a glance
Ideas worth remembering
Decompose subjective work before choosing an improvement method: explicit constraints can move toward deterministic verification, while aesthetics and preference still require data and human judgment.
Slop has three connected signatures: repetition, poor fit to context, and weak interpretation of user intent.
Small classifiers can each detect one design characteristic; combining their frequency turns a holistic impression into a structured prediction problem.
Useful creative variation is deliberate: preserve enough category expectations for the artifact to fit its purpose, then break selected rules rather than merely adding randomness.
Inference-time systems matter because that is where an application can recover intent, inject brand context, steer generation, and verify the result.
The near-term target is a higher quality floor, not the pinnacle of human craft: make generic output measurable, locatable, and correctable first.
A fuzzy domain creates two different engineering problems
Taste Labs founder Thais Castello Branco opens with a blunt mission: end AI slop. Design is the first test case for a larger problem—making models and agents better at subjective work such as design and writing, where correctness alone cannot define quality. The first move is decomposition: turn “good design” into smaller characteristics, then decide which improvement method fits each one. 0:13
Some criteria become close to deterministic once the task and context are narrow enough. Contrast, alignment, and palette selection can often be specified so that most experts agree. Aesthetics behaves differently: legitimate expert disagreement remains, so training has to rely more heavily on preference data rather than pretending there is one universal answer. This is less a split between objective and subjective domains than a routing decision inside each design task. 1:13
That routing spans two layers:
- Model layer: Evaluate failures, identify missing capabilities, and construct post-training data or reinforcement-learning environments aimed at those failures.
- Application layer: Supply the context, preferences, judgment, and verification that an off-the-shelf model lacks at the moment it serves a user.
A model that tends to collapse toward a familiar style may need better training, but a model cannot infer a company’s brand or a user’s unstated intent from nothing. Some failures therefore belong in the surrounding product rather than in the weights. 0:43
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Generation is cheap; judgment took a career to build
Math offers a clean target: a great answer can simply be the correct answer. A poem, coffee shop, website, or piece of art can feel special for several entangled reasons—uniqueness, craft, attention to detail, or authenticity. In creative work, the most likely answer may be precisely the wrong target. A memorable result often sits away from the average and calls attention to itself without feeling arbitrary. 2:25
AI has nevertheless achieved something remarkable: someone without design or engineering training can generate a slide deck, website, or web application with one action. As the cost of producing an artifact approaches zero, however, the cost of judging it does not. Designers build judgment through years of exposure, pattern recognition, restraint, and developing the confidence to depart from convention on purpose. 3:25
“Everyone should acquire taste” is therefore not a workable product strategy. Most people cannot spend a career developing expertise in every medium they may now generate. The practical challenge is to put enough design knowledge into models and application systems that a nonexpert can create something good—and articulate what they prefer—without first becoming a designer. 4:25
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Three signatures make slop easier to recognize than greatness
Slop predates generative AI; social media already rewarded rapid production and repetition. AI accelerates the phenomenon by making one-shot generation easy. Castello Branco identifies three recurring signatures:
- Repetition: The same palettes, structures, and visual patterns recur across outputs.
- Lack of fit: The result ignores who it is for, when it appears, and what situation it serves.
- Low intent: A shallow request passes through without the system helping the user clarify what they actually want.
These mechanisms reinforce one another. Repeated patterns become slop when they are reused without regard for context, while weak intent leaves the system little reason to select anything more fitting. 4:55
Consider two users asking for websites: one runs a pet shop and the other a finance firm. If both receive nearly the same layout and visual language, the observable failure is not merely sameness. The shared design has discarded category expectations, audience, tone, and purpose. A carefully made version might reuse sound fundamentals, but it would interpret those fundamentals differently for each business. The causal chain is straightforward: thin prompts provide little intent; the model falls back to frequent patterns; both outputs converge; and the resulting design fits neither context particularly well. 5:25
That diagnosis turns an aesthetic complaint into an evaluation agenda. Instead of asking one model for a holistic verdict on whether a page is “great,” the system can inspect repetition, fit, intent, and smaller design features separately. Measurement matters because a failure that can be located and described can become a training target, a runtime check, or a product intervention. 6:12
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From two million websites to small slop detectors
Taste Labs analyzed more than two million websites spanning roughly ten years and also generated a synthetic set for comparison with human-made sites. The historical data complicated the simple story that AI created visual homogenization: similar palettes and layouts were already spreading before widespread AI generation. Castello Branco suggests faster trend propagation as a possible cause. AI then intensified repetition across categories, making similar patterns appear even where the surrounding contexts differed. 6:41
The measurement pipeline begins by mining the collection for structured characteristics: colors, typography, layout, audience, and related features. Each feature can then support a small classifier—a “probe”—trained to recognize one characteristic. No individual probe has to define slop in full. The useful signal comes from combinations: when several slop-associated characteristics occur together and recur frequently, the system can estimate that a site is likely to belong to the slop category. 7:42
What question does the probe stack answer? It shows how a vague overall impression can be reconstructed from smaller observable signals rather than assigned in one holistic judgment.
Castello Branco reports that this combined-probe approach predicted slop very well and outperformed common LLM-as-a-judge methods that directly ask a model to distinguish high-quality human work from AI slop. The talk does not provide the dataset composition, held-out evaluation design, metric, score, or judge baselines, so the result establishes the team’s reported direction rather than a reproducible performance comparison. Its technical lesson is still useful: a collection of narrow detectors can expose recurring structure that a single broad aesthetic prompt misses. 8:11
Measurement is only the intermediate goal. When generation becomes abundant, judgment becomes the scarce resource: the ability to discern what is right, break a fuzzy problem into tractable pieces, and choose the appropriate intervention. Human judgment remains valuable, but tools can encode and apply parts of it repeatedly. 8:42
More than two million historical sites plus a synthetic comparison set.
The design is decomposed into structured features. Small classifiers detect individual characteristics, and their co-occurrence supplies a combined prediction.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Inference time is where context meets generation
Improving the base model is necessary, but it does not eliminate the inference-time problem. The application meets the user at inference time, where it can ask questions, interpret intent, retrieve context, enforce preferences, and verify an output. If that exchange remains shallow, a better model can still produce a polished answer to the wrong problem. Castello Branco therefore treats inference-time design as at least as important as model improvement for fighting slop. 9:27
For repetition, Taste Labs is developing what Castello Branco provisionally calls a “creativity API”: an inspiration system intended to push an agent out of its usual distribution. The important constraint is that creativity cannot be reduced to increasing sampling temperature. Randomness may produce novelty, but it can also produce work that simply feels wrong for the assignment. 9:57
A startup pitch deck makes the distinction concrete. The system should first understand what a competent pitch deck normally requires. It can then diverge deliberately on a few dimensions while preserving the expectations that keep the artifact legible as a pitch deck. Creativity comes from choosing which rule to break and why, not from discarding every category convention at once. 10:57
For fit, an existing brand supplies a valuable body of prior judgment. Dozens of design choices have already established which colors, typography, spacing, and visual details belong together for that company. Reusing that work gives an agent a context-specific target instead of asking it to invent “good design” from scratch. The probe approach can serve a complementary role as a gate: detect undesirable patterns before an agent ships them. 11:27
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn the brand into something an agent can follow—and be graded against
The first public product described is a brand API, then in beta testing with design partners. It takes a brand URL and extracts specific components suitable for an agent to follow. Structuring the brand serves two directions at once: the components guide generation, and the same components provide criteria for checking whether the result stayed on brand. Generation without this verification path would leave the system unable to say where or how it drifted. 11:57
What does this flow make visible? A brand specification is not only context injected before generation. It becomes the shared reference connecting input, output, and evaluation.
Users without an established brand need a different starting point. Taste Labs is also creating an index of prebuilt brand systems. If someone requests something “dreamy,” the application could retrieve a cohesive system already designed around that quality rather than generating an improvised palette, type hierarchy, and layout independently in one pass. Retrieval preserves a set of choices that were designed to work together. 12:58
The closing demonstration applies this mechanism to the General Intelligence Company of New York. The original brand provides the reference. A default request for a branded slide deck produces a middle result that departs from it; adding the extracted brand components produces a deck described as much more faithful, including small details that “feel right.” The observable change is higher resemblance to the original visual system. The causal path is extraction, structured guidance during generation, and comparison against the same brand target. 13:28
The immediate goal is deliberately modest. Human craft remains valuable, but the current question is not whether a model can reach its pinnacle. It is whether decomposition, measurement, context, and verification can lift the quality floor. Before debating machine taste at its highest level, the system has to stop shipping the same ill-fitting page to the pet shop and the finance firm. 14:04
An existing brand supplies the source material.
Extracted brand components guide the agent and provide the criteria used to inspect its output.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Ending AI Slop
Castello Branco’s companion talk goes deeper on routing design properties toward verification, reinforcement-learning environments, or expert preference data.
- Perceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.ai
A complementary treatment of why common proxy metrics struggle to capture human aesthetic perception.
- The Missing Layer: Design Taste in AI Agents
Extends the practical side of the problem with design guidance and visual references for reducing generic agent-generated interfaces.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Test.
- 0:13
Okay, amazing.
- 0:15
It's great to meet everyone. I'm Taís.
- 0:16
I'm the founder of Taste Labs. Uh for
- 0:19
those of you who don't know us, we came
- 0:20
out of stealth a few weeks ago and our
- 0:22
whole mission is basically how do we end
- 0:24
AI slop? I that's my personal enemy. Um
- 0:27
and so we really believe that to solve
- 0:30
this problem of slop, we have to like
- 0:32
decode subjective domains. Uh there's
- 0:34
been so much effort being put into
- 0:36
getting models and agents amazing at
- 0:38
things like coding and math. Uh and it's
- 0:40
time that we put all that same effort
- 0:41
into making them great at things like
- 0:43
design uh and writing. And so design is
- 0:45
this first pillar that we're starting
- 0:46
with and it's been it's been incredibly
- 0:48
exciting. Um
- 0:50
We work primarily in two ways. So we
- 0:52
work a lot with the frontier labs on how
- 0:54
do we evaluate their models, understand
- 0:56
where they're breaking, understand what
- 0:58
could be better about them, and then
- 0:59
construct the right either post-training
- 1:01
data or RL environments to basically fix
- 1:03
that problem. And part of this is like
- 1:05
how do you take something as fuzzy and
- 1:06
large as design and break it down to a
- 1:09
level that you can identify what is best
- 1:11
solved through each method. What are
- 1:13
elements of design that are almost like
- 1:15
once you kind of boil down the problem
- 1:16
becomes so specific that they almost
- 1:18
become deterministic. So for example, uh
- 1:20
if you're trying to train a model to be
- 1:21
good at selecting color palettes or have
- 1:23
contrast or alignment, those are things
- 1:26
that if you define the problem and the
- 1:27
context in a specific enough way, uh you
- 1:29
you can get to an answer that's like
- 1:31
pretty objective or that at least most
- 1:32
experts would agree to. But maybe other
- 1:34
things like uh aesthetics, you naturally
- 1:37
will see this expert disagreement. And
- 1:38
so then you want to lean on to things
- 1:40
that are closer to to data. So anyway,
- 1:42
we spend a lot of time thinking about
- 1:43
all those problems. Uh but on the other
- 1:44
side is also
- 1:46
without even touching the model layer,
- 1:47
right? How do we actually help agents
- 1:49
and app layer companies produce better
- 1:51
things? And there's a lot that goes into
- 1:53
that, right? You have these different
- 1:54
sets of problems at the application
- 1:55
layer because you're using an
- 1:56
off-the-shelf model that tends to
- 1:58
collapse in terms uh uh of style tends
- 2:00
to collapse to the mean. So, how do we
- 2:02
force that creativity back to the
- 2:03
system? How do we avoid these patterns
- 2:05
of slop, which we'll talk about a lot
- 2:07
today? Uh how do you understand like
- 2:09
user preferences or brand preferences
- 2:11
preference so that you can uh maintain
- 2:13
endurance to that style? Uh so, there's
- 2:15
lots of things that are actually need to
- 2:17
be solved as context or judgment or
- 2:19
verification at the app layer, which is
- 2:22
why we kind of work across both.
- 2:25
Maybe I'll start with more of a a
- 2:27
philosophical question of like how how
- 2:28
do you define something that is great?
- 2:30
Like how do you define greatness? And
- 2:32
for something like math, it's easier,
- 2:34
right? Because there's kind of one
- 2:36
objective answer, and uh great is the
- 2:38
same as correct. But then for something
- 2:40
like writing or design,
- 2:43
it's much harder, right? Like how do you
- 2:44
define what's like a great tweet or
- 2:45
what's a great art piece or what's a
- 2:47
great website? Um I don't know what's
- 2:50
the last time that you interacted with a
- 2:52
poem or walked into a coffee shop and
- 2:53
for some reason it kind of like hit
- 2:55
different and it felt
- 2:56
very special. Uh but probably it's a
- 2:58
combination of things that it it felt
- 3:00
very unique. It felt almost a little
- 3:02
different. It kind of called your
- 3:03
attention. Uh it felt like there it was
- 3:04
made with a lot of care and attention to
- 3:07
to detail and craft, and it almost had
- 3:08
the sense of of like authenticity. Um
- 3:11
and I think that's a lot of what AI is
- 3:12
missing today. It's like how do we take
- 3:14
uh things that are not necessarily
- 3:16
average, right? How do we produce things
- 3:17
that are purposely like out of
- 3:18
distribution? Um and slop is kind of the
- 3:21
opposite of that, right? I think it is
- 3:23
hard to define what is great sometimes,
- 3:24
but I think it's pretty pretty easy to
- 3:26
define what is slop in the sense that
- 3:27
most people would agree. I think the
- 3:29
sense of like repetition of kind of
- 3:31
soullessness is something that all of us
- 3:32
feel right now when using AI, and I
- 3:34
think it's quite magical, by the way,
- 3:35
that AI has gotten to a point that any
- 3:38
human on the planet that is not even a
- 3:39
designer, that is not an engineer, can
- 3:41
click a button and suddenly make an
- 3:43
entire PowerPoint or make a website or
- 3:45
make a web app. That's pretty cool. But
- 3:47
it comes with consequences, right? Uh it
- 3:49
comes with consequences of suddenly now
- 3:51
the cost of generation is basically
- 3:53
going to zero. Uh and But the average
- 3:55
person hasn't necessarily honed their
- 3:57
taste. Like I does think about the
- 3:59
amount of effort and work that a
- 4:01
designer puts in throughout their life
- 4:03
to like build up their taste, right?
- 4:04
Like there's all this process of like
- 4:06
getting exposed to many things and
- 4:08
learning to like spot patterns and
- 4:10
learning to develop a point of view and
- 4:11
like kind of do things in a in a
- 4:13
courageous way that maybe are a little
- 4:14
bit against the norm. Learning what not
- 4:16
to do and how to like have restraint and
- 4:18
that's very hard. Like the average
- 4:20
person doesn't necessarily have the the
- 4:22
time nor the skills to go and develop
- 4:24
taste in everything, let's say in
- 4:25
design. And so
- 4:27
um
- 4:28
I think it would be a bad case scenario
- 4:29
for us to just like be like, "Okay, the
- 4:30
way to fix slop is for everyone to have
- 4:32
taste." cuz I don't think that's
- 4:33
necessarily realistic. Um I think how do
- 4:35
we how can we understand this better so
- 4:37
that we can make even for the average
- 4:39
person the ability to create something
- 4:41
great and to understand maybe their own
- 4:42
taste um
- 4:44
easy more more more easy. So that's
- 4:47
that's a lot of what we're we're
- 4:48
focusing on. Um
- 4:49
So yeah, I mean this phenomenon of slop,
- 4:51
by the way, is not new. Uh if you were
- 4:53
in the internet uh as social media
- 4:55
emerged, you probably saw a lot of slop
- 4:57
before that. But I do think that AI has
- 4:59
been this kind of like accelerating
- 5:00
force, right? Of like being able to
- 5:01
create things very easily uh with a
- 5:03
click of a button and that like
- 5:04
thoughtlessness around it. And there's
- 5:06
kind of these three characteristics that
- 5:07
I I would say repeat in slop. Uh so A,
- 5:10
repetition. So you start seeing the same
- 5:13
thing many many many times. Um the
- 5:15
second is lack of fit, which I actually
- 5:16
think is is very related. So
- 5:19
fit is kind of this ability for
- 5:20
something to feel correct for a specific
- 5:22
context, right? For a specific moment in
- 5:24
time, for a specific person. Uh but
- 5:26
suddenly if you have repetition and
- 5:27
let's say one person asks for a website
- 5:30
for their pet shop and the other one
- 5:31
asks for a website for their finance
- 5:33
firm and somehow those designs converge
- 5:36
and look the same.
- 5:37
That's quite odd, right? Like if that
- 5:39
was in if you were actually crafting
- 5:40
that with care, that wouldn't you
- 5:42
wouldn't converge necessarily on those
- 5:43
things. And so this lack of fit and lack
- 5:46
of understanding of context is actually
- 5:47
huge problem that like leads to slop. Um
- 5:50
and the third is maybe low intent, which
- 5:51
is probably a mix of Yeah, you're going
- 5:53
to have a bunch of people prompting
- 5:54
really quickly and maybe just wanting to
- 5:55
one shot something. But I think there's
- 5:57
actually this like intent interpretation
- 5:59
piece that's missing in the systems that
- 6:00
we're building. Like how can you help
- 6:02
your user, right? Like how can you help
- 6:03
them better understand the intent that
- 6:04
they have um so that you can add more
- 6:07
color and add more context on onto what
- 6:09
you're trying to create.
- 6:12
Okay. And I I'm a big believer by the
- 6:14
way that you in order to fix something
- 6:16
first have to measure it and you first
- 6:18
have to understand it. I think that's
- 6:20
exactly why we're so focused on like how
- 6:21
do we
- 6:22
uh
- 6:23
turn these domains into something a bit
- 6:25
more verifiable so that we can attach a
- 6:27
measure to it. So uh
- 6:29
you'll you'll go on a little bit of a
- 6:30
research journey with me here now, but
- 6:31
we basically wanted to figure out can we
- 6:33
measure slop? Like can we actually
- 6:35
measure this quantitatively and spot
- 6:36
this and what does that like look like?
- 6:41
So we analyzed over 2 million websites
- 6:43
from the past like 10 years kind of like
- 6:45
way back machine style to try to
- 6:46
understand all the trends across like
- 6:48
design, how is the internet changing, uh
- 6:50
how are
- 6:51
how is like design changing over time?
- 6:53
And two things were interesting. And we
- 6:55
also met by the way then kind of
- 6:57
synthetically generated a set of uh
- 7:00
design websites so we could kind of like
- 7:01
compare like how does human-made sites
- 7:04
compare to AI-generated ones? And there
- 7:07
were a few things that were interesting.
- 7:08
So one was that you already kind of saw
- 7:11
a a bit of like a collapse
- 7:13
uh on the internet before even AI. So
- 7:15
you saw kind of the internet becoming
- 7:17
more homogeneous, using more similar
- 7:19
color palettes, using more similar
- 7:20
layouts, uh which is probably a function
- 7:22
of more uh
- 7:24
I would say this like kind of trends
- 7:25
spreading more and more and more
- 7:26
quickly, let's say. Uh but with AI I
- 7:28
think you saw this repetition happening
- 7:30
a lot more and being almost more like um
- 7:32
identified kind of regardless of
- 7:33
context. So even in completely different
- 7:35
buckets you saw patterns that were very
- 7:36
similar.
- 7:37
So we we built this I I call this
- 7:39
probes, but basically we uh we did two
- 7:42
things. So we did this like pattern
- 7:43
mining on all this data to understand
- 7:45
like what are features that we can
- 7:46
extract from all these sites. What are
- 7:47
all these characteristics that we can
- 7:49
make more objective, right? Colors,
- 7:51
typography, layout, audience. Like, how
- 7:53
can we like distill this down into
- 7:54
things that become almost like uh
- 7:56
structured? And then how do we uh train
- 7:58
up these like probes? So, think of these
- 8:00
as like baby classifiers. Like, how do
- 8:02
we uh train the ability to spot this one
- 8:04
characteristic?
- 8:05
And for all these slop sites, we
- 8:07
identify we started identifying like
- 8:09
what are the probes that basically mean
- 8:11
the site is very likely to be AI slop.
- 8:14
Um and especially when you start
- 8:16
combining them use and you see the
- 8:17
frequency of multiple of these happening
- 8:19
at once, it became very likely that you
- 8:21
could actually like measure
- 8:23
uh and predict slop. And we saw a a
- 8:25
super high basically ability to do that
- 8:27
prediction, which was really cool to
- 8:29
see. This performed better, by the way,
- 8:30
than like most LLM as a judge methods of
- 8:32
like asking an LLM to like judge if that
- 8:35
uh is like great human quality versus
- 8:36
like AI-generated slop. Uh so, that was
- 8:39
pretty cool to see. And I think kind of
- 8:40
shows this pattern that we see in AI
- 8:43
really being uh an actually quantitative
- 8:46
thing that we can see in slop, uh which
- 8:48
I find really cool. But, obviously, we
- 8:49
don't want to stop there, right? We
- 8:50
don't want to just measure slop. We want
- 8:52
to also solve it. And so,
- 8:54
um there's a few I I think I mentioned
- 8:56
this before, but like the as the cost of
- 8:58
production basically goes to zero, I
- 9:00
think the thing that becomes
- 9:02
expensive and matters more than ever is
- 9:04
judgment. Um I don't even want to use
- 9:06
the word taste here.
- 9:07
Is judgment. I think it's this ability
- 9:09
to discern what's right. It's this
- 9:10
ability to break down a problem so that
- 9:12
you can actually understand it and
- 9:13
create solutions for it. And so, yes,
- 9:16
there's the side of judgment that is
- 9:17
human judgment that I actually think is
- 9:18
more valuable than ever. But, there's
- 9:20
also the side of like how do we build
- 9:21
the right tools and systems to like fix
- 9:23
pieces of this problem, right?
- 9:27
So, yeah, how do we how do we fight
- 9:28
slop, my my enemy?
- 9:30
Um
- 9:31
And, by the way, I think there's there's
- 9:33
a lot of conversation going around how
- 9:35
do you fight slop at the model layer?
- 9:37
Like, how do we make models better? How
- 9:39
do we make models have a higher bar?
- 9:40
Which don't get me wrong, it has to be
- 9:42
solved and we're working very hard to
- 9:43
solve that, too. But I actually think
- 9:45
this problem of inference time is
- 9:46
equally, if not even more important.
- 9:49
Because that's actually when you
- 9:50
interact with the end user. And this
- 9:52
kind of back and forth of how do you
- 9:54
understand this context and intent
- 9:55
happens at the moment of inference time.
- 9:57
So, I don't think that we can ignore and
- 9:58
just make models better and not solve
- 10:00
this, otherwise slop will keep existing.
- 10:02
Um so, maybe breaking down a few of
- 10:04
those pieces and kind of
- 10:06
um
- 10:07
a few of the ways that we've thought
- 10:08
about solving this or a few solutions
- 10:09
that we built to solve this. But I
- 10:10
think, for example, for something like
- 10:11
repetition, one of the things that we're
- 10:13
working on is I I've nicknamed it, I
- 10:15
don't know if that's going to be the
- 10:15
official name, but like the creativity
- 10:17
API. How can we create a system that
- 10:19
almost becomes an inspiration machine
- 10:21
for your agent? So that it can produce
- 10:23
something that's actually out of
- 10:24
distribution instead of something that
- 10:25
is
- 10:26
in that same average and kind of mean
- 10:28
that we're seeing happen with like the
- 10:29
slop sites. Um so, this is one of the
- 10:31
ways that practically, if we can
- 10:33
intentionally produce something that's
- 10:34
out of distribution, you can improve
- 10:36
this like overall uh quality. And
- 10:39
by the way, I I don't think that this
- 10:41
can be something just like randomness.
- 10:43
So, it's not just about like turning up
- 10:44
the temperature of the model and and
- 10:45
kind of
- 10:46
fingers crossed hoping for the best. I
- 10:47
think it's much more like how do we
- 10:48
understand um even like what are rules
- 10:52
or expectations in specific domains?
- 10:53
Like let's say that you asked for a
- 10:55
slide deck for for the pitch of your
- 10:57
startup. Like
- 10:58
what is a what does a good pitch deck
- 11:00
look like? And then how do you almost
- 11:01
like intentionally break rules uh to
- 11:04
create things that are more creative,
- 11:05
right? Because usually creativity isn't
- 11:06
like randomness, isn't doing something
- 11:08
that completely feels off for that
- 11:10
situation. It's like you intentionally
- 11:12
maybe diverge on a couple of things
- 11:14
while maintaining kind of
- 11:16
um
- 11:16
adherence to to expectations of that
- 11:19
category, let's say for for others. So,
- 11:21
that's one of the things we're working
- 11:21
on. The second one of this problem of
- 11:23
fit, I think um it's interesting, but
- 11:26
brands, as probably a lot of you who are
- 11:28
designers know, takes so much effort to
- 11:30
create great brands. Like great brands
- 11:32
are the work of
- 11:34
dozens of designers putting in a lot of
- 11:36
like craft and thought and care.
- 11:39
And so we've almost like already
- 11:40
pre-done the work of defining what is
- 11:41
great for that specific company and then
- 11:43
we're not using it well. So this like
- 11:45
brand endurance actually think is a huge
- 11:46
problem and one of the things that can
- 11:48
very
- 11:49
more easily let's say like raise that
- 11:51
bar of quality. So I'll I'll touch on an
- 11:52
example on this one specifically. And
- 11:54
then same with like intent and judgment
- 11:56
I think the baby classifiers was a good
- 11:57
example. Um
- 11:59
like how it how we can actually like use
- 12:01
this to be even become a gate for slop
- 12:03
and and not let your agent ship slop.
- 12:06
But so the brand API is the first
- 12:08
product that we're releasing to to the
- 12:09
public. This is already in in beta
- 12:11
testing with a bunch of our design
- 12:13
partners. And essentially what it does
- 12:14
is it can take let's say a brand URL and
- 12:18
extract this into like very specific
- 12:20
components that are good for an agent to
- 12:21
follow. So basically how do we turn
- 12:23
something as fuzzy as a brand into
- 12:25
something so structured that it becomes
- 12:26
easy to
- 12:28
for your agent to follow that but also
- 12:29
for you to judge against it, right?
- 12:30
Because I think the piece that we can't
- 12:32
forget here is this judgment and
- 12:33
verification. So yes, this goes and
- 12:36
helps your agent to produce something
- 12:37
better.
- 12:38
But how can we also add a way for you to
- 12:41
judge okay, is the agent actually
- 12:42
staying on track? Is it actually
- 12:43
performing well to adhere to this brand
- 12:45
or how is it failing or where is it
- 12:46
failing? So this is the first flow I
- 12:49
would say that we we're seeing that is
- 12:51
really helping to improve quality.
- 12:53
And what's cool is of course we're
- 12:54
talking here about an example of a brand
- 12:56
that already exists. But let's say you
- 12:58
have an agent or you have an app and the
- 13:01
person that is using your app actually
- 13:02
doesn't have a brand. Let's say they're
- 13:03
an average consumer. Can we actually One
- 13:05
of the things that we're creating is
- 13:06
basically like a repository, like an
- 13:08
index of brands
- 13:10
of pre almost like pre-created brand
- 13:12
system so that if they want something
- 13:13
that feels dreamy, why not retrieve a
- 13:15
dreamy brand system that already has
- 13:17
been thought out to be cohesive instead
- 13:19
of doing like a generative approach the
- 13:21
moment of that might end up not so great
- 13:23
or might end up again in those pillars
- 13:25
of slop.
- 13:26
And I want to show you a real example of
- 13:28
this in action. So, um
- 13:29
there's this company that I think is
- 13:30
awesome called the General Intelligence
- 13:31
Company of New York. They have a sick
- 13:32
website, you guys should check it out.
- 13:34
Um but basically, if you ask Cloud
- 13:35
Design to create a slide deck uh in
- 13:38
their branding,
- 13:40
the the middle one is basically what it
- 13:41
comes up with. So, the one on the left
- 13:42
is is the original brand. Uh this is the
- 13:45
kind of the the default. And if you kind
- 13:47
of use this extraction actually in the
- 13:49
process, it creates something that's way
- 13:50
more high fidelity with the original. Um
- 13:53
and that even like in the details, I
- 13:55
would say like feels right. So, this is
- 13:56
just to show an example of it in in
- 13:58
action. Um
- 14:00
but yeah, I think we
- 14:04
I think all of us would agree that like
- 14:05
human human taste and kind of the peak
- 14:07
of human craft is always going to be
- 14:10
like deeply valuable. And that right
- 14:13
now, I think the challenge is we are
- 14:15
almost even not earning the right to
- 14:17
debate this like how can we have
- 14:19
uh models like reach this like pinnacle
- 14:21
of taste. I don't think it's about that
- 14:22
at all. It's like how do we first just
- 14:24
like
- 14:24
raise the bar. Like the bar is
- 14:26
currently, I would say, on the ground.
- 14:27
And so, I think all of this work that
- 14:29
we're putting into like how do we
- 14:31
decompose a problem and how do we
- 14:32
measure it is exactly so that we can at
- 14:34
least like improve this bar um of
- 14:36
quality. And I think we have to start
- 14:37
with that.
- 14:41
That's it. Uh thank you very much for
- 14:43
for the time. Uh this is this is
- 14:45
awesome.
- 14:46
>> [applause]
- 15:03
[music]