← All speakers

Bio, Work & Ideas

Sangwu Lee

Conference affiliation: AI Lead · Krea · 2026

Sangwu Lee is head of AI at Krea, where he develops image-generation models that give artists and designers greater creative control. His work on FLUX.1 Krea and Krea 2 challenges a central tradeoff in generative media: models optimized for reliable results often converge on the same predictable aesthetic.

Earlier, at the University of Rochester, Lee worked on computer vision and multimodal machine learning. His coauthored research on multimodal transformers incorporated visual and acoustic signals into pretrained language models; subsequent work explored humor recognition and video-based assessment of Parkinson’s disease.

At Krea, he helped develop FLUX.1 Krea with Black Forest Labs and coauthored the July 2025 open-weight release, detailing how aesthetic post-training could adapt a general-purpose image model for creative work. He subsequently led research on Krea’s own image foundation-model family as first-listed author of the Krea 2 technical report. The June 2026 release included openly available model variants and official inference code.

  • Creative exploration over aesthetic conformity. Lee prioritizes fast generation, stylistic range, and unconventional compositions because creative studios often need to discover what they want through iteration. He avoids training heavily on synthetic images that transfer another generator’s recognizable house style.
  • Dataset curation as creative direction. He preserves visually unconventional material that standard quality filters might discard, combines optical character recognition with vision-language captions, and uses hashes and compact classifiers to remove duplicates and unwanted imagery. When captions repeatedly omit details such as a painting’s white background, he treats the resulting hidden association as a training-data problem.
  • Sparse autoencoders for visual filtering. Lee applies sparse autoencoders to vision-model representations, producing unsupervised visual features that can identify watermarks, signatures, borders, and other recurring artifacts at scale.
  • Visual world knowledge with practical specialization. He checks coverage against prominent Wikipedia concepts, then develops capabilities for illustration, photography, and graphic design through progressively higher-resolution training, supervised fine-tuning, preference optimization, and reinforcement learning. His account of Krea 2’s research also identifies specialist-model distillation, richer visual conditioning, and simpler image-generation architectures as promising next steps.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Stylistic diversity depends on more than sampling: Krea’s training pipeline connects data curation, caption quality, scalable filtering and staged optimization to the images a model can imagine.

  • What if you do not know which image you want?
    0:15 ↗
  • Learn to denoise a compressed image
    3:27 ↗
  • Do not filter away the aesthetic
    5:16 ↗
  • A good image can still be bad supervision
    8:17 ↗
  • Spend expensive judgments where they matter
    10:01 ↗
  • Find unwanted features—and missing concepts
    11:55 ↗
  • Learn semantics first, then shape the distribution
    15:01 ↗
  • Translate user intent into training-distribution language
    17:35 ↗
  • Make the next experiment cheap
    18:30 ↗
  • Simpler models, richer descriptions
    19:11 ↗

References