AI Engineer World's Fair 2025

Advanced: Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han

Daniel Han2:42:28

Read the talk

From Reward Functions to Dynamic Quantization: Daniel Han’s Practical Guide to Efficient Language Model Training

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 946 seconds
From Reward Functions to Dynamic Quantization: Daniel Han’s Practical Guide to Efficient Language Model Training

Daniel Han explains how supervised fine-tuning, verifiable rewards, GRPO, careful sampling, selective quantization, and compiler optimizations fit together—and why reward design and efficiency matter more than algorithmic mystique.

From a talk by Daniel Han

At a glance

Ideas worth remembering

  • Use an instruction-tuned checkpoint or a small supervised priming stage before GRPO when a base model cannot reliably produce rewardable outputs; otherwise training can remain stuck with zero useful signal. 36:30

  • GRPO replaces a separately trained value model with statistics from multiple responses to the same prompt, while RLVR can replace a learned reward model with directly verifiable reward functions. 48:04

  • Treat reward design as the central engineering problem: combine correctness, format, and partial-credit checks carefully, and remember that a correct final answer does not prove the reasoning trace is valid. 1:30:35

  • Maintain sampling diversity, balance rollout count against memory and compute, and monitor answer-correctness rewards rather than assuming that improved formatting means the model has learned the task. 1:23:41

  • Apply dynamic quantization selectively: inspect activation and weight quantization errors, preserve sensitive layers at higher precision, and avoid assuming that only large-magnitude weights matter. 2:32:48

  • Distinguish demonstrated techniques from unresolved questions: Han treats whether reinforcement learning creates genuinely new capabilities, how far models should move from their starting checkpoint, and how subjective rewards scale as open or uncertain issues. 41:59

Training happens in stages, and each stage changes what the model can do

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 866 seconds
Training happens in stages, and each stage changes what the model can do

Daniel Han frames modern language model development as a sequence of optimization stages rather than a single leap from random weights to a capable assistant. Pretraining teaches next-token prediction over broad data; mid-training can emphasize higher-quality material and extend context length; supervised fine-tuning (SFT) teaches conversational or instruction-following behavior; and later stages include preference optimization and Reinforcement Learning with Verifiable Rewards (RLVR). He distinguishes RLVR from preference fine-tuning because it uses reward functions tied to outcomes that can be checked. 13:02

This progression also explains naming conventions that otherwise look inconsistent. Han identifies pretrained and instruction-tuned variants through labels such as PT, IT, Base, Instruct, and Chat, while noting that naming practices vary across model families. The practical distinction is that a base model primarily continues text, whereas an instruction-tuned model has already been adapted to answer requests in a more conversational format. 10:57

Although Han says reinforcement learning can be applied directly to a pretrained model, he argues that skipping SFT is usually inefficient. A base model may not reliably produce answers in the format a reward function can recognize, leaving training runs stuck with little or no useful reward signal. His recommended path is to begin with an instruction-tuned checkpoint or briefly prime a base model with supervised examples before applying reinforcement learning. 15:48

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:57 · section reference included

Reinforcement learning turns outcomes into changes in model behavior

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 1060 seconds
Reinforcement learning turns outcomes into changes in model behavior

Han introduces reinforcement learning through an agent, an environment, an action, and a reward. In a game such as Pac-Man, movement changes the environment, collecting an item can produce a positive reward, and encountering an enemy can produce a negative one. For a single-turn language model task, the prompt functions as the state and the generated response—including its reasoning trace, if present—functions as the action. 16:41

A simple arithmetic problem illustrates how much judgment enters reward design. An exact answer can receive a positive score, an incorrect numerical answer can receive zero or partial credit based on its distance from the correct value, and an invalid answer type can receive a stronger penalty. Han reports that distance-based scoring helped mathematical learning in his experiments, while emphasizing that binary pass-fail scoring is easier to implement and often more practical when meaningful distance is difficult to define. 19:31

The optimization target is not merely a raw reward in isolation. Han explains advantage as the reward relative to an expected or average baseline: positive advantage indicates an above-average action, while negative advantage indicates a below-average one. In conventional approaches, a separate value model estimates that baseline from the current state, and the language model’s own token probabilities provide the probability information needed to update the policy. 56:41

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:41 · section reference included

PPO adds guardrails; GRPO replaces a learned baseline with grouped samples

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 4330 seconds
PPO adds guardrails; GRPO replaces a learned baseline with grouped samples

Han describes PPO as a reinforcement learning optimization method involving a trainable generating policy, a fixed reference policy, and a separate value model. The reference policy represents the checkpoint the run started from, while the generating policy is updated during training. The value model estimates expected reward, but maintaining and training another large model adds substantial memory, compute, and potential estimation error. 22:52

The extra terms in PPO are presented as safeguards against unstable updates, overfitting, and reward hacking. A likelihood ratio compares the updated policy with the policy that generated an action; clipping limits excessively large changes; and a KL divergence term discourages the trained model from drifting too far from its starting checkpoint. Han stresses that these protections involve tradeoffs: restricting movement improves stability, but may also limit exploration of substantially different model behaviors. 1:09:06

GRPO removes the separate value model and estimates relative performance from multiple sampled answers to the same prompt. Han’s example generates several responses, scores each one, and uses the group’s mean and standard deviation to compute a normalized advantage; each prompt forms its own comparison group. When GRPO is paired with RLVR, a directly implemented reward function can also replace a learned reward model, reducing the number of heavyweight components required for training. 48:04

How it fits togetherHow GRPO builds a group-relative training signal

One question defines a comparison group.

Multiple responses to one prompt are scored, normalized within their group, and used to favor stronger answers.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

22:52 · section reference included

Reward engineering determines whether reinforcement learning teaches the intended task

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 7936 seconds
Reward engineering determines whether reinforcement learning teaches the intended task

Han repeatedly identifies the reward function, rather than the optimizer, as the difficult part of practical reinforcement learning. Mathematical answers can be checked directly, while code can be assessed through execution, imports, formatting, or other observable properties. More ambitious tasks expose the weakness of partial checks: a generated game might run and contain expected words or assets without actually being a good implementation of the requested game. 1:30:35

His notebook combines several reward components. A regular expression checks whether the model follows the requested reasoning-and-answer format; partial rewards recognize useful formatting progress before the full structure is correct; distance-based scoring gives higher marks to numerical answers closer to the target; and additional parsing handles formatted numbers. These examples show why reward shaping can help avoid an all-zero training signal, but also why each component must be tested to ensure that extraction and scoring work as intended. 2:09:03

Reward composition introduces additional judgment. If mathematical accuracy, code execution, and formatting are scored on different scales, their relative weights determine which behaviors training prioritizes. Han also warns that rewarding only a correct final answer does not establish that the intermediate reasoning was sound; likewise, an LLM as a judge can help evaluate subjective outputs such as summaries, but he cautions that repeatedly optimizing against another model’s judgments may eventually break down. 1:35:50

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:30:35 · section reference included

A practical GRPO run depends on priming, sampling diversity, and the right metrics

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 7707 seconds
A practical GRPO run depends on priming, sampling diversity, and the right metrics

The demonstrated workflow uses Qwen3, Unsloth, VLLM, and LoRA to run a resource-constrained GRPO experiment. Han sets a maximum sequence length of 2,048 for the example, notes that larger values increase memory pressure, and explains that 4-bit loading can reduce memory requirements. LoRA limits training to added parameter-efficient weights, while sharing VLLM weights between inference and fine-tuning avoids maintaining separate full copies of the model. 1:59:50

Before reinforcement learning, the base-model demonstration uses a chat template and a small amount of supervised priming so that the model can produce recognizable conversational and reasoning formats. Han initially describes a dataset containing 7,000 rows, then clarifies that the actual run used only a much smaller subset, later identifying 118 rows. He also says an existing instruction-tuned model can bypass this additional priming stage altogether. 2:06:02

Sampling determines whether GRPO has meaningful alternatives to compare. Han recommends avoiding zero temperature, suggests temperatures around 1.0 to 1.2 in the notebook and discusses higher settings elsewhere, and pairs higher temperature with min P around 0.1 to control low-quality outputs. More generations per prompt improve coverage but increase memory and compute; gradient accumulation can reduce the memory burden of larger effective training workloads, although Han notes that the usual batch-size equivalence does not straightforwardly hold for GRPO. 1:23:41

The training example begins with negative rewards, occasionally discovers high-scoring outputs, and trends toward more positive results over time. Han says the run took two hours and 54 minutes on the demonstrated free Colab setup, but cautions against treating formatting rewards as evidence of real learning: the important metrics are the reward components that test whether answers are actually correct. His example contrasts a base model that produces irrelevant text for the square root of 101 with a trained output that supplies reasoning and an approximately stated numerical answer. 2:15:52

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:23:41 · section reference included

Dynamic quantization and compilation extend the same efficiency-first philosophy

Selected presentation frame from [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han at 9466 seconds
Dynamic quantization and compilation extend the same efficiency-first philosophy

Han presents dynamic quantization as a selective compression strategy rather than a uniform reduction in numerical precision. He describes shrinking DeepSeek-R1 from approximately 730 GB to approximately 140 GB while acknowledging an accuracy tradeoff, and cites a Llama 4 Scout example in which a heavily quantized model remains close to a higher-precision result on the benchmark he discusses. His central claim is that some mixture-of-experts components can tolerate aggressive quantization, while attention layers, shared experts, and other sensitive components should retain higher precision. 2:31:38

The danger of indiscriminate quantization becomes clear in a vision-model example: applying 4-bit precision uniformly produces an incorrect image description, whereas preserving selected layers at higher precision restores the intended recognition behavior. Han recommends inspecting activation quantization error and weight quantization error to identify sensitive layers, because exhaustive testing of every possible layer combination would be prohibitively expensive. He also emphasizes that sensitivity patterns differ across models, including Qwen, Llama 3.2, and Pixtral. 2:34:07

Han further cautions that important weights are not necessarily the largest numerical outliers: a small value can still be critical to model behavior, making magnitude alone an unreliable guide to safe quantization. He discusses lower-precision formats including FP4 and MXFP4, while framing predictions about future GPU speedups as his personal view rather than established fact. His final practical recommendation is to experiment with Torch.compile, which he says can improve training speed and memory usage in some cases, while acknowledging that results vary and that compiler settings require tuning. 2:36:18

How it fits togetherSelective quantization preserves sensitive model behavior

Inspect activation and weight quantization errors.

Error measurements guide which layers remain higher precision while more tolerant components are compressed.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:31:38 · section reference included

Read the complete timestamped transcript
  1. 0:00

    [upbeat music] Hello, guys.

  2. 0:15

    Hello, hello, hello. Yes, sorry for being a bit late. There's a lot of traffic, [laughs] but, um, hello. Um, welcome to AI Engineer World's Fair. Thanks for coming to my session.

  3. 0:26

    Um, and yes, today we're gonna talk about the deep dive into RL, kernels, agents, and quantization. Um, you might know me or you might not know me, but I'm Daniel.

  4. 0:36

    Uh, my brother is somewhere. Uh, yeah, somewhere. But yes, um, thanks for coming again. Oh, we also have stickers and other, like, random stuff later. Um, that's after the talk.

  5. 0:47

    Um, so maybe you might know me, maybe you might not. Um, so we, um, on Twitter, we tweet a lot. Um, we did like a gradient accumulation bug fix last year.

  6. 0:57

    Um, we introduced something called Async offloaded gradient checkpointing. Um, we also work with the Hugging Face, Google, Meta, Mistral teams to, like, fix bugs in, like, their open source models like Gemma, Llama, Mistral, Phi, and more.

  7. 1:09

    Um, yeah, so like we... If you wanna follow on the latest [laughs] stuff of AI, better follow us. Um, we tweet about random stuff. Um, you might even know when next models might be released.

  8. 1:19

    You know, we sometimes tell people approximately, um, so that might be very interesting. Um,

  9. 1:25

    we also do, like, open source contributions to the entire open source ecosystem. For example, we contribute sometimes to llama.cpp. Um, we work with the Qwen team and Mistral on their releases.

  10. 1:35

    Um, we also do like, you know, Phi-4 bug fixes, Llama-4 bug fixes, which increase accuracy by a bit. Um, so definitely, you know, utilize some of the new, new uploads that we do, um, which fix bugs all the time.

  11. 1:48

    We just surpassed ten million monthly downloads on Hugging Face. Um, and yeah, we also have a GitHub package, um, with, with forty thousand GitHub stars. Um, and we do like...

  12. 1:57

    Essentially, we make fine-tuning faster and reduce memory usage.

  13. 2:01

    And yeah, that's the GitHub package. Definitely check that out. Um, there's like free Colab notebooks. I'm not sure many people know, but you have like free GPUs that Google offers.

  14. 2:09

    You can just use them. Um, please use them more. Um, and Kaggle, if not many people know, have thirty hours for free of GPUs per week, right? And there's no restrictions on it.

  15. 2:18

    You know, please utilize the free re- resources as much as possible. Um, yes, they won't be unhappy. Just use them all. [laughs] So yes, we have notebooks, so if you scroll down a bit on our GitHub pa- uh, page, there's like free notebooks, um, for C- Colab, Kaggle.

  16. 2:32

    We do reasoning, um, continued pre-training, supervised fine-tuning, and other stuff.

  17. 2:38

    We also upload models to our Hugging Face page. Um, for example, DeepSeek-R1-0528 released a few days ago. We upload 1.58-bit quants, which are like very small.

  18. 2:47

    They retain most of the accuracy, um, so, like, these can run on your local device. Um, even if you have, like, very low V- VRAM or, like, a not very good GPU, um, it will still work.

  19. 2:57

    And we will constantly upload models, so, like, sometimes people complain to us, you know, "Please stop uploading fixed models," you know, "It's kind of annoying." Um, but too bad, you know.

  20. 3:05

    Unfortunately, models have bugs, so we do have to fix them immediately. You will see sometimes, for example, the accuracy can increase by ten percent. Um, sometimes, you know, the large model providers won't tell you that they uploaded a fix.

  21. 3:17

    They're not gonna tell you. But, you know, be sure to download the latest models. You will get all the fixes.

  22. 3:23

    Now today, let's start off from history, right? Does everyone remember Llama? Um,

  23. 3:29

    although finally it got leaked, right? It was just a research paper, you know, Meta saying, "Oh, we trained Llama. Where's the weights?" You know, it's only research access, and then suddenly it got leaked, and it kind of just kind of spawned the entire open source movement.

  24. 3:42

    Um, you know, some of the people who are the, who are the authors, you know, are not part of Meta anymore, but Llama is extremely important for the entire ecosystem, and it was like the beginning of open source, kind of, um, for large language models.

  25. 3:55

    The most famous plot from the paper is this, right? So, like, you know, if you keep making the model train more, it just, the loss just keeps going down.

  26. 4:03

    Um, well, the question is, will the loss keep going down? That's the question, right? So Llama One was only trained at one point four trillion tokens, right? So now, one point four trillion tokens is actually very less.

  27. 4:14

    Um, most models are trained ten times more. Um, and you can also see from the trend that the bigger the model, the lower the loss, right? So seven billion is the blue line, um, and then sixty-five billion was, like, the red line, and you can see that, in general, as the model gets bigger, it, it gets smarter.

  28. 4:27

    Um, the training loss, the numbers are correct, so you should see generally these numbers from around maybe a bit higher than one. Um, if you see training losses when you do fine-tuning of, like, eight, thirteen, thirty, definitely something's wrong.

  29. 4:39

    Um, so you should get at least losses around two-ish, three-ish.

  30. 4:43

    And so now, uh, now Google, um, you know, Google's new Gemma 3 models are trained on fourteen trillion tokens, right? Which is much more. Um, Llama 4 is trained on thirty trillion tokens, right?

  31. 4:55

    So, like, literally, like, at least ten times more. Um, so Gemma is ten times more. Llama's, like, you know, thir- uh, thirty times more.

  32. 5:04

    Oh, yes, I forgot. You can also access the slides via the QR code if you want. It's also on the docs as well. Um, so if you go to the docs, there will be a link to all the slides.

  33. 5:11

    Um, I will probably post the slides anyways on Twitter and elsewhere so you can access them as well.

  34. 5:20

    Also, if there are questions for people, like, you know, raise your hand, ask. I will essentially do intermission between, like, some parts, and I will ask people if you have questions.

  35. 5:28

    Please ask. Um, last time the talk, many people asked questions. I will answer every single one, even if it's stupid. I don't care. Please ask. Um, sometimes I get stuff wrong, so yes, just ask questions.

  36. 5:38

    Um, are there any questions? [laughs] I mean, I'm assuming no. Okay. Okay, let's go to the... Okay, I will s- okay. So I, I don't know if people have seen this plot, very famous from Maxime.

  37. 5:50

    Um, he shows the open source versus closed source performance of popular benchmarks. I think that, uh, this is MMLU five-shot. Um, you can see the green line is open source models.

  38. 5:59

    Um, and then the red line, or the orange line is like, you know, closed source models. Um, and you can see that, in general, the slope of the open source models is more dramatic than the closed source models.

  39. 6:09

    And I would say, in general, that, you know, the open source models and closed source models, in terms of MMLU, they've kind of, like, reached the same. Accuracy, right?

  40. 6:18

    So, like, you can see that ... Okay, this is already outdated, but in general, you see like Llama 3.1 405 billion kind of reached, you know, GPT-4o level. Um, so open source models definitely have caught up to closed source models.

  41. 6:31

    However, there is a however. Um, recent, you know, like since September 2024, I would call this something called the o- open source drought. Um, no, no one wants to talk about it, but I will, right?

  42. 6:42

    So like September 2024, o1 got released, o1 preview. And to be honest, the open source community was shocked, right? So, like, suddenly the capabilities diverged, right? So there's something called the MMLU plateau, where most models, the open source models and the closed source models, they kind of converged.

  43. 6:57

    So the open source models was equivalent to the closed source models. But suddenly in 2024 September, you know, OpenAI released o1-preview, and it kind of shocked the entire community because the capability or intelligence kind of skyrocketed, right?

  44. 7:10

    So like with reasoning, long reasoning traces, it just was a total change of mindset. Um, and for four months, the open source community kind of died internally because there was nothing ...

  45. 7:21

    We can't replicate it. We don't know what to do. You know, do we do this? Do we do that? I don't know. Like, and so, like ... But then suddenly in 2025 January, DeepSeek-R1 came along, and they released R1, and that's when the entire world kind of changed their view, right?

  46. 7:34

    So, like, you can in fact train open source models to be as powerful as o1 or o3 or whatever, right? So, like, that's, that was what I call the open source drought.

  47. 7:44

    However, there was a previous drought even before that. Remember when ChatGPT got released in 2022 December, right? So, like, before even ChatGPT, most models were base models. Um, they were not really instruct fine-tuned that well, and so most large pre-trained models were actually useless, or they were terrible.

  48. 8:00

    But then suddenly ChatGPT came along, and they did better, you know, reinforcement learning from human feedback, better instruction fine-tuning, better instruction following, and it really changed the world, right?

  49. 8:10

    So, like, I think large lang- large language models were already here before 2022, right? It w- they were already there, but it was just ChatGPT which showed that if you have good data, right?

  50. 8:20

    Good instructions, good answers, good supervised fine-tuning, um, and good reinforcement learning, you can actually make the model very useful. Um, and yes, again, open source had a delay, a very long delay, right?

  51. 8:30

    Until, like, Llama 1, I guess. Um, and so always, I would say open source always tries to catch up to the closed source models. Um, the next question is, what is after reasoning?

  52. 8:40

    Um, is there going to be something else? Um, I think that's a very good question. My personal take is it's going to be very hard. Um, I think reasoning was, like, the last, most ...

  53. 8:49

    The DeepSeek R1 paper said that most likely the model already has these reasoning capabilities, and we just need to accentuate them. Um, and so, like, I'm not sure if there's going to be some other new, you know, like, step function where we get to, like, the next capability.

  54. 9:03

    Um, but in my view, I think, like, every single time the closed source models will always do, like, a step function. Um, but who knows? Maybe, maybe now it will plateau forever.

  55. 9:11

    I don't know. So that's like ... You have, like, you know, you have, like, long discussions about, you know, if AGI is gonna come or not, like, but who knows?

  56. 9:17

    Um, the talk is not gonna be about that, but, um, yes. Next. Um, so I call the first jump the SFT or RR, um, RLHF jump, right? So, like, that's essentially if you do good supervised fine-tuning, you get this large jump in performance.

  57. 9:30

    And then the second jump is called the RL jump, right? So, like, this essentially can increase performance dramatically if you employ second methodologies like RL, right? So, like, but the question is, like, what's the next jump?

  58. 9:39

    Um, I don't know. So I'm not sure if you guys saw this picture before. Um, it's very widely known in the community about, you know, by Yann LeCun. He essentially showed this cake, um, and essentially unsupervised learning or, like, just pre-training in general, um, you know, is just a cake, not that good.

  59. 9:57

    Um, and then supervised fine-tuning is kind of the icing on, on top of the cake, so, like, it's a bit better. Um, and then the reinforcement learning is the cherry, right?

  60. 10:04

    So, like, I'm not sure people like the cherry, but, like, some people like the cherry. Um, and so, like, the goal is how can we get the cherry? Um, but the problem is there is so less data about this, right?

  61. 10:13

    Reinforcement learning is, like, very, very, very less data. And so the problem is most mo- large model labs will train these large pre-trained models, and then they will iteratively refine it to make the model better through supervised learning, through reinforcement learning.

  62. 10:27

    Interestingly enough, this slide was actually shared last year, very popular, but actually this was from 2016 September. Um [chuckles], so I had to dig this up on YouTube. And so Yann LeCun actually talked about this back in 2016, so literally nearly 10 years ago.

  63. 10:41

    Um, I was like, "How can it ... Wait a second, that's 10 years ago." Uh, very long. Um, so this slide was actually very popular on Twitter. Uh, I, I think it was in November of last year, people kept tweeting about it.

  64. 10:49

    I don't know. I saw this, I was, like, shocked. Um, but yes, so this encapsulates, like, the current AI, um, boom.

  65. 10:58

    And so, like, firstly, like, when we talked about these large models, remember they started from a base model. Um, and so we call these training stages, right? So when you have a base model, um, you then convert it to a chat model, right?

  66. 11:08

    So, like, for example, ChatGPT is not a base model. It is a instruct fine-tune model or, like, some sort of fine-tuned model from a base model. So actually, OpenAI does have a base model somewhere sitting in their server somewhere.

  67. 11:19

    Um, they're probably not gonna serve it ever, but it is somewhere on the computers, and they essentially fine-tuned it to make ChatGPT 4. Um, Claude 4 has most likely a base model, and then they fine-tune it to become Opus, right?

  68. 11:31

    So, like, Gemini also has a base model, and they convert it to Gemini 2.5 Pro. So this, this phase, when you convert a base model to a chat model, is the fine-tuning phase.

  69. 11:40

    Um, and then the question is, like, you know, what do we do in the arrow, right? Like, you know, is it reinforcement learning? Is it supervised fine-tuning? Is it, like, some other special sauce?

  70. 11:48

    I don't know. But, like, we essentially will discuss about these, um, topics. Um, any questions first?

  71. 11:56

    Okay. So for example, in open source models, you might have seen Gemma 3 PT, um, Gemma 3 IT, Llama 4, Llama 4 instruct, Qwen 3 base, Qwen 3, um, Mistral Small base, Mistral Small instruct, Llama 2, Llama 2 chat, right?

  72. 12:12

    These, like, terminologies. To be honest, I think the open source community should standardize a meth- like, terminologies, like instruct or chat or, you know, no, no, not even a word, like, you know, IT and PT.

  73. 12:23

    Maybe they should, like, standardize it a bit. Um, but in general, if you see IT, it means instruct and instruction tuned. PT means pre-trained. Um, instruct just means, you know, instruction fine-tuned.

  74. 12:33

    Qwen-3 just removed it entirely. It's just called Qwen-3. Um, and then the base model's called with a base. Um, and so like essentially these na- naming methodologies, um, if you see on Hugging Face, um, hopefully you will now know, recognize these different types of, um, models.

  75. 12:49

    And so, generally what we say for like, you know, reinforcement learning and fine-tuning is I would say it's called re... Fine-tuning's everywhere. Um, you start off with pre-training, you then convert it into a supervised fine-tuning model via supervised fine-tuning, SFT.

  76. 13:02

    It's called SFT. You also might hear like IFT, which is instruction fine-tuning. They're the same thing. Um, and then we call something post-training, the post-training phase. Um, but actually recently it actually kind of changed.

  77. 13:12

    Um, so I don't know if you guys have been keeping up with the latest stuff, terminology. Um, I actually don't really like terminology anymore, but like we have something called pre-training, which is you take like all of Wikipedia, all of the web, you know, everything, all of the data you can ever see, shove it into the model,

  78. 13:26

    predict the next word. That's called the pre-training s- uh, pre-training stage. We then have something called the mid training stage, um, which essentially gives you higher quality data. Like for example, you can weight Wikipedia more because it's high quality.

  79. 13:38

    Um, you can essentially do long context extension as well. You shove this during the mid, mid pre-tra- uh, mid training stage. So if your context of your model is very short, you want to extend it to very long context.

  80. 13:49

    You shove this during the mid-training stage. And then the second stage is the supervised fine-tuning stage where you wanna convert the model to a chat model. And then we have the post-training phase, which is like press- preference fine-tuning, DPO, RLHF, and stuff like that.

  81. 14:03

    And then we have this new thing called reinforcement fine-tuning or RLVR. Uh, if no one knows what RLVR stands for, it stands for reinforcement learning with verifiable rewards. And this is like a new paradigm, um, not the same as preference fine-tuning or DPO, where we consider reward functions to make models much better.

  82. 14:22

    Um, and so this is how I would v- envision like, you know, the whole training phases of models.

  83. 14:31

    Another way to put it is we have some random initialization of the model, like some random weights of the model, right? So like seven billion parameters, literally random numbers, right?

  84. 14:40

    Like GPT-4, GPT-4, I don't know, 1.4 trillion parameters, like just random numbers. And then somehow we move in the space, like the black line, right? So pretend this is like some high dimension of 1.4 trillion dimensional space, and then we somehow move in this space, and then we get the final model, right?

  85. 14:54

    That's the green dot. The question is how do we move in this space to get to the final model? That's the question.

  86. 15:01

    Most people, what they do is firstly you start from a random initialization. You do the pre-training phase, which is very long. You get to this dark blue dot, right?

  87. 15:09

    That's called the pre-train model. And then you do some supervised fine-tuning, instruction fine-tuning to get the blue dot. And notice the line for the light blue line is very short because it is very short.

  88. 15:19

    Right? There is not that much data for supervised fine-tuning. And then somehow we get the blue dot, and then we keep doing more iteration to get to the purple dot, which is through preference fine-tuning.

  89. 15:29

    And then finally we get the g- the green dot, which is reinforcement learning via, you know, verifiable rewards, like, you know, o3 or o1. And so like the goal is somehow we have to move from the black dot to the green dot, and essentially all of large language models, all of AI is just an optimization problem, right?

  90. 15:45

    So like how do we make this easier to get to the green dot? You could, you know, you could kind of theoretically guess, you could just... Why don't you just go from the green dot, uh, black dot to the green dot, like skipping all of the dumb phases?

  91. 15:57

    You know, just skip it entirely. Yes, you could do that, but it's not gonna be very efficient. You're going to be waiting there for like, you know, I don't know, millennia.

  92. 16:05

    Your loss is not gonna go down. So the tricks that we found in reinfor- um, AI is like you have to do these phases to get to your final green dot.

  93. 16:13

    Um, there is like a new methodology where you can actually bypass this supervised fine-tuning stage and the preference fine-tuning stage and directly go to the green dot. There is a way, and that's the dark red line.

  94. 16:23

    Um, I think DeepSeek, uh, DeepSeek-R1-Zero kind of like sh- showed that. You can use a pre-trained model, a base model, and directly do some reinforcement learning with, uh, verifiable rewards and just skip it entirely.

  95. 16:33

    Um, so that's like a new paradigm that people wanna focus on. In my view, I think you should still do the light blue, the purple, and then the green.

  96. 16:41

    I don't think so you should like directly skip over to the green. If you wanna waste resources, you can skip to the green. Um, but I don't think so large model labs wanna s- waste resources.

  97. 16:50

    Um, hopefully not. So I, I don't know if people have seen this diagram, agents in the old sense. So like everyone keeps connecting agents with reinforcement learning. Okay, but like why?

  98. 17:02

    Um, so in general, an agent is you have some sort of environment. You have like the agent doing something in the environment. You get like an action. You do the action, and then you get like some, some sort of, some sort of reward.

  99. 17:13

    Um, and the reward is R. S is the state, so the current, what the environment currently looks like. Um, and then essentially RL tries to optimize this loop. You're trying to maximize the reward, um, giving some sort of action.

  100. 17:26

    Um, and that's why like, you know, that's why RL and agents are kind of connected. Assume the agent is the lang- language model, right? So assume the agent is in fact the language model, and so this, the environment is kind of fishy.

  101. 17:38

    Like, you know, it's, it's hard to say what exactly is the environment. It's more like the language model's inference space. That's the environment kind of. Um, but like pretend this was a game, right?

  102. 17:48

    So like the agent was the computer. The environment is like Mario, for example. Right? And so like you're playing a Mario game automatically, and your goal is to win the game.

  103. 17:57

    And so like the whole goal of RL is to maximize reward.

  104. 18:03

    Another one is like Pac-Man, right? So like you have the yellow Pac-Man. Um, you can either go up, down, left, or right, right? So like up, down, left, or right.

  105. 18:10

    That is the action. That's the A. The yellow... The, I think, I don't know, orange or something, yeah, the orange little things are like rewards, right? So like if you eat, if you eat a, you know, a yellow, a red, an orange dot, you will get positive reward, so R+.

  106. 18:25

    Right? So like you have R+'s. If you eat a very big one, you'll get like very large reward. But if you encounter one of the enemies, you will get minus reward, right?

  107. 18:33

    So like the que- the an- the question is how do we maximize the reward based on this environment?

  108. 18:41

    For language models, there is a trick. The trick is this loop kind of changes, because we don't actually have a continuous loop. Um, the state does not actually change over time, right?

  109. 18:52

    So like, for example, in a game, in a game, if you do an action, the whole state changes, right? The environment totally changes, and so you have to, like, continuously keep a history of the past steps.

  110. 19:02

    But in language models, there is no history, right? So like if you do a prompt, "What is two plus two?" If you ask another question, "What is four plus four?"

  111. 19:08

    It's, like, totally not relevant to your previous prompt. Okay, well, fine, it is kind of relevant, but you... Like, it's not directly correlated, and so, like, you can actually delete one of the lines, um, that's the next prompt.

  112. 19:19

    You can delete it entirely. So for example, what is two plus two, right? So, like, essentially, you have all of these options. It could be zero, it could be one, it could be two, it could be infinity, it could be B, it could be D.

  113. 19:30

    I don't know. It could be anything that you like. A symbol. And so the "what is two plus two" is the state. So, like, what is the question? That's the question, right?

  114. 19:38

    So, like, the question. The reward... For example, if you choose the one ... If you choose four, your reward is plus one. If you choose anything else, your reward might be negative infinity, zero, whatever number you like.

  115. 19:49

    You can come up with any number you like for reward. It doesn't have to be plus one. It can be plus 10. It can be plus 100. And you can do anything that you like.

  116. 19:57

    Um, you can also do distance-based scoring. For example, is, you know, is choosing the number five better than zero? So that's a question. What, what do you guys think?

  117. 20:08

    Is choosing the number five better than zero, or is it worse?

  118. 20:12

    I think it's better.

  119. 20:13

    Yes, okay, better. So what would you do for the reward then? Pretend the model outputs five for what is two plus two.

  120. 20:20

    Something less than one.

  121. 20:21

    Zero.

  122. 20:22

    Zero? Someone said zero? Okay, zero is fine, because it's wrong. Yes. Like, if you... The answer five is wrong, so you should probably give a reward of zero. But is there a better answer?

  123. 20:32

    Less than one.

  124. 20:32

    Okay, yes, less than one. So, like, some sort of, like, maybe zero point eight. I don't know, right? So correct. So, like, you could do, like, the answer divided by the correct answer.

  125. 20:40

    Right? So, like, if it's five, you divide it by four. You could do some sort of re... Okay, may- no, wait, that's wrong. It's five minus four divided by four.

  126. 20:46

    Um, so, like, something, some reward like that. Um, pretend if you're... Pretend if the model says A, what is the reward?

  127. 20:53

    Minus one.

  128. 20:53

    Minus one? Okay. Or it could be minus 10, because that's very bad, right? So you, you should not output a letter. It should be some sort of number. So that's how you design reward functions, right?

  129. 21:02

    You, you... We just design a reward function. Shove this into, you know, take your reward function. It's like just if statements. Shove this into a language model, fine-tuning phase, and there we go.

  130. 21:11

    You have o3. Okay, well, you won't have o3, but you know what I mean. Um, essentially, o3 is a collection of all these reward functions, right? So, like, for example, this, "What is two plus two?"

  131. 21:21

    Is one question. Remember, it doesn't have to be, "What is two plus two?" It is a general maths question. What is 10 plus 20? What is 10 times 200 divided by 10?

  132. 21:30

    You know, whatever maths equation you ever want. And this function can take your question and convert it into a number. And o3 is just a collection of all of these reward functions.

  133. 21:43

    And so the goal of RL is to make the good ones more good. You want the good rewards increase in value. So for example, the four, you want the four to appear more, but you want the three to decrease.

  134. 21:55

    You know, you don't want the three to keep appearing in your answer, but you want the D and the B to be very, very, very heavily penalized. That's the goal of RL, right?

  135. 22:03

    So, like, we don't actually have the answer, right? Okay, this, this question's very easy. What is two plus two? Obviously, it's four. Yes. Okay, that's very easy. But pretend you have some sort of complicated question, like, for example, how do I win the stock market?

  136. 22:15

    For example, like, dumb, dumb thing. Okay. How do I win the stock market? You don't know what actions you're gonna take, right? Like, but the point is you have the result, right?

  137. 22:22

    You have the result, like profit or loss. But the question is, we don't know how do we get to the good profit. And so the question is, how do we maximize good actions as much as possible and decrease bad actions as much as possible?

  138. 22:34

    And that is RL. And so OpenAI released something called RLHF in, uh, ChatGPT, and they sh- oh, actually, I think it was InstructGPT. But anyways, for ChatGPT, they showed that you just need some training data, you need some data.

  139. 22:49

    You interact with the agent, which is the language model. You then get some actions, which is your answer to the language model. You feed this into a reward model, and then you get some reward, and then you keep ite- you know, iteratively doing this step, and you'll finally get ChatGPT.

  140. 23:05

    So remember, the tr- the base model you start with, so ChatGPT base, you convert this into ChatGPT 4 via this method.

  141. 23:15

    To expand on a bit, if you guys have heard of PPO. Right, what is PPO? Um, essentially, PPO is just... You just expand the box for the agent, right?

  142. 23:24

    So the language model is like an agent. You expand it, and there's just three models inside of it. Um, there is a generating policy, the reference policy, and then there's a value model.

  143. 23:32

    Um, and that's all. It, it's not that special. Um, we will talk about each of these things separately, um, but PPO is just a optimization algorithm to make RLHF work better.

  144. 23:45

    GRPO, which is the algorithm behind DeepSeek-R1, smartly deletes one of the fact, uh, deletes one of the things, the value model. It just gets rid of it. Um, now, why would you delete it?

  145. 23:56

    We will talk about this, but the trick is, if you delete a model, the value model, you just save parameters. You save compute, and it's much more efficient, right?

  146. 24:03

    Remember, each of these models is kind of like a large language model, right? So, like, you know, pretend you're generating... Your model is, like, already one point four trillion parameters.

  147. 24:12

    What are you gonna do? Make another one point four trillion parameter model for the value model? So we just get rid of it, delete it, um, and that is GRPO.

  148. 24:18

    That's the only difference between... Okay, well, there's other differences, but the biggest difference is we get rid of the value model. Any questions?

  149. 24:27

    Yes.

  150. 24:28

    Um, so you, you talked about negative reward.

  151. 24:32

    Yes.

  152. 24:32

    It's confusing because in pre-training, isn't... It's like, um, it's-

  153. 24:37

    Probabilities?

  154. 24:38

    Well, it's like negative rewards, right? And then... So I guess it's the... Why is the phrase reward used when it can be positive or negative?

  155. 24:46

    Versus comparing it to pre-training where it's always negative

  156. 24:51

    Do you mean pre-training as in, like, the negative log likelihood or, like, some prob- ... So the, when doing pre-training, the goal is to maximize probability

  157. 24:58

    Yeah

  158. 24:59

    So like, you know, you output a, some number from zero to one, the probability of the next word, and you want to maximize that. For RL, you want to maximize reward.

  159. 25:08

    So if it's a negative reward, you still want to maximize that. So if it's negative one, you just want to make this negative one go away, and you wanna be positive, in the positive range.

  160. 25:18

    But also rewards can actually be just negative, right? So for example, if your, your reward function can be negative 10 and negative one. The good one is negative one, the bad one is negative 10.

  161. 25:28

    Your goal is to move towards negative one as much as possible, because your goal is to maximize it. So I would say the reward is a misnomer. You can just add 10 to everything, and then the, you know, the sc- it scales the numbers away.

  162. 25:39

    Um, does that kinda make sense? Or ...

  163. 25:41

    It does.

  164. 25:41

    Okay.

  165. 25:42

    It was a question of nomenclature, I think.

  166. 25:43

    Okay.

  167. 25:44

    Yeah.

  168. 25:44

    Yeah. Most people, I would say, like, to be honest, they don't like negative rewards. Actually, in the RL space, people just like to do positive reward. I don't know, I like negative rewards.

  169. 25:53

    Um, I, I feel like it's more ... For me, it's more intuitive. Um, yeah. Any other questions? Yeah, yeah.

  170. 26:00

    So you've got your language model, and then you've got generating policy and reference policy-

  171. 26:03

    Yes

  172. 26:03

    ... right? What, what models are being used there? Is it the same as your language model, or is that another model?

  173. 26:08

    Trick. Yes, very good question. There are some tricks that you can employ. Most people just make them at the same model. The reference model is, like, the beginning of the model.

  174. 26:16

    The generating model is the model that you're updating. So, like, essentially the reference model is, like, the base model. Okay, that's probably not a good... Okay, fine. Just keep it as the base model.

  175. 26:24

    The base model. The generating policy is actually the model where you update it. So, like, the model that you update the, the, the base model's update. So every single time you get a base model, you update one, update two, update three.

  176. 26:35

    That is the generating policy. But we will talk about this. So it's, it's essentially the same model, but there is updates to the model. So the reference policy is the model that is not updated.

  177. 26:46

    The generating policy is the model that is actually updated, and they're both the same model, so the, the language model. One of them's updated, the other one's not updated.

  178. 26:53

    But we will talk about that. Um, y- yes.

  179. 26:56

    So the actions, is it typically one token or is it more, uh, a series of tokens, like, uh, an actual full completion in a latent space? Or, like, how do you do it in, in context of-

  180. 27:07

    That's a good question. In the Pac-Man case, the action will be a string of actions, right? So, like, you can go up, then down, then left or right. You know, some sort of, like, long history.

  181. 27:15

    For, in the language model space, we just generally ... Th- this is called single turn and multi-turn. Generally speaking, currently, single turn is what most people do. It's just one action.

  182. 27:26

    So the action will be, um, the action essentially is saying, "What is two plus two?" And the answer is four, and that is the action. The action is actually the inference space.

  183. 27:35

    So, like, what is the actual chain of thought? That is the action. Um, and so, like, it's just one, though. But it is the total sum of the chain of thought.

  184. 27:42

    So, like, if you have, like, what is two plus two? I think the answer is four. You know, let me do some working out, blah, blah, blah, blah, blah, blah, blah.

  185. 27:48

    That is the entire thing. If ... Does that kinda make sense?

  186. 27:51

    Yeah, it makes sense.

  187. 27:52

    Okay. Yes.

  188. 27:54

    Uh, the, there was a Claude conversation around how, like, to finish a poem's next line, it had to, like, think ahead to the last letter, to the last word to match the previous one.

  189. 28:05

    Do you think, like, a reward focus versus next word completion allows you to do more think ahead, uh, inference?

  190. 28:14

    Y- you could

  191. 28:16

    ... really greedy when you're trying to, uh, understand the reward function instead of maximizing the next word

  192. 28:24

    I think for pre-training specifically, there are research base, papers which show that actually pre-training doesn't just predict the next word. It does try to predict many words ahead. And so, like, yes, maybe reward mo- In reinforcement learning, you can see, you can accentu- it essentially accentuates a pre-training behavior.

  193. 28:40

    So maybe this behavior already exists in the model, we just see it more often. So maybe I would say that,

  194. 28:47

    I would say that the model itself, it already has this capability. We just want to make it more obvious. And so maybe the model already even knows how to do that.

  195. 28:54

    It already knows it predicts 10 words ahead or 20 words ahead. It already knows how to do that. But we just wanna make it more obvious. I'm not sure if that answers.

  196. 29:03

    I mean, would it be safe to say, like, so if it's a-

  197. 29:05

    Yeah

  198. 29:06

    ... poem generation, you want, like, every last word, every ... You want almost a circuit for the last word to match the previous line's last word.

  199. 29:14

    Okay.

  200. 29:14

    It's more likely to build that circuit with generative, uh, GRPO instead of, like, the classic ones which really rely on next 100 word predictions or next 500 word predictions.

  201. 29:27

    I think for reinforcement learning specifically, yeah, I, I guess yes. Um, I think it essentially, your goal is to maximize the reward. And so, like, however, whatever way you try to get there, it's different from generally pre-training.

  202. 29:39

    Pre-training's just maximizing next, the probability of the next word. But reinforcement learning is you're trying to maximize reward. The question is, how do you actually maximize the reward? Do you, like, do chain of thought?

  203. 29:49

    Do you do, uh, what you describe, like thinking about the next, you know, in the future or something? I don't know. Like, I mean, the question is, like, what is the reward function actually doing?

  204. 29:57

    I don't know. What is the language model actually doing? I don't know. Um, I don't know if that answers your question. Like, it's, to be ... Yeah.

  205. 30:05

    Yeah.

  206. 30:05

    Yes. Yeah.

  207. 30:07

    Um, I was curious about when you were talking about, you know, the arithmetic, whether five is a better answer than, say, negative-

  208. 30:12

    Yes

  209. 30:12

    ... 100 or something-

  210. 30:13

    Yes

  211. 30:13

    ... for two plus two. And, like, given that there are these, like, closed circuits between all these different related, you know, mathematical functions you can do on numbers in space, like, whether it is, like, in the literature or in, like, the current state of the art-

  212. 30:27

    Mm-hmm

  213. 30:27

    ... better to train it so that a closer prediction is more accurate, or whether just saying, like, the right answer is right and everything else is wrong, which in some logical sense is true-

  214. 30:35

    Mm-hmm

  215. 30:35

    ... like, gen- tends to produce more performance-

  216. 30:38

    Yes

  217. 30:38

    ... in, in that space.

  218. 30:40

    Yes, you're correct. You should have data which is, like, getting more accurate data. Uh, is that what you're trying to say? Like, you should have data which is, like, for example, what is two plus two?

  219. 30:46

    You should get more data which is, like, four. It shouldn't be, like, five or 10 or minus 100.

  220. 30:50

    Well, like, saying that t- fi- like, the pr- the practice

  221. 30:55

    Yes

  222. 30:55

    ... like, in general, is that what people do in previous studies?

  223. 30:58

    Oh.

  224. 30:58

    Do they tend to say like-

  225. 30:59

    Yes

  226. 30:59

    ... like when there is, like, exactly one correct answer and everything else in a mathematical sense is equally wrong because it's not that answer?

  227. 31:05

    That is a good question. I don't know. Um, I think large model labs won't tell you exactly what they do. In our experiments, when you use our notebooks, we actually show that if you do distance based, so the closer your number is to the actual number, you will get better results.

  228. 31:19

    But generally speaking, it's easier to just say, "Five is wrong." Just give it zero reward. Everything is zero, and then the good one is, like, one. It's actually much easier to do.

  229. 31:28

    Um, for example, if you wanna do execution of code, how do you actually do distance-based scoring? Right? So like, if you ask it to create a Flappy Bird game, you just have to find an output.

  230. 31:37

    But you don't actually know how to verify, like, you know, oh, is this Flappy Bird ga- game better than the previous Flappy Bird game? It's only in the mathematical sense you can, like, do distance-based scoring.

  231. 31:46

    Um, I'm assuming large model labs, they probably do the zero one better. Like, the majority of them just do, like, yes or no, yes or no, binary. Um, but in our experiments, for math specifically, you should do distance-based scoring.

  232. 31:58

    Um, it makes the model learn faster.

  233. 32:00

    Thank you.

  234. 32:01

    Yeah.

  235. 32:02

    Uh, for verifiable domains like math, um, is it actually tenable... Because two plus two makes sense, but it's not gonna scale for, like, really large numbers or large multiplications.

  236. 32:12

    Mm-hmm.

  237. 32:13

    So are we gonna end up... Is the end game using tool use to calculate that, or the model could potentially be trained to solve that?

  238. 32:21

    That is a good question. In the olden days, before this paradigm came along, we would think that you can just use a tool, like a calculator, to... You should actually...

  239. 32:29

    I, I would say you should still use a calculator to calculate two plus two. Right? You should not use a language model. But with RL, you know, VR, the trick is we actually found that actually, wait a second, if you just do two plus two, or you do another question like ten times ten, or you do some

  240. 32:41

    sort of complicated mathematical expression, you know, like the derivative of X squared or something, I don't know, like some random mathematical equation, it randomly learns to actually solve that equation without actually doing over-fitting.

  241. 32:53

    And so, like, I would say that with RL, you can actually make the model actually learn how to do multiplication, how to do addition. So it's actually in the model.

  242. 33:01

    Um, does that kinda make sense?

  243. 33:03

    Yeah. Would we use this in production or it's-

  244. 33:05

    Oh, yes, yeah. People use that in product- I... Okay, maybe don't use it in production [laughing]. You know, you're not sure if the answer is correct. Maybe it will say three plus three is seven.

  245. 33:13

    Okay, I don't know. It's pop- it's possible. Right? So like, essentially... But it's getting better. Um, maybe in the future, all of mathematical equations can just be done by a, a model.

  246. 33:23

    Um, I think in the... Maybe a few months ago, maybe like Septem- you know, before o1 got released, actually not even o1, a few months ago, people would still say, "Use a calculator," you know, like some sort of tool calling.

  247. 33:34

    Yes, you should probably still do it. Um, but imagine, you know, as time goes on, as models get better and better and better in terms of, like, training data, just, just for the maths equation, you know, two plus two, um, imagine in the limit as we get all of the world's data for just this maths question, right?

  248. 33:48

    Two plus two, four plus four, ten times ten, or whatever, it should, in theory, solve them all. In theory. Um, yes, the... It's always in theory. Um, but yes, you don't need the tool calling.

  249. 33:58

    It's not necessary. Um, yeah. Y- Y- Yes.

  250. 34:02

    My question's about the reward model. In practice, are people using large language models as a reward model?

  251. 34:07

    Good question.

  252. 34:08

    For inverse reinforcement learning, or is it hand rolled or-

  253. 34:11

    Yes, I ac-

  254. 34:12

    ... reliable solver?

  255. 34:13

    Good question. Oh, I was gonna go in the next, next slides we'll be talking about that. Um, y- yes.

  256. 34:19

    How does this change with, uh, multi-turn?

  257. 34:22

    Multi-turn? W-

  258. 34:23

    How does this, how does it... Does this change with multi-turn? I mean, you showed just a single turn, right? So in a multi-turn, you can sort of loop out like a tree in

  259. 34:32

    Yes.

  260. 34:33

    How does this-

  261. 34:33

    You could do multi-turn. It's a bit more complicated. Um, you just imagine... There are tricks you can do. You imagine that your current step is good. Imagine it, and then you just continue doing inference.

  262. 34:44

    You, you append, like, your next question. Like, for example, "How am I gonna... What is two plus two?" You say, "Okay, let me think about this question. What is two plus two?

  263. 34:53

    Blah, blah, blah, blah, blah, blah, blah, blah, blah. The answer is four." And then the question is, what is your next question? Maybe the user interacts with it and says, "Okay, I, I don't think your question's correct.

  264. 35:02

    Oh, I don't think your answer is correct." And then the model says, "Oh, okay, let me rethink about this. Blah, blah, blah, blah, blah, blah, blah. I still think the answer is four."

  265. 35:08

    Um, so you could chain this all together and shove it into the, you know, the whole RL step. You could do that. Um, it's a bit more complicated. I think the diagram will be a little bit more different.

  266. 35:17

    Um, yeah.

  267. 35:18

    Yeah, I guess a follow-up to that, to that would be, um, do you, do, do you assume, like, a, a loop is a single turn, or a loop is a lot of turns, and then you only give a reward at the end?

  268. 35:29

    Or you give, like, sub-rewards or-

  269. 35:31

    Very good question. So there is... In the DeepSeek-R1 paper, you could do sub-rewards, or you could just do the reward at the very end. I think sub-rewards might actually do better in general, but the question is sub-rewards is very hard to calculate.

  270. 35:43

    You would rather just wait, you know, all until the very, very end and just give a reward. That's probably the easiest. So it's more about efficiency. It's all... To be honest, all of AI is about efficiency.

  271. 35:53

    What is more efficient? What is more... It's all optimization. Um, so the answer is, like, I would suggest people just to, like, shove a reward at the very end.

  272. 36:01

    Um, uh, yes.

  273. 36:02

    Once you give your reward signal, is it just the REINFORCE algorithm with the gradients to go back the weights?

  274. 36:08

    We will talk about that, yes.

  275. 36:09

    Okay.

  276. 36:09

    Yes. Mostly, yes. Correct.

  277. 36:12

    REINFORCE.

  278. 36:12

    Yes. So we will talk about REINFORCE. We'll talk about PPO, GRPO, and stuff like that. Um, y- yes.

  279. 36:16

    You mentioned with the slide with the squiggly lines that using RLVR, uh, you're all, like, to skip the intermediate steps is-

  280. 36:23

    Yes

  281. 36:23

    ... a waste of resources. Can you just talk about, like, why it's a waste of resources and, like, how you would choose to use, like, more resources?

  282. 36:28

    This one?

  283. 36:29

    Yeah.

  284. 36:30

    The problem is, if you skip from pre-training to the RLVA st- uh, RLVR stage, it's relatively hard because your model doesn't actually know how to do instructions, right? So like, you have this base model.

  285. 36:41

    You ask the question to the base model, "What is two plus two?" It's not gonna say, "I think the answer is four." It might... You might be lucky, somewhere in your pre-training data, somewhere on the web, someone asked this question, "What is two plus two?"

  286. 36:53

    And then, you know, the qu- the answer was like, "Oh, the answer is four." But you have to be lucky. Um, so like- So the problem of this is the whole trick of SFT is you wanna force the model to answer what is two plus two in an instruction way, right?

  287. 37:08

    So, like, it will tell the model what is two plus two. You want it to say the answer is four. You don't want it to, like, blabber on and, like, regurgitate...

  288. 37:15

    Like, get some Wikipedia article and shove it as the output. So the whole point of SFT, preference fine-tuning, and stuff like that is to make the model forced to make a more optimal, to, like, output conversation style.

  289. 37:27

    Um, if you wanna skip, it's also fine.

  290. 37:30

    It's just not efficient, um, because, like, you, you could do this. Um, I'm assuming large model labs are trying to do this. Um, so it's not like a you should or you shouldn't.

  291. 37:38

    They are trying. Um, does that-

  292. 37:41

    Yeah.

  293. 37:41

    Okay. Any... Are there... Yes.

  294. 37:44

    Yeah, I have a question about a couple of places. One question is that is this RLVR, is it a online policy optimizer or offline? Basically, does the reference model change after each block of spaces?

  295. 37:58

    The reference model does not change. So the reference model is just a model that you didn't train. Um, it's like the... It's like the, it's like the base model or, like, the SFT.

  296. 38:05

    Whatever checkpoint you started with, it doesn't change. You could change it. I think that'll be too expensive though. I think if you change it, that'll be more complicated. Remember, all of [laughs] AI is about optimization and efficiency, so I feel like you, you don't have to.

  297. 38:19

    You could. I, I don't know if there are papers talking about it though. Um, maybe OpenAI does it. I don't know. Um...

  298. 38:25

    So, so then the other question is that

  299. 38:27

    do we need less, uh, examples of, you know, training data for, you know, any of this RLVR-

  300. 38:36

    Mm-hmm

  301. 38:36

    ... compared to, like, weak pre-training and anything, you know, compared to SFT. Do we, do we need less training data or what do you think?

  302. 38:44

    So the trick of RL is you just need a reward function. You need to make that. And you don't need data. You don't need the answer of the data.

  303. 38:53

    Oh, actually, you do need the answer. You don't need the chain of thought. You just need lots of questions, like what is two plus two? What is four plus four?

  304. 38:59

    Remember, you can actually automatically generate this, right? So, like-

  305. 39:01

    The number of samples. Like, do you need less number of samples compared to SFT?

  306. 39:05

    You should do as much as possible. You... Most large language models, I think for, like, you know, o3 or o1, I don't know what is the percentage of compute.

  307. 39:12

    Maybe they spend, like, five percent or less. But the goal is what happens if you spend double the compute just on RL, right? So, like, previously, if you do fourteen trillion tokens on pre-training, make RL fourteen trillion tokens, and then the goal of large, large labs is to just do that.

  308. 39:28

    So currently, it's very less, but over time it will increase. Yeah.

  309. 39:34

    But, but compared to, like, SFT, you know, like they're saying here, the number of samples will be much lesser, right?

  310. 39:40

    Currently, yes.

  311. 39:41

    Because, because it's expensive. Like, we need, like, big models.

  312. 39:45

    Correct. It's expensive. But over time, I think, like, maybe by, I don't know, next year or, like, this year, large model labs, their goal is to do this phase the most.

  313. 39:54

    That's their goal. Because remember, you can automatically generate questions now. What is two plus two? What is two, two times two? What is ten divided by ten? I don't know.

  314. 40:02

    Generate as many math questions as you like. But remember, you can also generate, you know, like, coding questions. You can generate any questions that you like, or you can use the supervised fine-tuning data itself for the RL step.

  315. 40:13

    You can do that as well. Um, does that kind of make sense? Okay. Any... Okay.

  316. 40:19

    How do you protect the SFT from being, like, screwed up by-

  317. 40:23

    Oh, good que- yes. We will talk about that

  318. 40:24

    ... say, like, you have two plus two equals four.

  319. 40:26

    Mm-hmm.

  320. 40:26

    Five minus one is also a good answer. Two plus two equals five.

  321. 40:29

    Correct.

  322. 40:30

    But I don't want more equations. I want the answer.

  323. 40:32

    Very good.

  324. 40:33

    So is there, like, techniques to make sure that we're not violating our instructions to be concise?

  325. 40:37

    Yes. We will also talk about that.

  326. 40:38

    Okay.

  327. 40:39

    Um, clipping and stuff like that, yes. Yes.

  328. 40:42

    Is there any research that's been done on how, um, I guess RL specifically, but maybe training more generally affects specific circuits in the model? So, um, for example, you know, there's a circuit that, like, says two plus two is four.

  329. 40:56

    Mm-hmm.

  330. 40:56

    It just knows that.

  331. 40:57

    Yes.

  332. 40:57

    But can you incent... Is there anything that's, like, incentivizing the model to, like, learn addition, the concept generally?

  333. 41:06

    You-- That is a very good question. I don't know. I think that's, like, the, during the pre-training phase. Essentially somewhere, somewhere in the internet, someone wrote what is two plus two somewhere, and then somehow maybe someone did a formulation, like, you know, some sort of derivation of, like, what is two plus...

  334. 41:21

    Okay, prob- I don't think anyone has done... Uh, pretend there is some derivation of, like, some complicated math equation, and so the model somehow learnt to predict all of th- that entire trace.

  335. 41:31

    And if it keeps seeing this, it would, like, accentuate the fact, "Oh, okay, I've seen this before. Let's make this even more, um, more prevalent in the model." So somewhere in the model, it has learnt two plus two is equal to four somewhere.

  336. 41:43

    Yes. Uh, I, I guess what I'm saying is, s-s-so we, we can use, like, the RL stuff specifically. Um, you're saying that if you get, like, super bad reward or super small or super good reward, um, you can-

  337. 41:56

    You're weighting it

  338. 41:57

    ... you can kind of poke the model into a certain direction.

  339. 41:59

    Correct. So there is actually two schools of thought. The first one is the model already has this knowledge, right? It already knows what is two plus two, and you're just...

  340. 42:08

    RL just tries to, like, maximize what... If it sees two plus two is equal to four, it tries to weight this factor more, weight this circuit in the model more.

  341. 42:18

    So the model already learns, learns. But then the second school of thought is like, okay, maybe the model doesn't know actually, and RL actually learns a new thing. Um, I'm more in the first one.

  342. 42:27

    The model probably already learns.

  343. 42:28

    Yeah.

  344. 42:28

    It already knows it, and we're just maximizing the... You're just trying to make it more accentuated. Um...

  345. 42:34

    Exactly. So the extension of my question is basically, I want... I'm wondering if you can steer, or, or maybe there's like research of this, that you can steer the model away from the two plus two circuit to the always do the, like, do addition circuit, if that makes sense.

  346. 42:50

    You mean, like, just do addition?

  347. 42:52

    Yes.

  348. 42:52

    You could. I guess what you could do is, like, get the language model, see which weights are changed during the RL phase, which weights are changed, and you just give it what is two plus two?

  349. 43:01

    What is two plus two? What is two plus two? You just keep doing this question, and you can see which of the weights are changing, and essentially you can extract this from the model.

  350. 43:08

    You could I don't know if there's research about this, but I'm just making s-- I'm just making stuff up on the spot. You could do that, if that makes sense.

  351. 43:15

    Uh, maybe that's a research question someone should do, research paper.

  352. 43:19

    A-any other... Yes.

  353. 43:21

    If you go back to GRPO slide, like, uh, more and more in here, what about, uh, basically quality network, uh, that they save, you know, money in compute by just making equal solutions that they don't know, like you have to answer exactly four in your scenario for R1, right?

  354. 43:41

    That's one of the optimization-

  355. 43:43

    Yes, we will also talk about that. Yes, yes. Yes. Okay. N- I- I-

  356. 43:47

    Two more questions.

  357. 43:48

    O-okay, yeah, sorry.

  358. 43:49

    Do we expect usually changes, the RL changes all the parameters in the model, or is it, like, more targeted, like this set of RL would be this set of, like, targeted parameters, um, as, as a subset?

  359. 44:03

    You, you could. There's like two... Large model labs will most likely change all of the parameters. Every single parameter is changed. Um, but there are papers which show that actually not all of the parameters are actually changing that much.

  360. 44:14

    Some of them are changed by, like, zero. Like, the majority of updates to the model is, like, zero, and only some very small updates to the model are seen.

  361. 44:21

    And so that's kind of the circuit idea where, like, the model already knows how to do whatever question you give it. You just-- Most of the updates are, like, zero.

  362. 44:29

    Um-

  363. 44:29

    Is there any literature saying which one is better? Does it-- Is it better to aim for changing all the parameters versus a subset?

  364. 44:35

    You can also do, like, LoRA, you know, parameter efficient fine-tuning. You can do other things. You don't have to fine-tune every single thing. Um, but I think majority of large language model labs, they just do everything.

  365. 44:47

    Um, otherwise, this, again, becomes a optimization problem. You know, what do we-- which layer do we select, and stuff like that. So it gets more complicated. But yes, you could do parameter efficient fine-tuning.

  366. 44:56

    Actually, we're gonna share a notebook for that. So you can actually do it on your own computer. Um, yeah. There was... Yes, one more question. One more. Yes.

  367. 45:04

    Uh, yes. I wanted to add a side note about the model steering. It's called model abliteration.

  368. 45:10

    Oh, okay.

  369. 45:11

    You can steer a model towards a specific output that you want it, uh, to find that. But I also-- My question was, I wanted to, I guess, highlight what you said about the model can build capabilities outside of what it was pre-trained on.

  370. 45:23

    For example, uh, there was a paper from Google Research that talked about how the KL divergence constraint, for example, limits the capabilities that can be learned by the model, and you can only have the previous priors of the model that pre-trained.

  371. 45:35

    I did some experimentation with small language models and found that to be true.

  372. 45:40

    Mm-hmm.

  373. 45:40

    For example, I found that when performing o1 reasoning, it limits my sequence. So I wanna, I wanted to understand what are the implications for, for example, working with smaller language models?

  374. 45:50

    How important is this base model? And I guess, how can we work towards that kind of not make RL, uh, leading to this solution?

  375. 45:58

    Good question. For the KL divergence term, do you mean, like, removing it? Would that make it better?

  376. 46:04

    Oh, I was saying, uh, the KL divergence term... Well, I, I guess they removed it in the GRPO paper, um, because they found it didn't have an effect on the end result.

  377. 46:12

    But-

  378. 46:13

    Because the whole point of the KL divergence term is to, like, not make it too stray away from the supervised model.

  379. 46:18

    Reference-

  380. 46:19

    Okay. Maybe I'm not kept up to date with ref-- with research papers. But anyways, I... So your q-- your point was like, if we remove the KL term, it will be better, it learns new capabilities?

  381. 46:30

    Well, it was more... No, it was more, uh, there was some-- There was, like, a proof paper-

  382. 46:35

    Mm-hmm

  383. 46:35

    ... that described you can't build additional capabilities than what you had previously built in the model, almost your point about amplification capabilities.

  384. 46:44

    Mm-hmm.

  385. 46:44

    So I wanted, I wanted to understand what are the implications of that, and how can we prevent RL from being, uh, akin to distribution, for example? Or, or how can we kind of...

  386. 46:54

    It's almost like the essence of your talk, but I guess understanding how can we build on top of that.

  387. 47:02

    Do you mean, like, do you want to have more capabilities into the model?

  388. 47:06

    Yeah. So how, how can... Are there, are there strategies you thought of for-

  389. 47:09

    Uh, strategies? [laughs] Hard to say. I'm assuming the large model labs will probably know strategies. Uh, we will show, like... I, I will show examples of how to, like, make RL better, like how to, like, reach higher reward faster.

  390. 47:24

    I'm not sure about new capabilities. It's actually very hard. It's actually very, very, very, very hard to es- elicit new capabilities in the model. Um, the question is, like, is this new or not new?

  391. 47:35

    I think that's the question. Like, is this actually part of the model or not part of the model? Um, and most research papers are, like, hand-wavy. They say, "Oh, most updates are sparse."

  392. 47:45

    You know, like, so most likely it's not, you know, new capabilities. But what happens if, you know, one year later, all of the model updates are, like, not sparse?

  393. 47:54

    Is this considered new capability? I don't know. Like, you know, th-those are the questions. It's more like... I don't know if that answers your question. I probably didn't answer your question, but...

  394. 48:03

    No worries. Uh-

  395. 48:04

    Maybe we-- Maybe the other parts of the talk maybe might answer some part. Um, yeah, okay. I will keep going on. More questions later. Um,

  396. 48:12

    okay. So, like, the reward model, right, was actually a language model that-- or, like, some sort of model, some neural network, some AI model that predicts the reward. In RLVR, we delete this entirely, and we just call it the reward function.

  397. 48:26

    So, like, the ground truth reward, you know, if it's correct, you plus one. If it's bad, it's just zero, right? So you essentially delete another part. GRPO essentially deletes another part, right?

  398. 48:35

    So, like, you remove... Remember, GRPO, you delete the value model, get-- totally remove it, and then you delete the reward model, and it's just a reward function.

  399. 48:44

    And yes, as a reward model, you could use LLM as a judge. So you could ask a language model itself to say, "Is the answer good or bad?" You could do that.

  400. 48:53

    You could do regular expression check. You know, like, is the formatting of the answer good or bad? Is the maths equation good or bad? You know, is the final output good or bad?

  401. 49:03

    You can do distance scoring, stuff like that. You can also execute the Python code, and then you can see if it actually executed, right? So, like, is there, like, import errors or, like, format errors or, like, some sort of Python error?

  402. 49:14

    And you can use this as a reward And so this blue box, the reward, can be anything that you like. It just needs to output a number, you know, minus one, plus one, I don't know.

  403. 49:25

    It just has to be a number. In fact, you can make a dumb reward, just everything, it just does random. You know, plus one fifty percent of the time, minus one fifty percent of the time, I don't know.

  404. 49:34

    And confusingly, a paper recently showed that actually random rewards works. Um [chuckles] so like, yes, go ahead. You can try it. Um. [laughs]

  405. 49:44

    But also, why... Did someone say why? Um-

  406. 49:47

    Yeah. Why?

  407. 49:48

    Probably read the paper. [laughs] [laughs] But, but, why? To be honest, actually, I think the paper might be a bit...

  408. 49:54

    There was an-

  409. 49:55

    I don't actually believe it. There was a... Yes, there was an update showing that actually this was wrong. Um, that actually it's because the model, they don't-- The benchmarks are incorrect.

  410. 50:04

    So when you say that you actually increase accuracy, like from twenty percent to fifty percent, but actually the model itself was already fifty percent, they just didn't check the accuracy of the correct model before.

  411. 50:14

    Um, so there was a recent rebuttal to those types of papers. Um, but you know, interesting results. Um,

  412. 50:22

    yeah. [laughs] I don't know. [laughs] Okay, so remember in RL the goal is you don't know the best action to take in the space, right? When you're doing Pac-Man, I don't know if going left, right, up, or down is the best.

  413. 50:37

    I don't know. But at the very, very, very end, you will either d- you know, win or like, you know, get some reward, or you will die. Yes. Um, but the goal of RL is to maximize the best action you can ever take, right?

  414. 50:50

    So like what is a better action than all of the other bad actions? So RL just tries to maximize the best-- well, not the best action, the better action.

  415. 50:59

    Normal pre-training, you already know what is the best answer. It's like you already know what is the next word, right? You-- If you wanna predict, you know, "Hello, my name is Daniel," you already know the next word is gonna be Daniel, right?

  416. 51:09

    So you already know it. But RL, you don't know in advance what is the actual correct reward. Um, so you c- the only thing you can do in RL is to maximize the, you know, one of the better options.

  417. 51:21

    And so yes, okay, now more maths. Um, the goal is to maximize this equation. Um, that's the goal of RL.

  418. 51:31

    So what is this equation? The J is like the total gradient. Um, well, actually it's, it's more like the, it's more like we wanna maximize this. It's not actually the total gradient.

  419. 51:41

    Okay, maybe I misread that. [laughs] Anyways, pretend I didn't write that. Um, the-- We wanna calculate the gradient with respect to the policy language model, and the action is given a state, and the R is a reward.

  420. 51:55

    If you wanna write this down in like English, it's like we wanna take the derivative of the log probability of the action given the state times the reward. Now, I don't know if you guys understand what that means, but I did like a example.

  421. 52:08

    Um, Pac-Man case, okay? So you are Pac-Man. The red is your enemy. You don't wanna go there, right? So like you definitely don't wanna go to the red thing, right?

  422. 52:16

    But you wanna eat the two gray dots, right? So that's-- Remember, you can only go up, down, left, or right, right? You only have four actions. Remember, the action space is just up, down, left, or right.

  423. 52:27

    So if you do rewards, I just randomly made some rewards up. If you go to the red thing, you will get minus ten reward. Or actually it should be minus infinity, you die.

  424. 52:35

    But anyways, minus ten. If you eat the gray dots, you get plus one or plus one. And if you go up, it's just zero reward. There's nothing there.

  425. 52:43

    Now, when you get this language model or like some sort of model, it has to tell you what is the next action, right? It tells you what to do to the, the next action.

  426. 52:52

    For now, we will just assign every single action up, down, left, or right as one quarter probability. Right? So like you go up twenty-five percent of the time, left twenty-five percent of the time, and so on.

  427. 53:03

    So these are your numbers, right? This is the entire state.

  428. 53:06

    So the goal of RL is you want to do that red, going towards, um, the right. You want to do this-- You wanna go towards the right less. You wanna do a much less, right?

  429. 53:16

    So like you wanna push the probability of the zero point two five of the right much less. And you wanna go bot- like, you know, down and left much more.

  430. 53:26

    Right? So you wanna push the probabilities much more. And the top are not really that important. And so RL essentially, your, your goal is to avoid doing the bad thing, and you wanna do the good thing much more.

  431. 53:37

    That's kind of RL. If you convert this into a table, right, you have the probability of the action given the state. Right? Remember up, down, left, or right, we just assigned twenty-five percent chance.

  432. 53:49

    Right? Just, just pretend twenty-five percent chance. The reward which we can calculate, right? We calculated it. We calculated the reward. We just made some numbers up, right, as zero, one, one, and minus ten.

  433. 54:01

    The probability times the reward, we get some numbers, right? So like zero, zero point two five, zero point two five, minus two point five. And then if you take the log of the probability times the reward, you get some number.

  434. 54:11

    Right? So like zero, minus zero point six, minus zero point six, and six point zero two. So from this table, does anyone know which row do we want to maximize?

  435. 54:21

    What is the goal-- like, what do we wanna maximize?

  436. 54:25

    Which row?

  437. 54:28

    Bottom. Bottom.

  438. 54:30

    You want to maximize the bottom row? What is the reward of the bottom row?

  439. 54:35

    You wanna minimize it.

  440. 54:36

    Correct. You wanna minimize the bottom row. Remember, the reward is minus ten. We do not want to maximize the last row, because the last row is the worst. And so that means the six point zero two, we want to actually decrease this number dramatically, right?

  441. 54:46

    That's way too large. We want to decrease it. The other rows we want to maximize. And so the goal is, okay, we just take the sum of all of that, right?

  442. 54:54

    We take the sum of the four numbers, and it's four point eight.

  443. 54:59

    And so remember, okay, let's try, right? So like by hand, by hand we shall... Remember the right, remember all the probabilities are one quarter, zero point two five. By hand we shall do the bad action even more, right?

  444. 55:12

    We'll actually do the worst thing. What happens to the... What happens, right? So like the reward, the probability times the reward is now minus four. It used to be minus two point five.

  445. 55:21

    And so the reward, the log probability times the reward, the sum actually decreased. Right? It decreased to 2.58. Before it was 4.81.

  446. 55:31

    Is 2.58 smaller or bigger than 4.81?

  447. 55:35

    Mm-hmm.

  448. 55:36

    Obviously smaller, so actually this is worse. You should do, not do this, right? This is actually bad. Remember the goal is to maximize, maximize this equation. Maximize it, right?

  449. 55:46

    Maximize. And so 4.81 is actually better. The original state is actually better than the 2.58. So this, the thing that we just did is worse. So do not do this.

  450. 55:58

    However, let's do the right thing less, right? Let's not go to the right and actually maximize the rest.

  451. 56:05

    You shall see that if you do the log probability times the reward, sum them all, you will get 8.9, which is a larger number. And so the goal is to maximize this as much as possible.

  452. 56:15

    You could say, "Wait, wait, we know the answer," right? You should go towards the right. Just make this 100% probability. Let's just, you know, and you'll have an infinite reward.

  453. 56:24

    Okay, okay, not infinite, but you'll get maximum reward, right? Why don't we just do that?

  454. 56:29

    But you should not do that because you're actually forced-- If you do this, your model will be like learn, "Oh, okay, let's just keep going right. Let's keep going right," and it just gets stuck, and it just becomes very bad for optimization.

  455. 56:41

    So definitely don't do that. Now there is someone who's, who talked about REINFORCE. Um, we don't just multiply the reward, right? Remember this equation we did. Where is it?

  456. 56:52

    The probability of the action given the state times the reward. We don't actually multiply the reward. We should not do that. Um, you actually multiply by something called the advantage.

  457. 57:01

    Um, and what is the advantage? The advantage is a reward minus the average reward, the base reward. So you shouldn't actually just see you want to maximize reward. You want to actually maximize the reward, but also looking at the average reward across the entire model.

  458. 57:18

    So it's called the baseline. And so this B, this baseline model, is what is the value function, the value model. Remember GRPO deletes the value model? This was the value model, and this value model essentially estimates what is the average reward if we just see the current state.

  459. 57:35

    It does not take in-- It does not look at like, you know, what is the next, uh, next step. It does not look at, you know, what is the next action.

  460. 57:41

    It just takes a snapshot of what you're currently-- So essentially it looks at this, it looks at this, and just guesses what is reward, right? It doesn't-- You're, you're not supposed to give it the rewards.

  461. 57:52

    You're not supposed to give it minus ten, plus one, plus one or zero. It just looks at the current state and produces a number, and this number is called the average reward.

  462. 58:03

    And so the goal is now we don't actually want to maximize this, you know, just the reward. We want to maximize the advantage as well. Um, so like we multiply all this together and the goal is we want to maximize this new equation.

  463. 58:15

    Does anyone have any questions? There's lots of maths, but questions? Yeah. Yes.

  464. 58:20

    So in terms of probability, how do you-- Is the large language model emitting an estimated probability, or this is a known probability of all the possible states? But how do you get that in practice in, in the normal world-

  465. 58:32

    So large-

  466. 58:32

    It's not a game

  467. 58:33

    So a model, a large language model predicts the next word. So for example, you take the entire Wikipedia, and then you like chunk it into small little tokens, and then the output is just what is the next word.

  468. 58:44

    So for example, my name is Daniel, but it could also be my name is Michael, my name is Bob, my name is whatever, whatever, right? You have all of these probabilities for every single word in the entire language possible, like one hundred and twenty-eight thousand words.

  469. 58:55

    You assign a probability for every single one.

  470. 58:58

    And so it's-

  471. 58:58

    So the probability is based on the token.

  472. 59:00

    Yes, correct.

  473. 59:00

    But they'll be... Okay.

  474. 59:02

    The trick of this for language models is you can utilize the probabilities directly.

  475. 59:06

    Right.

  476. 59:06

    That's the trick, and so like that essentially makes everything easier.

  477. 59:11

    An-any other question?

  478. 59:13

    Yeah.

  479. 59:13

    Yes.

  480. 59:14

    So it's kind of the intuition behind this, you're kind of almost like normalizing the gradient to kind of give it a little perspective [inaudible].

  481. 59:21

    Yes. Yes. Yes, correct. Correct. Y-yes.

  482. 59:26

    What about multimodal models? Uh-

  483. 59:29

    Oh.

  484. 59:29

    They have, uh, uh, one model that is not only based on text, but it can be also based on vision.

  485. 59:35

    Multimodal models. Do you mean like doing RL multimodal models? Oh, that is more harder. I would say you, you could, you could look at the Sudoku puzzle, just convert the text model into a vision, just cheat.

  486. 59:49

    I guess you could do that. You could like say, oh, you know, I guess you could give it the Pac-Man, you know, give, give it the Pac-Man thing and tell the model, "What should I do next?"

  487. 59:59

    You, you could. I mean, vision, it's, it's kind of the same thing, but it's more--

  488. 1:00:04

    I-- Does o3 do vision plus reinforcement learn? I, I, well, I think it does. Um, y-yes, you could. I think for open source, I don't think I've seen... Yeah, I don't think s-- Yeah, I don't think I've seen open source models do that very well.

  489. 1:00:14

    Um, it is still very hard. Um, yeah. Any other questions?

  490. 1:00:21

    N-no? Okay.

  491. 1:00:23

    You can answer this.

  492. 1:00:25

    Someone did ask a question? Oh. Oh.

  493. 1:00:27

    So what is the-

  494. 1:00:28

    Yes. So sorry.

  495. 1:00:30

    So the average, what is the average, average of all the actions? Um-

  496. 1:00:36

    Oh, what is the B? What is the base model?

  497. 1:00:38

    Yeah, I understand average, but like what is the average on reference model for what action, you know?

  498. 1:00:43

    So it just like, your goal is to, uh, your goal is you see this current state of the model, like whatever the environment currently looks like, and you just wanna produce a number that approximates what is the total average reward.

  499. 1:00:55

    Okay, I'll give you an example. Pretend you're playing chess or like Go or remember AlphaGo. You look at the board, the current state of the board, and just say, "What is the probability of the white player winning?

  500. 1:01:06

    What is it?" You're not supposed to do any prediction. You just have to predict what is the probability of the white player winning by just looking at the board.

  501. 1:01:14

    That's kind of the average reward.

  502. 1:01:17

    So it's always low.

  503. 1:01:19

    It's always low? Yes. Yes, correct. But remember, at the very, very, very end phases, like, you know, you might get higher reward, but that's the goal. You wanna, you essentially wanna predict what is the probability.

  504. 1:01:28

    Always, you know, i-- for example, in chess- I'm sure there are, like, some steps you can take to make the reward higher. The question is, when the model sees this, you need to-- essentially, you need to say, "Is this board better than the previous boards?"

  505. 1:01:42

    And so this model, you have to train as well. You have to train this model. It needs to output a probability of it winning. That's for the chess example.

  506. 1:01:50

    Uh, does that kinda make sense or... No?

  507. 1:01:55

    Yes?

  508. 1:01:55

    But, so there's a policy model and reference model. Do we use reference model for this?

  509. 1:02:00

    The value-- No, no, no. The value model's totally different. It's an-- There's three models.

  510. 1:02:04

    Okay.

  511. 1:02:04

    There is a value model which predicts the average reward of the state. The reference model is just the model that you started with, and then the policy or the actual model that you're changing is the out-- the final result of your model.

  512. 1:02:16

    Like, the actual chat model. So there's actually three models.

  513. 1:02:20

    So you would use policy model to get B as well?

  514. 1:02:24

    This one, the B.

  515. 1:02:26

    Okay. Got it.

  516. 1:02:26

    You will see the current state. You will see the current state, and then you will see, okay, what is the actual... I think you do use the policy. No, you, no, you just look at the current state.

  517. 1:02:34

    You look at the current state, and then you output a probability of whether this chessboard is good or bad.

  518. 1:02:39

    I see.

  519. 1:02:39

    Of some, like, zero point eight percent you're gonna win. Um, okay. Yeah, something like that. Y-yes?

  520. 1:02:46

    So when we're estimating the advantage in PPO, uh, they tell you to take, like, the ratio of the new policy and the old policy to determine whether it's within-

  521. 1:02:55

    Oh, yes, yes. Yeah, we'll talk about that. This is just a general, uh, simpler formula.

  522. 1:02:59

    Okay.

  523. 1:02:59

    Yes. We'll talk about that.

  524. 1:03:00

    Okay, cool. I guess my question will be relevant then if you want to ask it then or no?

  525. 1:03:05

    Oh, you can ask. Yeah, go.

  526. 1:03:06

    Okay. Uh, I guess because the policy is predicting the next token-

  527. 1:03:10

    Mm-hmm

  528. 1:03:11

    ... we're-- that, that's the probability we're trying to reduce or increase based on the reward.

  529. 1:03:15

    Yes.

  530. 1:03:16

    Are you taking the ratio per token on the advantage while you're estimating the advantage, or are you taking the ratio across a turn?

  531. 1:03:27

    That is a good question, and that is an active area of research because you could either normalize by all the tokens or the entire just one turn. Remains to be seen which one's better. [laughs]

  532. 1:03:36

    Um, it's actually still people talking about that.

  533. 1:03:38

    Because if you keep adding it, uh, uh, sorry, keep taking the,

  534. 1:03:43

    the multiplication of all the tokens across your context window, you'll probably get a super low number, right?

  535. 1:03:49

    Correct. Yes. So generally speaking, normally people just assume this, assume this rollout is correct, assume this chain of thought is correct, and they just do the very end. But then you do have to multiply probab-- Wait, I don't...

  536. 1:04:01

    Yeah, you do have to multiply probabilities. So there is a multiplication somewhere. You will get very small, yes. Um, but you know, you'll get very small, but the numbers are relative, right?

  537. 1:04:10

    So everything is very small, but then the smaller one-- the bigger ones are still very small, but it's still better. So they're all relative.

  538. 1:04:18

    Yeah.

  539. 1:04:18

    Any... Was there one more? Yes.

  540. 1:04:21

    Yeah. It's relating to the REINFORCE, uh, slide.

  541. 1:04:24

    Yes.

  542. 1:04:24

    Um, so are you referring to the classical, uh, reinforcement learning algorithm, REINFORCE, or is it, like, some new thing?

  543. 1:04:32

    Oh, no, no, it's very old. Yes.

  544. 1:04:33

    Okay.

  545. 1:04:34

    Very, very old. Y-yes.

  546. 1:04:37

    Um, I wonder if you can, like, give some, like, advice on, like, how to think about this as a sort of framing or abstract level about, like, error propagation between, like, if you have a trained model which does the scoring or does the value function or whatever-

  547. 1:04:52

    Mm-hmm

  548. 1:04:52

    ... that itself is trained from data, it has some, like, error margin. Um, and that,

  549. 1:04:58

    you know, you have some softmax function, for example, that, like, only one in a hundred times will produce the wrong thing, but, like, it has that probability.

  550. 1:05:05

    Yes.

  551. 1:05:05

    Like, how do we think about the, like, development over time of these models and, like, to what extent that error propagation is something that you can observe and measure-

  552. 1:05:14

    Mm

  553. 1:05:14

    ... or, like, you know, systematize and engineer around? Um, like, just, like, I don't really understand, like, what the state, the mindset is in this sort of-- in this process right now around that.

  554. 1:05:24

    In my view, I think all of these formulas are just

  555. 1:05:30

    made up. And so, like, the goal is to maximize reward.

  556. 1:05:33

    Yeah.

  557. 1:05:33

    But the question is, you need to, like... You can't just maximize reward because otherwise you might make the model really silly. Like, you might say, "Okay, what is two plus two?"

  558. 1:05:41

    It just says four. Pretend your data set was just, what is two plus two, right? So, like, you literally just cheat. What is two plus two? What is two plus two?

  559. 1:05:47

    What is two, two plus two? Just make the model just say four, four, four, four, four, four, four. Just four forever.

  560. 1:05:51

    Yeah.

  561. 1:05:51

    Do you want this as a model? Definitely not. So, like, we want-- It needs to learn, okay, if I give it the next question, "What is eight plus eight?"

  562. 1:05:59

    It should not just say four.

  563. 1:06:01

    Right.

  564. 1:06:01

    Or what is two minus two? It shouldn't say four. And so the goal of all these algorithms is to somehow force the models not to, like, overfit to your question.

  565. 1:06:10

    And so, like, these formulations are trying to do these things to, like, not overfit.

  566. 1:06:15

    Yeah. Well, I'm thinking about, like, the chess example for-

  567. 1:06:17

    Okay

  568. 1:06:18

    ... you were saying. The thing which scores the board and produces this, like-

  569. 1:06:21

    Number

  570. 1:06:22

    ... a number-

  571. 1:06:22

    Yeah

  572. 1:06:22

    ... which is like, is this good or bad?

  573. 1:06:24

    Yes.

  574. 1:06:24

    Like, sometimes these well-trained models have these theoretical novelties that, you know, where they say, like, make this move, and it's like not-- the new state is not obviously good or whatever, but they somehow have, like, figured out to do this.

  575. 1:06:35

    Yes.

  576. 1:06:35

    Um, and suppose that your training mechanism for the value function model, you know, hasn't picked up on something like that.

  577. 1:06:43

    Mm-hmm.

  578. 1:06:43

    In fact, there's, like, some error in the tendency of the value scoring model.

  579. 1:06:48

    Okay.

  580. 1:06:48

    Like, its probability of producing, like, a, you know, some sort of, like, perfect scoring of the board's position-

  581. 1:06:54

    Yes

  582. 1:06:54

    ... for, like, black versus white-

  583. 1:06:56

    Yes

  584. 1:06:56

    ... you know, is not always exactly right.

  585. 1:06:58

    Yes, that's so-- Always not right. Yeah.

  586. 1:06:59

    And, like, that will bubble up into your training process-

  587. 1:07:01

    Yes

  588. 1:07:01

    ... in some sense slowly.

  589. 1:07:02

    Correct. Yes.

  590. 1:07:03

    And I'm like, how do you think about that? Like, how do you-- Like, what is the, um-

  591. 1:07:07

    The, the value model, you have to train it together. So it's like a combination of the entire algorithm. So the value model predicts what is the probability, but you actually have to train this as well.

  592. 1:07:16

    And so that is actually the problem. Some people, you could train this, you could train this separately. You know, you can, like, get all the chess possibilities and then output what is the final number.

  593. 1:07:26

    I think that's what some people do. You could train this in tandem with the model, actually. I think that's actually more harder. I don't know if that- You-- So there is always error in the value model, always.

  594. 1:07:35

    But you have to train this model as well, so you will reduce the error. But there's always error. So I think there's like some numbers you can like force the value model to be like less, less prominent.

  595. 1:07:43

    Like don't forcibly utilize the rewa-- uh, the value function. But in GRPO, we just get rid of the rewa- reward model, uh, the value model anyways. Um, so totally gone.

  596. 1:07:53

    Um, no, no, so you don't need to worry about that anymore. Um-

  597. 1:07:55

    Great.

  598. 1:07:55

    Okay. I will keep going on. Let me just check time, actually.

  599. 1:08:01

    Okay. Okay, so remember, the goal is the advantage-- We want to maximize advantage, not reward anymore. Advantage is the reward minus the average reward or the base, the base reward, right?

  600. 1:08:16

    If the advantage is less than zero, it means that it is worse than average. If, if the advantage is more than zero, it means that it's better than average.

  601. 1:08:25

    And so the goal is we want to do the action more if it's better than average on general.

  602. 1:08:31

    Now to PPO, right? So like, I don't know if you guys have seen the PPO formula. It is ugly, but this is the PPO formula, right? So like it looks-- it's more confusing because there's like a clip, and then there's epsilon and blah, blah, blah, whatever.

  603. 1:08:45

    But we could just sc- strip everything away. It's just, it's just the probability of the action given the state times advantage, right? We literally just discussed about this. Okay, minus a log.

  604. 1:08:57

    Uh, okay, the log's gone. But anyways, it's just that. And then the rest is, the rest is trying to reduce overfitting.

  605. 1:09:07

    And so remember-- So essentially this-- There is a thing called the division of the old model, and essentially it's the model that created the action. And the goal is we now want to maximize this likelihood ratio.

  606. 1:09:19

    We don't just want to maximize the probability of the model. We don't want to maximize the probability of the action given the highest reward. We actually want to maximize the likelihood instead.

  607. 1:09:28

    But what is this likelihood? So I did some numbers. I just made some numbers up. So this is the pack-- So pretend the, the top, the, um, numerator is zero point zero one, and the denominator is zero point zero one.

  608. 1:09:42

    Zero point zero one divided by zero point zero one is one.

  609. 1:09:46

    If the denum-- if the top is zero point zero one and the bottom is zero point nine nine, remember, these are all probabilities, you divide the top from the bottom, you'll get zero point zero one again.

  610. 1:09:56

    If the top is zero point nine nine and the bottom is zero point zero one, you'll get ninety-nine, and so on, right? The last one's one. And so the goal is zero point zero one divided by zero point nine nine is zero point zero one.

  611. 1:10:10

    This means that the action that you do is actually very likely, right? Because the bottom, the bottom thing is zero point nine nine, but we actually don't like this, right?

  612. 1:10:18

    Remember, the top is zero point zero one. We do not like this. So the ratio is zero point zero one. And then the bottom is this, this action, the bottom-- the denominator is zero point zero one.

  613. 1:10:29

    It is actually not likely, but we actually like this because the top number is zero point nine nine. And so when you multi-- when you do the division, you get ninety-nine.

  614. 1:10:36

    So this is actually good. And so actually, we're not actually trying to maximize the probability. We're actually trying to maximize the likelihood now.

  615. 1:10:45

    And so the question is, why don't we just maximize the probability, right? The first equation. Why do we need to do the division thing? Because if we maximize just the top, you will have reward hacking.

  616. 1:10:56

    What is two plus two? It might say, "To solve this question, we need to do blah, blah, blah, blah, blah, blah, blah, blah, blah." And then suddenly it says, "Hello, hello, hello, hello, hello.

  617. 1:11:03

    Hello, hello, hello," and then it says, "Four." Is this good? I don't think so, it's very good. We don't want it to say hello, hello, hello something, or like some weird trace in the reasoning model.

  618. 1:11:13

    It does something weird. We don't want this to happen. And so actually, this, this, you know, hello, hello, hello, hello is actually very not likely. And so the goal of the division is to reduce these issues.

  619. 1:11:28

    The epsilon part is called the trust region. Essentially, we don't want to make-- we don't want to do large steps for, um, PPO, right? We don't want to do large steps.

  620. 1:11:37

    And the, the trick is we want to restrict them, right? So like you don't want to overfit the model, so now we restrict the model. And so epsilon could be like zero point two, zero point one, zero point three.

  621. 1:11:47

    And the one minus epsilon is zero point eight. One plus epsilon is one point two. And the trick is we just want to not move the direction of the gradient that much, right?

  622. 1:11:57

    We don't trust the model that much. We don't trust the algorithm that much, so we want to constrain it.

  623. 1:12:05

    And then also the PPO, there's also a KL term. Um, there's another term. Um, essentially, what this does is we want the model to be as close to the supervised fine-tune model as much as possible.

  624. 1:12:15

    We want it to be-- We don't want it to go so far away from the base model, um, or the pre-tra-- or the supervised fine-tuning model. So essentially, if it deviates too much, we want to tax the, uh, uh-- we wanna tax it.

  625. 1:12:28

    And so this beta is like zero point zero five, and the KL divergence is the dist-- Okay, it's not a distance. It's like the distance between the current model and the pre-trained model.

  626. 1:12:38

    And essentially, we want to also shove this into the equation. Um, so you can see with PPO, there's many moving parts. It's-- Who cares about the equation? It's not that complicated.

  627. 1:12:48

    The point is all of these extra add-ons are just to reduce overfitting and not to make the model like randomly go to some weird state, um, that like, you know, overfits to like s- your questions.

  628. 1:12:58

    And so the trick of PPO is they just added all these terms in to make training more stable.

  629. 1:13:04

    And so the final equation is like this, this again. Um, hopefully you will... To be honest, no one even cares about the formula. It's not that important. Um, but I just tried to like break it down into pieces.

  630. 1:13:14

    And the go-- Remember, the goal is to maximize this equation, right? We wanna maximize it. And normally, I just like to think about this one, right? You just need to learn this one, right?

  631. 1:13:25

    You want to maximize the probabi-- So it's just this equation. Um, remember we did the table? Just this is enough. You don't need to learn the rest of the formulas.

  632. 1:13:34

    It's not very interesting. Um, yeah. Any questions? Yes.

  633. 1:13:42

    Like, the reason why PPO was preferred over, like, REINFORCE or whatever is because of the stability of-

  634. 1:13:48

    Yes, correct

  635. 1:13:49

    ... previous one. Could you go more in depth, like, what causes this instability? Why this, uh, minimizes the instability?

  636. 1:13:57

    So the biggest problem is pretend you are like... Pretend you just started RL. Like, you, you have the pre- you have, like, the base model. You have, like, a supervised fine-tuned model, and then you do RL.

  637. 1:14:08

    The gradient updates at the very beginning are gonna be gigantic, right? You're like, "What is two plus two?" It says four. But if it says five, you want to, like, penalize it dramatically.

  638. 1:14:17

    And so the problem is you don't actually wanna do large steps, and so the goal is you wanna constrain it. And so the constraint factor is, like, you know, if the, if the, if the num- if the gradient update is extremely large, you just want to, like, constrain all the numbers, if that makes sense.

  639. 1:14:32

    Yeah.

  640. 1:14:32

    The goal is just to constrain the update, not to make it too large.

  641. 1:14:35

    And then, so that's the clipping. What about the ratio?

  642. 1:14:38

    Oh, the, the ratio, it's the KL divergence. Oh, sorry, not the KL, the likelihood. To be honest, I think I need to do more research. I would ask Gemini exactly [laughs] what it is. [laughs]

  643. 1:14:51

    That's my answer. I'm probably not the best person to answer every single question. Yes, any other questions? Yes.

  644. 1:14:58

    So for the denominator, you know, so that, this, this probability coming from, like, the reference model, is that the process that corresponds to it, or how we can... That you see on the down there.

  645. 1:15:10

    Yes.

  646. 1:15:10

    Um, it's the old one.

  647. 1:15:12

    Yes.

  648. 1:15:13

    So is that, uh, for the correct answer of, uh-

  649. 1:15:18

    It's the model that actually created the action. And so the top one is all of the numbers that are actually, like... How do I explain this? The bottom one is the model that created the action.

  650. 1:15:28

    So for example, you- the model says you wanna go up, down, left or right.

  651. 1:15:31

    No, is that a, is that a correct action or the, the action-

  652. 1:15:35

    Oh, it just created the action. Like, it could be anything.

  653. 1:15:38

    Okay.

  654. 1:15:38

    It could be the... So it's the ma- it's the maximum. It's whatever action the model says currently.

  655. 1:15:43

    I see.

  656. 1:15:43

    It might be wrong, it might be good, it might be bad. It's just any action.

  657. 1:15:46

    Okay.

  658. 1:15:48

    Any other questions? Okay.

  659. 1:15:55

    One question.

  660. 1:15:56

    Uh, yes.

  661. 1:15:58

    Uh, has this been tried in latent space instead of the probability, uh, uh, PPO or for different probability as, uh, people try to simulate thinking in the latent space?

  662. 1:16:11

    Or REINFORCE, PPO and REINFORCE.

  663. 1:16:16

    I don't think so. Can I answer that question? I don't know. That's why-- [laughs] I don't know. Maybe research papers show it. I, I'm not sure. Any other? Okay. I will...

  664. 1:16:25

    Okay, so GRPO, the trick from PPO is we re- remember, remove the value model. We get rid of it entirely. We do not want to estimate the average reward.

  665. 1:16:34

    It's totally removed. Um, and the reward model is now removed as well for a reward function.

  666. 1:16:43

    So we get... Yeah, we remove it, right? Remember, the value model's removed. B is a reward, value model. We get rid of it entirely. But what do we replace?

  667. 1:16:52

    So the trick of GRPO is we do rollout or inference sampling. We get the answer, what is two plus two. You literally make four inferences. You just literally call the model four times.

  668. 1:17:02

    It could say the answer is zero, the answer is one, the answer is two, or the answer is four. You can do, do... You do, like, you just literally call the model four times.

  669. 1:17:11

    And you take the reward, right, zero... What is two plus two? The correct answer is four. So you want the last number to be one, but the rest is all zero.

  670. 1:17:23

    And the trick is you literally just take the statistics of your current rollout. You take the statistics of all of this. You literally take the reward minus the mean divided by the standard deviation.

  671. 1:17:34

    You get the Z-score. And this is your ro- this is your base model. This is your value model, right? This is-- There's no more value model anymore. It's just a number.

  672. 1:17:44

    And so I did this on the table as well, right? What is two plus two? If you think it's zero, uh, remember the prediction could be zero, one, two, or four, and your reward could be zero, zero, zero, or one.

  673. 1:17:56

    And if you take the mean or the average of all the rewards, you get zero point three seven five. If you take the standard deviation, you take zero point four three three zero one, and then you do the reward minus the mean divided by the standard devia-deviation, you get some numbers.

  674. 1:18:09

    Remember, the number four is correct. That is why, you know, reward minus the mean divided by the standard deviation is one point four four. It's the largest number. And so that is why we need to, like, max- we need to essentially maximize that good answer, and we want to reduce the bad answers.

  675. 1:18:27

    But why is it called group relative in GRPO? Because it's not just one question, it's many questions. It could be what is two plus two? What is four plus four?

  676. 1:18:35

    Okay, but my graphs are all the same. My plots are all the same. But anyways, imagine there's like four different tables. What is two plus two? What is four plus four?

  677. 1:18:43

    How do I create this Python function? You know, whatever. And there'll be four tables. And so the goal, group relative just means you-- for each question, we take the statistics within each group.

  678. 1:18:56

    For example, what is two plus two? You create four-- You literally call the model four times, and you get some, you know, answers. What is four plus four? You call it four times, right?

  679. 1:19:05

    Create Python code, you call it four times.

  680. 1:19:10

    Yes, there are other factors of GR-- So essentially we already explained what GRPO is, right? Everything you need to know about GRPO we already explained. In to- in the total mathematical formula, it looks kind of like this.

  681. 1:19:19

    Um, there's some rearrangement. For example, the minus beta, the KL divergence, is just taken out of the reward function. That's the only other difference.

  682. 1:19:30

    Um, hopefully, it makes more sense about the parts of the GRPO formula and stuff like that. It's actually not that complicated to understand. The majority is just trying to reduce overfitting, right?

  683. 1:19:39

    That's the whole goal, right? Minus beta times the KL divergence is to reduce overfitting. One minus epsilon, one plus epsilon is to reduce overfitting. The division reduce overfitting.

  684. 1:19:49

    Everything's reducing overfitting, right? So that's all of machine learning and AI. It's just to make the training more stable and to reduce overfitting.

  685. 1:19:58

    I would highly suggest there is-- these are the two things that I really highly suggest. Um, Nathan Lambert's Policy Gradients, um, book. It's online, though. Very, very, very, very helpful.

  686. 1:20:10

    Um, and, um, Yannic's video on GRPO and stuff, very, very, very helpful as well. Um, yeah. And then now I will go into a Colab demonstration of GRPO. Um, uh, before that, like, does anyone have any questions?

  687. 1:20:23

    Let me just check time. Questions, yes.

  688. 1:20:26

    So looking at the formula, the GRPO formula, it looks like everything there is about exploitation and not so much exploration. Yeah, you can say that the exploration is between the new model and the old model, but how about, uh, if the models can look in a totally different part of the search space?

  689. 1:20:47

    Because maybe the answer is there or, or, or better answer is there.

  690. 1:20:52

    Because it, it looks to me that all of, all, all, all, all of these is, make more salient whatever information is in the base model in the end.

  691. 1:21:00

    Yes.

  692. 1:21:00

    That can mean memorization, but you make more salient the right memorization. In this case, the question of the two plus two.

  693. 1:21:10

    To answer your question another way, I think it's actually because GRPO itself is the problem. Remember, the goal of all these algorithms is to force the model not to detract too much from the original model, right?

  694. 1:21:20

    With this, like, minus beta KL divergence, you know, one minus epsilon. All of this is trying to make the model not go towards too far away from the original model, and I think that's the problem because you're essentially forcing the model not to go too far.

  695. 1:21:33

    And so maybe there might be some new algorithm, I don't know, something, some other formulation which, you know, you want to go very far away. You could do that.

  696. 1:21:44

    Um, I, I don't know if there are any...

  697. 1:21:47

    I don't know if there's any research papers about that. I, I don't know.

  698. 1:21:49

    Okay.

  699. 1:21:49

    But you could-- Yes, you could do that.

  700. 1:21:51

    Thank you.

  701. 1:21:52

    Yes. Any other questions? Yeah, yeah. Yes.

  702. 1:21:55

    You probably heard that Yann LeCun is saying that, like, don't believe in LLMs, right?

  703. 1:21:59

    Okay. Yep.

  704. 1:22:00

    And that's why he's pushing towards the JEPA-

  705. 1:22:02

    Yes. Yes. Yes

  706. 1:22:03

    ... for the energy-based models.

  707. 1:22:04

    Yes.

  708. 1:22:04

    So what do you think about that?

  709. 1:22:07

    What do I think about it? I can't really comment, but I mean, he definitely is-- You should listen to what he says. Um... [laughs]

  710. 1:22:13

    But you don't see anything like movements in open source, uh, in that direction?

  711. 1:22:16

    I don't think so. I think open source, we kind of got captivated by RL, GRPO. I don't think so open source people are doing whatever he's talking about, unfortunately.

  712. 1:22:27

    Of JEPA?

  713. 1:22:27

    I think he needs to talk about them more. Yeah, JEPA. Yeah.

  714. 1:22:30

    Yeah. Yeah.

  715. 1:22:30

    And energy-based models. I-- unfortunately, I don't think so open source... Yeah.

  716. 1:22:33

    Okay.

  717. 1:22:33

    Maybe we should talk about them more, but yeah, I, I don't think so. Um, yeah. Yeah. Yes.

  718. 1:22:38

    Yeah. So for group relative, uh, so GRPO,

  719. 1:22:43

    I was-- So it's, it removed the value function of like-

  720. 1:22:46

    Yes

  721. 1:22:47

    ... the GRPO.

  722. 1:22:47

    Yes. You got rid of it.

  723. 1:22:48

    And then the... So for the rewards, the sampling of these rewards, I always thought it was very interesting, like, because you, you, with the LLM, you can actually create different types of samples.

  724. 1:23:01

    Yes, correct.

  725. 1:23:01

    My question is like, are there strategies related to understanding whether you want to have a larger variance or you want to have a larger group? And-

  726. 1:23:10

    Yes

  727. 1:23:10

    ... making sure that this mean convert, uh, is more related to like the, the value function, the true value function of like that LLM, let's say, let's just say at that state.

  728. 1:23:20

    Because I'm not sure, like... Because it seems like when you, you constrain the group to be very, very small, you're going to almost, uh, certainly go through some kind of bias in the unobserved variables unless you have like maybe a large enough sample that it might--

  729. 1:23:35

    So one of these larger traces would be way, way more valuable than, let's say, traces you would never get if you had a very small batch size.

  730. 1:23:41

    So this is more about an optimization question. So in theory, you should make the-- You should make-- For example, I just selected four, right? What is two plus two?

  731. 1:23:50

    Create four examples. You should do, you know, as many as you like. You know, three thousand, whatever number you like. You should do as much as possible. But remember,

  732. 1:24:01

    AI is about optimization. This is gonna take forever. You know, it's all about efficiency. So probably don't do as many as you like. Um, but you should. In the limit, you should do that.

  733. 1:24:10

    But you know, everyone can't just wait there waiting for the computer to spin. Um, so yes, you, you should do as much as you like. Um,

  734. 1:24:17

    that's it. But yes, also for like recommendations, when you do inference sampling, you should set, set temperature as like, you know, one point two, one point five, set min P to be zero point one.

  735. 1:24:28

    You know, something like that. If you set temperature to be zero, you will have the same answer every single time. So definitely don't do that. But you should have high temperature numbers to make the model, you know, produce new output as much as possible.

  736. 1:24:40

    Maximize variability?

  737. 1:24:42

    Yes.

  738. 1:24:43

    Uh, or distribution.

  739. 1:24:44

    You should try your best to maximize variability. Your outputs should not be all the same. Um, if it's all the same, I don't think the model is going to learn.

  740. 1:24:51

    So you should make it as different as possible. So that's why you should set temperature to be one point two, one point five, whatever, some large number. Don't do too large though.

  741. 1:24:59

    Um, any other questions? Yeah. Yes.

  742. 1:25:03

    Yeah. I'm, I'm curious of like the situation where like all the rewards are, are null or zero, basically.

  743. 1:25:11

    Mm-hmm.

  744. 1:25:11

    And the model doesn't have any way of learning anything.

  745. 1:25:13

    Yes, correct.

  746. 1:25:14

    So you're starting, like you mentioned, like yeah, you could kind of pivot them all away from their behavior maybe. Yeah. You were probably gonna talk about KL later. Um, but like, yeah, if the model doesn't have the capabilities of actually-

  747. 1:25:28

    Yes

  748. 1:25:29

    ... answering the question-

  749. 1:25:29

    Yes. Yes

  750. 1:25:30

    ... like all the first steps are only-- all the, all the rewards are basically giving no signal at all of the underlying thing you're trying, trying to train. Like how do you deal with that?

  751. 1:25:41

    I, I have faced that a lot of times with GRPO. Like moving to a larger base model helps sometimes

  752. 1:25:53

    Yes, that's a good question. So essentially, you're saying the model, if the model starts off with, like, no reward, like every single update is like zero, zero, zero, zero, zero, zero, zero, zero, zero, it's not gonna do anything.

  753. 1:26:05

    Yes, that happens all the time. But by chance, just by chance, you know, you have like some small, little, little, little probability, just by chance you will have some reward.

  754. 1:26:15

    So that's the trick. You will see this

  755. 1:26:18

    after 10,000 inferences, what is two plus two? Suddenly, the model says four. Suddenly, okay, just suddenly, just by random probability. Let's make this more. That's all. That's all of GRPO.

  756. 1:26:29

    Yeah, like this is addition, like it's a super simple task. If it was like a proof of mathematical concept, it may never come up with the right solution.

  757. 1:26:37

    Yes, may never, but remember, you're not doing this one question. You're also shoving this together with other questions. What is two plus two? What is four plus four? What is m- the, you know, derive the derivative of blah.

  758. 1:26:48

    Do this Python function. This step is very large. You essentially shove this all together, and the trick is, in general, it works, in general. Maybe by bad luck it might not work.

  759. 1:27:01

    But I feel like the bad luck won't last forever because remember, you're changing the samples, right? So like the question, what is two plus two, you're changing that. The next phase will be some other question.

  760. 1:27:09

    And so the trick is just by chance you will have a good reward, just by chance, and we just force that to be more.

  761. 1:27:18

    Does that kinda make sense? So it's all-- To be honest, it's all luck. Yes. It's all luck. We're just guessing or we, you know, we're, we're praying that there's gonna be some positive reward somewhere in the model.

  762. 1:27:29

    There will be negative reward, right? So like if your model is really, really bad, you, you can do negative reward. And so you just don't wanna do the negative one.

  763. 1:27:37

    You just wanna do the negative one less.

  764. 1:27:39

    And by miraculous probability, you know, just rely on probabilities, you will get a good reward somewhere, just by chance.

  765. 1:27:48

    Does, does that kinda make sense? I mean, all of the large model labs are just literally relying on the fact that that's what they're doing. They're just guessing. We're just praying for the GPUs to work, and then suddenly the reward comes out.

  766. 1:28:00

    I'm being serious. That's exactly what they do. They're just waiting for the algorithm to work, and then suddenly, oh, okay. That's why people do random seeds as well. So for example, the initialization of the model might not be good, so you just kill the training run.

  767. 1:28:13

    You do like five hundred training runs. Oh no, four hundred and ninety-nine of them are like zero reward. Oh, just kill them all.

  768. 1:28:19

    No, like-

  769. 1:28:20

    Don't release them

  770. 1:28:21

    ... to extend to that topic, like maybe it's, it's worth it. Like if... I have seen online trainings are like GRPO, most of them stays flat forever.

  771. 1:28:29

    Yes, very common.

  772. 1:28:30

    And, and spending some more time on doing a small SFT step, so the model may get like five percent of those rewards correct-

  773. 1:28:36

    Yes, that's the co- yes, I was gonna show you guys that

  774. 1:28:39

    ... before it take off.

  775. 1:28:41

    Exactly. So like you could force the model to answer some question. Like for example, you ask the question, what is two plus two? It's very easy, it's four. You just force it to learn, oh, okay, it should be four first, and then you do other steps.

  776. 1:28:54

    That is actually why... Remem- Okay, I have to go back to all the slides. I don't remember... Okay, where is it? How do I... Okay, I'll exit. Uh, it's the same as this problem.

  777. 1:29:05

    Where is it? This one. Right, someone asked about why don't you just start from the blu- you know, the pre-trained model to go to the green one. It's the same thing.

  778. 1:29:13

    Essentially, the trick is we want to do some supervised fine-tuning to make it know some instructions, so it knows something, and then you wanna go to the reinforcement learning phase.

  779. 1:29:23

    But if you wanna start from nothing, like just the pre-training phase, that's the hard part, right? Your reward might be zero, zero, zero, zero, zero, zero, zero, like zero forever, and then suddenly one, you know, suddenly.

  780. 1:29:32

    No, but I start from ping most of the time. Like I always-

  781. 1:29:35

    Oh, if you start from ping, ge-

  782. 1:29:36

    ...

  783. 1:29:36

    Okay.

  784. 1:29:37

    Different models.

  785. 1:29:37

    That's just unlucky then. [laughs] I, I think like if you, if you see z- zero rewards, most likely either, one, your reward function's not that good. Two, yes, doing priming or like, you know, make the model learn a little bit about your data is actually does wo- does help.

  786. 1:29:50

    Um, so there are tricks to make it work. But generally, I would just say it's bad luck, just bad luck. Um, and unfortunately, you can't do anything. It's not your fault. [laughs]

  787. 1:29:59

    It's, yeah, just unfortunate. Uh, yeah, yes.

  788. 1:30:03

    How should we think about like when is... Or like what to expect from GRPO? Like are we... Is, is it gonna be just that this is the way open source catches up with closed source models?

  789. 1:30:13

    Or is it something... Is another tool for like your average or like your, your competent ML engineer to be able to like specialize a model for a task? Like I, I just don't...

  790. 1:30:22

    Yeah, is, is there consensus on like where, what this is gonna bring us to? Is smart people like you gonna give us a really good open source model, or is it that, you know, we should think about this as a new tool for specialization?

  791. 1:30:34

    The algorithm is not special. The hard part is actually the reward functions itself and the data that you're gonna shove into the model. That's the hard part. So like I, I think there is a misconception like, you know, the algorithm is important.

  792. 1:30:47

    No, it's not. It's useless. Who cares about the algorithm? You can literally just use, you know, the general, the function which I gave. You can just use any algorithm that you like.

  793. 1:30:54

    But the problem is actually the reward function itself. I gave you some examples, you know, like what is two plus two? The answer is four. Yes, you can do distance based, but that's just one example.

  794. 1:31:05

    Can you... Like, can someone make a function-- Can someone make a reward function for like trading? Stocks. That's-- You do that, do that, right? And then there you have like a model for trading.

  795. 1:31:14

    Go ahead.

  796. 1:31:14

    So you, so you think it's going to be more that because people are able to create reward functions like a little bit more easy, like similar to like you could create a prompt.

  797. 1:31:21

    It's, it's an easier thing for most people to, to like iterate on, um, than-

  798. 1:31:26

    It's actually quite hard

  799. 1:31:27

    ... coming up, coming up... It's, but it's easier than coming up with the algorithm itself or whatever.

  800. 1:31:31

    Actually, I think col- collecting... So in the olden days, large model labs will ask like, you know, large data providers like Scale or whatever to create data. Like, what is two plus two?

  801. 1:31:42

    You literally have someone sit there and write, "Okay, the answer is four." But then you also have to do the chain of thought. Like, "Oh, I think the answer is four because of blah, blah, blah, blah, blah, blah, blah."

  802. 1:31:50

    Or like, you know, "This is my working out." You literally have to ask someone sitting there to make the data.

  803. 1:31:55

    Mm-hmm.

  804. 1:31:56

    The trick is ge- no more. You don't need the data labeling step anymore. It's totally gone You have the answer and you have the question. The middle step is totally removed, but you still need to make the reward function.

  805. 1:32:08

    You need to ver- you need to say, "Is the answer four good or bad?" For maths, it's very easy. For code, for code, it's somewhat easy. You know, you can check, oh, did you import the correct function?

  806. 1:32:20

    Did you import the library? Did, did your code execute, you know, some other reward functions?

  807. 1:32:25

    Mm-hmm.

  808. 1:32:26

    But it's still hard to verify if your actual function's correct. For example, let's say your question wa- let's say the task was create the Flappy Bird game. How do you actually know that the output's good?

  809. 1:32:36

    How do you actually know? We don't. You, you could, again, ask a human to verify, or you know, [smacks lips] let's test the Flappy Bird game.

  810. 1:32:43

    Mm-hmm.

  811. 1:32:44

    And then give it a good reward or bad reward. Or the trick is, did the game actually run? If it ran, plus one. Did you see the word Flappy Bird in the, you know, Flappy Bird inside of the functions?

  812. 1:32:56

    If yes, plus one. Did you see the, you know, the image of the Flappy Bird sprite be used? If yes, plus one. Yeah, something like that. So like you don't-- It's still...

  813. 1:33:07

    I would say the hardest part is writing the reward functions. And for open source specifically,

  814. 1:33:13

    if, you know, the whole open source community starts writing reward functions, we can probably beat o3, o1, like, you know. Okay, plus compute, you still need compute. That's the problem.

  815. 1:33:21

    You still need compute. But if you write good reward functions, you'll probably catch up in no time.

  816. 1:33:26

    So the end state here is that you want an open version of the closed model. It's not... [stutters] I guess my original question is like, is this a thing that m- people are gonna use like prompts to have specialized models separately, or is this more that we want one or many good open models?

  817. 1:33:45

    It de- Okay, that is a good question. It depends on which school of thought you're in. If you're in the thought that large language models already have the capability and you're just trying to accentuate it, then there'll be just one model, yes.

  818. 1:33:57

    But this model can only learn some facts, because otherwise you're like over fitting and like it's not... But if you're in the second ha- camp, that the model actually learns something new...

  819. 1:34:07

    Oh wait, did I say it right? I think it's the opposite way around. Um, the first one is you have many models. I so- sorry, I said it wrong.

  820. 1:34:12

    The first one is you have many models, because the model doesn't actually learn that much. But if the second one is the model actually learns everything, then you have this one gigantic model.

  821. 1:34:19

    I think OpenAI probably ascribes to that point. You know, like most large model labs think that actually RL can get you to AGI, right? It will know everything about everything.

  822. 1:34:28

    Any single question you ask, it already knows. And so like, that's, I think that's where they're trying to go for. For open source, it's more harder. I think the open source community consensus for now, for now, is the model, it already knows your questions, and you're just trying to accentuate it.

  823. 1:34:44

    And by doing reward functions, you're trying to weight the circuits more. You're trying to weight the model to know how to do these equations and stuff like that. So I think like

  824. 1:34:53

    the goal of open source is, you know, if the entire community comes up with good reward functions, write them all, [smacks lips] then the problem is we need compute. That's the second problem, right?

  825. 1:35:02

    If you shove both of them together, you will get O- O10 or something. I don't know. Um, right, imagine if every single person writes a reward function once per day.

  826. 1:35:10

    Okay, that's probably too hard. Once per day, we have like, you know, seven billion reward functions, more than OpenAI can ever come up with, and you will defeat OpenAI.

  827. 1:35:18

    But, you know, you need the compute part. That's the only problem. Um.

  828. 1:35:23

    Thanks.

  829. 1:35:23

    Okay. Any other questions? Yes.

  830. 1:35:26

    Oh, sorry. Okay. I'm curious because when you said that like, um, yeah, we know GRPO is gonna get traces that may get correct, so those get rewarded.

  831. 1:35:35

    Yes.

  832. 1:35:36

    Um, how do you feel about, like, saving those traces-

  833. 1:35:38

    Mm.

  834. 1:35:39

    -and just doing SFT on those traces?

  835. 1:35:40

    Very smart. That's what-- Yes, yes, yes. I don't know if large model labs do that. You could do that, yes.

  836. 1:35:45

    'Cause I feel like, yeah, we're just mining for those-

  837. 1:35:48

    Good examples

  838. 1:35:49

    ... traces that are correct.

  839. 1:35:50

    The only problem I would say is pretend the question was what is two plus two, right? And then the model says... [smacks lips] Okay, the, pretend the model just says this, okay?

  840. 1:35:59

    It says, "Oh, let me work out what is two plus two. I think it is, oh, you know, the number two means two apples, and I want to add two more apples.

  841. 1:36:08

    I think the answer might be three. Hmm. But let me rethink about it. Wait a second, it's like four." Is that a g- Should you fine-tune on that? I mean, you, you could, but maybe it's like cheating.

  842. 1:36:19

    Maybe it just says four by chance. Maybe like... Okay, I'll give you another example. Let's say that by chance... Okay, the question was what is two plus two? It says gibberish like, um, "I like to go to Paris for fun or whatever.

  843. 1:36:32

    I don't know. I like to go to the, you know, this event," blah, blah, blah, blah, blah, blah, blah, blah, blah, blah, blah. And then suddenly it says four.

  844. 1:36:39

    Yeah.

  845. 1:36:40

    Just by chance. Remember, we're still rewarding this. We're literally rewarding this as good. But this is not good.

  846. 1:36:46

    Yeah, but that's how we reward that slide back as-

  847. 1:36:49

    No, so we, we reward it at the very end. Remember, we see the number four, it is good. The question was what is two plus two. The model can generate anything it likes.

  848. 1:36:58

    It could-

  849. 1:36:58

    And we can reward that path too, right?

  850. 1:36:59

    Well, you, you could, but that gets harder. So the trick is people don't actually reward the steps in between. They just do the final step because otherwise it gets too complicated, right?

  851. 1:37:07

    So like, what is the intermediate steps reward? It's way too complicated. So what you do is you just reward the final step. So if, if you see the number four, it's good.

  852. 1:37:17

    But we don't know how we got there. So you should-- So yes, you could maybe at the very, very, very end step of RL, you can then use some data to do, do, um, fine-tuning.

  853. 1:37:29

    Yes, you could. But I think like in general, it's because we don't know what the process is in between.

  854. 1:37:33

    When you say we don't train on that thought or random-

  855. 1:37:37

    Oh, no, we don't train on the thought. Yes, we don't. But rem-

  856. 1:37:40

    Is that a problem if we train then?

  857. 1:37:42

    We don't train on the intermediate step in between. You don't have to.

  858. 1:37:46

    No, you don't have to. Yeah.

  859. 1:37:47

    You, you could. You could. But remember, we don't know if the trace is a good or bad. We don't know. So you can't just take this trace and then do supervised fine-tuning.

  860. 1:37:57

    Because pretend the answer was four is good, but we don't know the intermediate steps unless if you read the data, right? You could ask s- you know, some human labelist to like, "Oh, you know, please verify if this trace is good."

  861. 1:38:08

    You could, you could. But then that kind of defeats the whole purpose of RL. So, like, you don't, you don't wanna do this. Um, does that kinda make sense or not really?

  862. 1:38:17

    Yeah.

  863. 1:38:17

    Okay. Uh, yes.

  864. 1:38:18

    Can we see the demo?

  865. 1:38:20

    The what, sorry?

  866. 1:38:22

    The demo.

  867. 1:38:22

    Oh, yeah, yeah, yeah. We... Yes, we will do that. Yes. Yes. Uh, yes, uh, que-

  868. 1:38:25

    Um, if we have multiple categories of reward function, like one for math, one for Python-

  869. 1:38:29

    Yes

  870. 1:38:29

    ... what's the normalization among the rewards that are, like, best practices?

  871. 1:38:34

    W- what do you mean by normalization amongst rewards?

  872. 1:38:36

    So, like, if we have, you know, for four, you know, for the math one, two plus two, four gets a one, five gets 0.5, A gets zero. Then you have one for Python that might just be a does this compile, it's a binary, um, reward function.

  873. 1:38:49

    Yes.

  874. 1:38:50

    If we're running this in one big clump-

  875. 1:38:52

    Yes. One batch. Yeah

  876. 1:38:53

    ... like, what... How do we normalize the fact that-

  877. 1:38:55

    Good question

  878. 1:38:56

    ... we may have totally different types of rewards?

  879. 1:38:57

    Correct. It could be like minus 10.

  880. 1:38:59

    One row at plus 100 to negative infinity, the other person wrote a binary-

  881. 1:39:03

    Very good question. That's your choice. Unfortunately, that's the problem of RL. It's all about human choice. Like, you, you will have to decide. You know, is the Python one more important than your, um, than your maths, what, what is two plus two?

  882. 1:39:15

    Then you can weight it more. For example, your make two plus two the reward as minus one and one, and the Python function as, like, 1,000 and zero, right?

  883. 1:39:24

    You, you have to decide on the weighting functions. That is, that is your choice. Um, unfortunately, it is a...

  884. 1:39:32

    It's kind of like an art. You could, like, dumbly, you could just do everything as the same scale. I think that's what most large mo- I think that's what the large model labs probably do, is, like, everything, all the reward functions have the same scale.

  885. 1:39:43

    You know, plus one, minus one, plus one, minus one. Not plus 10 and then minus 1,000.

  886. 1:39:49

    So it's up to you.

  887. 1:39:50

    Okay.

  888. 1:39:51

    Yes.

  889. 1:39:51

    And what about stuff like, you know, is this a good summary? Is this a poor summary? How, how are we, how are we creating the reward function?

  890. 1:39:57

    That is the question. So, like, now you want to, like, analyze... Okay, that's where the LLM as a judge comes in. So, like, there is a school of thought that you can use a language model itself to, to make a number.

  891. 1:40:08

    You can ask ChatGPT, "Is this a good summary or is this a bad summary? Please give me a score from minus 10 to 10." Ask ChatGPT. You, you could do that.

  892. 1:40:17

    That's called the LLM as a judge thing. Um, there is a paper which shows that you can do this for some time, but then it breaks down. So you can't just keep calling the language model, you know, like...

  893. 1:40:28

    It's kind of like cheating, if I would say. Like, you're trying to call ChatGPT to train ChatGPT. Like, it will work for some time, but then it will break down.

  894. 1:40:37

    There was a paper, uh, I, I need to find the paper, but the paper showed that if you keep doing this, your actual reward actually goes backwards. Um, so, like, you will get more and more and more reward, and then suddenly it just, I don't know, by bad luck, again, it's always about bad luck, the reward just

  895. 1:40:50

    goes back. Um, so... Yes, I'm being serious. Like, all of AI's about bad luck and good luck. Um, and, you know, optimization, trying to do efficiency, um, that's what everyone...

  896. 1:41:00

    Yeah, un- unfortunately. Um, so, yes, you can use LLM as a judge. Um, is that kind of-

  897. 1:41:06

    Okay. If I use LLM as a judge, doesn't it just end up being a teacher and student model or?

  898. 1:41:11

    Yes. Y- but that's why, like, that's why sometimes... Essentially, the problem is, I, I need to find the paper. If you keep doing this, it will actually do bad.

  899. 1:41:19

    I mean, intuitively, it kind of works, but then at some point... There is actually another way. You could ask a language model to generate reward functions. That is actually another school of thought.

  900. 1:41:28

    You can actually ask a language model to generate 10, 7 billion reward functions. But the question is, are, is the reward functions good or bad? I don't know. Um, so, like, you now you need to, like, rely on the fact that the models are good or bad.

  901. 1:41:38

    You could then ask another language model to verify the reward functions. So, yes, you could do this. Maybe that's what OpenAI's doing. I don't know. Maybe OpenAI's goal this whole time is, like, "Oh, let's generate all these reward functions, verify each of them, and then shove it into the function.

  902. 1:41:51

    Let's see what happens." Maybe that's what they're doing. I, I don't know. Um, but yes, it is a student teacher. Um, uh, yes.

  903. 1:41:57

    What's your opinion on how to make reward models at scale efficient for the open source world? Is it a bunch of distributed experts? Is it inverse learned? Is it, is it something else?

  904. 1:42:08

    Do you mean, like, how to make reward functions more efficient in general?

  905. 1:42:11

    No, I mean scalable in the sense of going over many different domains-

  906. 1:42:16

    Mm

  907. 1:42:16

    ... comprising, uh, past just coding and math.

  908. 1:42:19

    Yes. The majority of reward functions currently, that's why it's called verifiable rewards, is, like, maths and coding. To be honest, I think coding's also hard. I th- I don't know why people lump...

  909. 1:42:28

    So coding, you can't actually verify technically it's correct. You can just say it ran, or the output is most likely correct, right, for some functions. But for example, the Flappy Bird game, tell it to create a Flappy Bird game.

  910. 1:42:40

    How do you actually verify if it even is the Flappy Bird game? I don't know. But you could, right? That's the whole point of the LLM as a judge.

  911. 1:42:47

    You could take the output of the Flappy Bird game, ask the language model, "Does this look like the Flappy Bird game?" And if it's yes, okay, plus one. If no, minus one.

  912. 1:42:56

    You could do that. Um, but you can only go so far. If... Does that kind of-

  913. 1:43:01

    I'm just wondering what, what your opinion is to scale beyond that.

  914. 1:43:04

    Oh. Hard to say. I think most... I, I think large model labs, they're currently just trying to use their own model to, to r- to, like, literally reward, reward it.

  915. 1:43:17

    Like, a- as I, like, described. You know, ask it, "Oh, does this look like the Flappy Bird game?" If yes, plus one. If no, minus one. And I think maybe that's what...

  916. 1:43:25

    I think large model labs, their view is if you keep doing this, you'll get to AGI. That's their view. I mean, if you think about it, it... You could, maybe.

  917. 1:43:33

    But then I always go fall back to, oh, but you might be bad luck, it's not gonna work. Um, so, like, I, I think in general, I think bad luck will just, it won't work.

  918. 1:43:40

    Um, you will only get so far, and then suddenly it just doesn't work.

  919. 1:43:44

    Does that... Okay. Any other? Y- yes.

  920. 1:43:47

    Um, how do you think about the scope of the task that you're, um, instilling to the model? For example, you can, you can train for the math, you can train for coding, you can train for planning.

  921. 1:43:58

    Do you think if you keep it very, um, task diverse, does it work well or, or should we specify the specific task?

  922. 1:44:05

    That is your choice again. So, like, if you want to specify, for example, you just want to make a legal bot, you're given some sort of court case, and if it's like, you know-

  923. 1:44:14

    ... the plaintiff wins or the defendant wins, I don't know. You could just do law. Yes, you could. You could do that. But in my view, you should combine it with other sources.

  924. 1:44:22

    You should combine it with some maths. You should combine it with some programming because the point is you don't want the model just to know, like you just, you don't want the model just to like overfit just to just law.

  925. 1:44:32

    Maybe maths might be helpful just by, you know, by chance again, maybe. You know, maybe coding might be helpful. Uh, probably not, but like, you know, in general. Um, so like you should combine other source, um, other domains together.

  926. 1:44:45

    Um, I feel like, you know, all the large model labs, their goal was to do every single domain possible, right? Like mine every single reward function in the whole world, make all the reward functions, shove it into the model, and just learns.

  927. 1:44:56

    Um, so like, yes, you should do more domains if... Yeah.

  928. 1:45:02

    There's another... Yeah.

  929. 1:45:03

    Yeah. I have a follow-up question to the practical point of view. So, so what, what is the latest research in should we fine-tune... Say, say let's say for incredibly one, a smaller model.

  930. 1:45:17

    So should we fine-tune the smaller model first or, or go with feature distilling to smaller model? Is that better for this particular thing, you know, uh, using GRPO and Q-learning?

  931. 1:45:30

    So the notebook I will share will showcase you should probably do some supervised fine-tuning first. It's called the priming stage. Um, otherwise, you'll... Remem-remember the plot over here, the, the, the, this one, right?

  932. 1:45:42

    You don't wanna be in the situation where like you're starting from like some bad, you know, pre-trained state and you're trying to go to the RL stage. Very not efficient.

  933. 1:45:50

    But remember, AI is all about efficiency. You don't wanna do this step. So we do have to do some priming, you know, the SFT stage and the other stages, if that's your question or...

  934. 1:45:59

    So should I, should we use bigger, bigger model first?

  935. 1:46:03

    Oh, if you want to, yes. The bigger the model, the better. Yes.

  936. 1:46:06

    I mean, it's, it's extra cost, you know. So this would... If you, if you can just do, if you just train the smaller model, do you think it's still efficient?

  937. 1:46:15

    What's your-

  938. 1:46:16

    That's the trick. So essentially, the research papers show that small models actually do work, confusingly enough, because essentially these small models, it just does longer thinking. It does longer reasoning traces if the model's smaller.

  939. 1:46:30

    And if it's a larger model, maybe the m-reasoning trace is like smaller in general. So like I, I feel like the small models actually do work. They do break down though.

  940. 1:46:38

    If you wanna do like very complicated reasoning traces, then maybe the small models might not work because, you know, there is only seven billion parameters. There's no-not that much space you can move.

  941. 1:46:46

    Um, and so the large models, you just have more space to move around. And so that's why large models are better. If... I don't know if that maybe kind of...

  942. 1:46:54

    I don't know if that answered your question, but-

  943. 1:46:55

    No, yeah, yeah. That makes sense. So another question I have is one task that... It's like just why we, if we want to fine-tune for our use case-

  944. 1:47:04

    Mm-hmm

  945. 1:47:05

    ... like we need one task. Can we use the-- We can also use one, the distilled DeepSeek model, you know, which is already fine-tuned in other tasks. So we can, instead of working with base model, we can take that and then fine-tune.

  946. 1:47:19

    Yes, correct. Exactly. Yes, that's what you should do. Yes. You can take a distilled model, like a already reasoning model, and then further fine-tune it. Yes, you could. Um, I would say it's a bit more complicated because you could do that, but remember, the reasoning model itself is already a reasoning model, and you're trying to fine-tune it

  947. 1:47:36

    to become other re- Like you're trying to do some other domain. It might be easier, it might be harder. It's all about luck again. I don't know. So you have to try.

  948. 1:47:46

    It's all trial and error. Um, yeah. Yeah, yes.

  949. 1:47:50

    Two questions for you. One, like, um, in sort it's pretty empirical.

  950. 1:47:54

    Mm-hmm.

  951. 1:47:55

    So you can just like try and see what works and what doesn't and see what sticks.

  952. 1:47:57

    Yes, correct.

  953. 1:47:58

    Okay. And then the other side, uh, how are you keeping up with all the papers and all the content that's being put out? Like I'm sure it's a lot.

  954. 1:48:05

    How are you learning, like some... The principles that you're like using to follow like latest work?

  955. 1:48:11

    To be honest, I don't-- You don't need to follow. That's my view. Don't try to follow the latest research because sometimes there may, like in the next day, it's like rebuttal of the previous paper, and then the next paper says, "Oh, it's a rebuttal of the rebuttal."

  956. 1:48:20

    I don't know. So I would not try to keep too much up to date with the latest research. I think the field has kind of matured, and it is mostly stable now.

  957. 1:48:30

    You might have like some algorithm increasing accuracy by one percent or two percent or some efficiency. Remember, all of the papers are about efficiency. It's always about efficiency, making the training more stable, reducing overfitting.

  958. 1:48:41

    It's always these similar, similar papers. Um, so I would say don't... You can keep up to date with papers. Twitter is very good as a resource. Sometimes I tweet about papers, although I, I don't suggest...

  959. 1:48:54

    The Nathan Lambert paper-- The Nathan Lambert, um, book is very good. He keeps updating it. Where is it? Where did I put it? Um, this one. The RLHF book.

  960. 1:49:04

    That is very good, so definitely read that. Um, he updates it all the time. And so maybe follow Nathan Lambert. He's actually a very good tweeter on like the latest research.

  961. 1:49:13

    Um, so he's very useful. In general, there's a lot of noise in the RL space as well. You don't know if the research is good or bad. Like, you know, rebuttals on top of rebuttals.

  962. 1:49:23

    So I would suggest people just to like try... It's just trial and error, right? Try to see if a reward function is good or bad. You know, is the loss not...

  963. 1:49:32

    You know, is the reward just zero, zero, zero, zero, zero? Like unfortunately, something's wrong or you just... It's bad luck. Try again. Um, so it's just empirical. Yes. Um...

  964. 1:49:42

    Are you gonna put these slides up anywhere?

  965. 1:49:44

    Sorry?

  966. 1:49:44

    Are you gonna put these slides up?

  967. 1:49:45

    Oh, yeah, yeah, yeah. Yeah, yeah. Um, yes, these slides should be up. I was supposed to make a Bitly link. Um, I'll probably do that later, but I will share the slides.

  968. 1:49:52

    Yes.

  969. 1:49:53

    Like can you just post them on Slack, the engineer Slack?

  970. 1:49:55

    Oh, yeah. Okay, I'll do that then.

  971. 1:49:56

    Appreciate it.

  972. 1:49:56

    Okay. Okay. Any... Uh, yes.

  973. 1:49:59

    Yeah, I have, I have two questions. So, um, the idea that the capabilities already exist and-

  974. 1:50:05

    Yes

  975. 1:50:06

    ... during, uh, reinforcement learning, you're not adding new capabilities. When you're talking about the value model and the reward model, is that something that you are creating in the-

  976. 1:50:17

    In the old PPO sense, you are-- the value model is a new model. The reward model is a new model. Yes. Remember, in GRPO, we delete the value model.

  977. 1:50:26

    The value model's totally gone. We, we create the value from just statistics from the distribution. We essentially just create, you know, four... What is two plus two? Just create four examples, four trials, and then find the mean, find the standard deviation, and the data is your va-- data is your value model.

  978. 1:50:41

    It's not even a model anymore. And then the reward model is no more as well. It, it's just reward functions. And that is why we call it reinforcement learning with verifiable rewards.

  979. 1:50:51

    It's not normal RL anymore. It's like you replace a reward model as well if... Does that kind of make sense?

  980. 1:50:57

    That does make sense.

  981. 1:50:58

    Okay.

  982. 1:50:58

    And my follow-up question is-

  983. 1:50:59

    Yeah

  984. 1:50:59

    ... so is it sound intuition that then there are some circuits within the existing model that have w-this, uh, idea of a reward, and so you're kind of weighting those, uh, circuits to approximate the reward that we are getting during training?

  985. 1:51:17

    Yes. So the-- [lips smack] Again, there's two schools of thought. The first one is like

  986. 1:51:23

    the, the question, what is two plus two? Somewhere in the model, somewhere, I don't know where, right, this high dimension of one point four trillion parameter space somewhere, it knows to calculate it as four.

  987. 1:51:34

    Right? There is some sort of circuit inside the model, and then the goal of RL is just, just to maximize this circuit somehow by via, well, by these formulas and stuff.

  988. 1:51:42

    You're just trying to maximize it. But that's one school of thought. The other school of thought is, oh, you know, RL's actually learning something. You know, like, it's actually learning how to do two plus two is equal to four, and it's not actually in the model.

  989. 1:51:54

    Thank you.

  990. 1:51:55

    Okay. Any other-- Yeah. Yes.

  991. 1:51:56

    When you say capabilities-

  992. 1:51:59

    Mm-hmm

  993. 1:51:59

    ... of, of a model that already has them inside-

  994. 1:52:02

    Yes

  995. 1:52:03

    ... you mean knowing actually the, the answer to a question or knowing how to reason to get to a question?

  996. 1:52:11

    That is a good question. Maybe both. I, I think it depends. It probably knows how to do the re...

  997. 1:52:18

    It might-- For example, a contrived example. You get all of the entire world's data, right? Like, you know, thirty trillion tokens, get all of the data, and you just make a question that is not part of the data, right?

  998. 1:52:29

    You, you could do that, right? What is ten billion, you know, some random number times some random nu-- You can make a math equation which is not in the data.

  999. 1:52:37

    But somehow the model has learnt to do multiplication, has learnt to do addition somewhere. And so maybe this circuit for addition, for multiplication, you know, for whatever, some, you know, many, many, many circuits of these, like, functions, we just wanna accentuate them all, and that is what kind of RL is trying to do.

  1000. 1:52:58

    If... And yes, there's a reasoning circuit. So like somehow the model also learns how to do reasoning, and so we also want to make that more important. And so, you know, two plus two is very important, addition is very important, multiplication is very important, and so on.

  1001. 1:53:09

    We're just trying to like make all of these circuits more prevalent. But that's only one half of the AI community, right? That's only one half thinks like that. The other half is like, "Oh, but the model's actually learning.

  1002. 1:53:19

    You know, we're actually training the model to learn, and the base model actually doesn't know how to do reasoning." If... Does that kind of-

  1003. 1:53:26

    Okay.

  1004. 1:53:27

    Okay. Yes. Okay. Hopefully. Okay. Any other questions? Yes.

  1005. 1:53:32

    So last question on that. So for, uh, RL framework, what frameworks do you suggest? Like, I, I see the notebook, we use, uh, a new base one, but there is also popular one, VERL, V-E-R-L.

  1006. 1:53:44

    Yes.

  1007. 1:53:45

    So which one-

  1008. 1:53:45

    There is VERL, there is TRL. Um-

  1009. 1:53:47

    Yeah

  1010. 1:53:47

    ... Unsloth, like we, we also make a... It's not called a framework. It's more like we showcase that you can do, uh, GRPO and reinforcement learning with very low resources.

  1011. 1:53:57

    So we are the only package which allows you to do GRPO on like a free colab. And so, like, that's the only difference between us and everyone else. So for VERL, VERL was very good for large training runs.

  1012. 1:54:08

    But for now, Unsloth, like if you want to do like small experimentation, you want to try stuff out, you don't know what reinforcement learning is, you don't know how to make reward functions, you don't even know what reward function to do, you should utilize our notebooks.

  1013. 1:54:20

    And that's what I was gonna demo. Um, yeah. Yes.

  1014. 1:54:25

    So one question is, uh, RL, like let's say you don't want to use a thinking model or do chain of thought. You just have a regular task, right?

  1015. 1:54:32

    Mm-hmm.

  1016. 1:54:33

    Like which say calls some tool, et cetera. So, uh, what are your, like, uh, if you've done any experiments on that, what is the, if, uh, effectiveness of that?

  1017. 1:54:41

    Like using RL to just improve, like, let's say, tool use for your use case.

  1018. 1:54:46

    Yes. Yes, you can do that. Exactly. It should increase accuracy by quite a bit. If your accuracy, if your accuracy before with tool use was not very good, RL should definitely help.

  1019. 1:54:54

    And if you're like R-- The trick of RL is it reduces overfitting. I think that's the trick because you do, you do multiple inferences. You don't know which one's correct, but you're trying to like maximize some good ones, right?

  1020. 1:55:05

    Some good ones. The problem with general fine-tuning is you're kind of like overfitting the model. And so the trick of u-- the trick of reinforcement fine-tuning is you can essentially reduce overfitting.

  1021. 1:55:16

    So the model actually learns how to do tool calling, not just, "Oh, you know, I see someone's trying to do a restaurant order. I just want to call DoorDash to do order," or something, right?

  1022. 1:55:26

    But it actually learns, "Oh, okay, because the person wants to order food, I should or-order DoorDash." It's like reverse thinking. Um-

  1023. 1:55:33

    Yeah.

  1024. 1:55:33

    So it should definitely help.

  1025. 1:55:34

    Yeah. Because in some of my experiments, like unless, like I try not explicitly asking the model, like a regular instruct model, right? Go and open it to, like, just do some task, right?

  1026. 1:55:45

    And it does not, or unless it, uh, it does not automatically start the thinking process unless you explicitly prompt it that, "Okay, first, uh, think and then answer."

  1027. 1:55:54

    Yes.

  1028. 1:55:54

    I was, I was trying to check if like j-just without explicitly asking it to, uh, start the thinking process, just to improve the tool use accuracy itself, does it start thinking?

  1029. 1:56:05

    And that, I think if that did not happen because it, uh, yeah.

  1030. 1:56:09

    You don't need to do... So GRPO, generally people utilize GRPO and like reinforcement learning algorithms to create the thinking process.

  1031. 1:56:15

    Mm-hmm.

  1032. 1:56:16

    That's because it's like a artifact of GRPO. Just by chance, they see the reasoning process. You don't need it. So maybe by chance, by luck-

  1033. 1:56:24

    Somehow it learns how to do tool, tool calling, and it's not some thinking process. It could be some weird symbols maybe. You know, I don't know. It could be like using some other ...

  1034. 1:56:33

    You know like sometimes models have like different languages suddenly? It could be like that. You know, randomly it learns how to do tool calling. It made some new programming language internally, I don't know, but it could have done that.

  1035. 1:56:43

    In this case there was no thinking process.

  1036. 1:56:46

    Yes.

  1037. 1:56:46

    It just directly gave the output because, like, let's say you sample like 10 trajectories, and none of them have a thinking process. Then it's never, it never explores those paths at all.

  1038. 1:56:54

    For now it will never. But remember, it's all about luck. Over time you will have a thinking process. Just by miraculous chance there was a thinking process somewhere, and then, "Oh, you should do this more," and it would just do this more.

  1039. 1:57:06

    But to make that more probable, like we should prompt it, uh-

  1040. 1:57:09

    You can prompt it. So essentially in the system prompt, you can say, "Please put your working out between this box." You could force the model to create the working out.

  1041. 1:57:19

    You could do that. Y- you could. Um, but is this the most efficient? I don't know. Y- you could. You could say, "Oh, please create a new language that I don't understand which does tool calling," and then it does some weird symbols and then, oh, it do- does tool calling.

  1042. 1:57:36

    I don't know. But yes, you, you should prompt it. It should make it more effective.

  1043. 1:57:40

    Mm-hmm.

  1044. 1:57:41

    Okay. Any other ... Yes. What's the secret sauce of Unsloth to make the training fast and, uh, efficient? Oh, we, we utilize Triton kernels. We do like kernel optimizations.

  1045. 1:57:52

    We reduce memory usage by 70%. Um, there are lots of optimizations that we do to make training faster, and more memory efficient. Can you talk about it later? Yes.

  1046. 1:58:00

    Later. Yes. Yes. It's not the main focus, but yes, we will talk about that. For vLLM specifically, we use ... Okay, that's actually the notebook. Oh, before that, do we have any other questions?

  1047. 1:58:09

    I, I ... Yes.

  1048. 1:58:11

    Real quick, like, with all the focus on test time scaling these days, and also papers, uh, claiming that even the test time reinforcement learning can get efficiencies in performance-

  1049. 1:58:19

    Mm-hmm

  1050. 1:58:20

    ... do you see any kind of reinforcement learning that kind of improves the test time learning capabilities in terms of-

  1051. 1:58:27

    The ... Actually, the DeepSeek paper talks about this. There is pass@K and majority@K. I think they said that if you do test time, test time scaling, I think it improves majority.

  1052. 1:58:37

    I, I think that was correct. You need to read the ... I'll have to reve- revert back to the DeepSeek paper. But they did say ... Remember, test time scaling is different from reinforcement learning.

  1053. 1:58:45

    They are different methodologies. Test time scaling is calling the model 10,000 times, and then, you know, then you just check, you know, by average what ... Okay, you, you ...

  1054. 1:58:55

    For example, you ask ChatGPT, "What is 2+2?" It might say 4. You know, it says 4, 4, 4, 4, 4, and suddenly it says 5. I don't know, by chance it says 5.

  1055. 1:59:03

    It says 0. And you just take the most likely answer. That's called test time, test time scaling.

  1056. 1:59:08

    Right.

  1057. 1:59:08

    And then reinforcement learning's kind of different. It's more like, oh, we want to like make ... We want to actually train the model to actually do the whole trace, and you don't, you don't need to do ...

  1058. 1:59:17

    You don't need to output 10,000 examples and get the best answer. You just do one.

  1059. 1:59:22

    There's also test time optimization. We have few, for example, kind of optimizing your, uh ... Things like DSPy that, that does, like, you know, prompt optimizations things like that, right?

  1060. 1:59:31

    So there's optimizations you can build on a test time-

  1061. 1:59:33

    Yes. Correct

  1062. 1:59:34

    ... as well.

  1063. 1:59:34

    Yes.

  1064. 1:59:35

    Is there anything that is feeding into the reinforcement learning methodology itself that improves the test-time optimization?

  1065. 1:59:41

    Yeah, so the trick is you do the RL step, and then you do test time compute. You will actually make the accuracy much higher.

  1066. 1:59:47

    There's no pre-time there.

  1067. 1:59:49

    You ... That's a good question. You could. Kind of GRPO's kind of like that, right? So GRPO you do test time scaling in, in the actual reward function. You literally call the model, "What is 2+2?"

  1068. 2:00:00

    Do test time scaling, and then you aggregate the results. So it's like GRPO itself is doing test time scaling internally. You could

  1069. 2:00:09

    ... That sounds like a new research paper. You could do that, I guess. Yes.

  1070. 2:00:13

    Okay. I will have to go to the notebook now. Um, so in order to access the notebook, you can go to our GitHub page, um,

  1071. 2:00:23

    which ... Until my internet actually loads. Um, so if you go to Unsloth, right, the GitHub page Unsloth, there is a button called Qwen3 GRPO, right? And you can click Start for free.

  1072. 2:00:33

    That's how you get the notebook. Or you can go to our docs, um, which have the notebook. Um, so remember, go to the GitHub page and then click Start for free, um, yes, for Qwen3 GRPO.

  1073. 2:00:43

    And then you will get this notebook. Um, generally it's dark mode, but I know in presentations, you know, presentations, I don't think so people can see that, so I will change this to li- light mode.

  1074. 2:00:55

    Um, so we utilize vLLM behind the hood. So vLLM ... Does, does anyone not know vLLM? I think that's a good question. Who does not know vLLM?

  1075. 2:01:09

    Okay. 100% you must use vLLM, right? So, like, for all open source, how do you serve a large language model? Please use vLLM or SGLang, um, or, you know, uh, I think Hugging Face has one as well.

  1076. 2:01:21

    So, like, these are the best open source libraries to serve open source models. Um, you know, like you have a GPU, how do we actually serve Llama 3? You know, how do we serve Llama 4?

  1077. 2:01:30

    You use vLLM to serve it. The trick of Unsloth is ... So Unsloth is a package for fine-tuning, for GRPO, for reinforcement learning, for whatever you like, continue pre-training, whatever.

  1078. 2:01:40

    And the trick is we just optimize it. We make it much faster, you know, use two ti- use, um, 70% less memory, um, make it fit on a free Colab.

  1079. 2:01:48

    Um, remember, please use free Colab resources. You know, Kaggle has 30 hours for ... I already said this again. Ka- Kaggle has 30 hours of, for free per week of GPUs.

  1080. 2:01:58

    Please utilize them. Um, they won't be unhappy. You know, please utilize them. Um, and yeah, so you install Unsloth and vLLM. Um, and we have this thing called the FastLanguageModel class, which essentially you can call a model.

  1081. 2:02:13

    Um, for example, we, we will now utilize the Qwen3 base model, right? So, like, remember I told you not to do this, right? But we are ... Anyways, we're gonna do it.

  1082. 2:02:21

    Um, this plot, where is the plot? Um-

  1083. 2:02:26

    We are going to do the dark blue dot to,

  1084. 2:02:31

    to the, um, the dark green. Um, yeah, that's what we're gonna do. [laughs] We're actually not gonna do what I suggested not to do. Um, but anyways, um,

  1085. 2:02:41

    we're gonna do that. Um, you also have to se- set a max sequence length. So for example, if you wanna make it longer, you can set it for longer, right?

  1086. 2:02:47

    If you want longer reasoning traces, you can increase the maximum sequence length. We set it as two thousand and forty-eight. Um, if you set it for larger, the free GPU will run out.

  1087. 2:02:57

    Um, so that's the problem. You can also load in four bits. So if you wanna do four-bit quantization, you can make the model go to four bits. You can reduce memory usage by quite a bit, so you can do that as well.

  1088. 2:03:08

    Um, and remem- we are utilizing LoRA, which is a parameter efficient fine-tuning method. Um, you don't need to fine-tune every single weight inside the entire model. Um, this will be very, very, very costly.

  1089. 2:03:19

    Instead, we add small weights to the model to fine-tune it. Um, and so that's the trick that we do.

  1090. 2:03:25

    And because we utilize vLLM directly, we do a trick. We actually re- reduce memory usage by fifty percent. The trick is we share vLLM's weights directly. Um, other training frameworks, what they do is they have to copy vLLM's weights, um, because you have to...

  1091. 2:03:39

    You have the model for fine-tuning, and you have the vLLM weights. The trick that we do is we actually share the vLLM weights directly, so you can reduce memory usage by a further fifty percent.

  1092. 2:03:50

    We use something called the unsloth gradient checkpointing, which reduces memory. Okay, essentially everything in AI is about reducing memory usage, more efficiency. You know, everything that we do is just efficiency.

  1093. 2:03:59

    Um, so everything that we set is for efficiency purposes. Um, your LoRA rank for, if you do LoRA, please set it to be the alpha to be two times the LoRA rank.

  1094. 2:04:08

    Um, it, it speeds up training dramatically, so please do that.

  1095. 2:04:12

    Some, lots of stuff, right, compiling. We do, like, automatic compiling and stuff like that. You don't need to read the, all of this. Um, and here is the bulk.

  1096. 2:04:20

    This is the most important part. Someone was asking about, you know, about a prompt. You make a system prompt, right? You are given a problem. Think about the problem and provide your working out.

  1097. 2:04:31

    Place it between reasoning start and reasoning end, right? So, like, the reasoning start is start working out and end working out. So it should look something like this. Um, if you, you know, it does, uh, start working out and end working out.

  1098. 2:04:47

    Right. You are given a problem, think about the problem, provide your working out. Place it between start working out and end working out. Then provide your solution between solution start and solution end.

  1099. 2:04:56

    So it should look something like this. Um,

  1100. 2:05:00

    and this is the system prompt that we're going to use for reinforcement learning. Remember, you can customize this to however you like, [snaps fingers] right? You don't have to say you are given a problem.

  1101. 2:05:08

    You are given a legal case, right? Think about the case and provide, provide your, I don't know, legal thinking. I don't even know. I'm just making stuff up. I don't know.

  1102. 2:05:19

    Whatever. Provide... Place it between, I don't know. It can be literally anything. S- What, um, thinking. I don't know. I don't... You can even make spelling mistakes. It doesn't really matter, right?

  1103. 2:05:31

    End thinking. The whole goal of RL is you can design your reward, you can design the system prompt to whatever you like, right? I think that's the main problem, is like people think, "Oh, you must follow DeepSeek think," right?

  1104. 2:05:42

    Think, right? You need to... People see in DeepSeek think, right, and </think>, right?

  1105. 2:05:50

    You do not need to follow, uh, you do not need to follow this, right? You do not need to follow this at all. You can make it up entirely, right?

  1106. 2:05:58

    This is customizable to whatever you like. The hard part is because we are using a base model, remember this is a base model, you have to make a chat template as well.

  1107. 2:06:09

    Um, this is the more annoying part. You can just copy and paste this chat template. You do not need to do anything else. Um, just, you know, literally copy and paste it.

  1108. 2:06:18

    Um, the base model does not have a chat template, right? A base model, when you call it, you can't actually call it for conversation. You can't... It's not ChatGPT, right?

  1109. 2:06:27

    It's just a base model. It doesn't do anything. So you need to specify a template for it to understand how to do conversations. And so this is kind of like a template that we did.

  1110. 2:06:36

    It's a very, it's very generic. You can just copy and paste this. It should be the same for anything.

  1111. 2:06:44

    And then we show... So after you, uh, do the chat template, we show an example of how to actually utilize the chat template and the tokenizer, right? So for example, if you ask what is one plus one, you do the reasoning process.

  1112. 2:06:54

    It will say, "You are given a problem. Think about the problem and provide your working out," blah, blah, blah, blah, blah. And then the question, this is the question.

  1113. 2:07:01

    "Your question is what is two plus two? Remember, the answer is four." And so start working out. This is what you give to the model. You give all of this to the entire model,

  1114. 2:07:12

    right? All of this, um... Okay, I can't really highlight. All of this you will give to, into the model, and the goal of RL is you want to create the working out process automatically, right?

  1115. 2:07:20

    The RL algorithm will automatically create the, uh, you know, the working, the working out or the thinking process. And then finally, it will say four. Well, hopefully it will say four.

  1116. 2:07:29

    And the goal is if you see the wor- if you see four, you want to make the reward higher just for that.

  1117. 2:07:35

    Now, someone was talking about, um, you know, fine-tuning with the, um, you know, instruct fine-tuning first. Again, remember, we go back to this diagram. We wanted to start from the blue dot to go to the green dot, but we found it doesn't actually work. [laughs]

  1118. 2:07:50

    So don't actually do this. The trick is we go back to this diagram. We actually want to take the pre-trained model, do some fine-tuning, do some supervised fine-tuning, and then go to the green dot.

  1119. 2:08:03

    And so this part we show that you should actually do some supervised fine-tuning. You need to do some fine-tuning to prime the model, right? The goal is you want to learn, you make the, you want to make the model not just output zero reward forever.

  1120. 2:08:15

    And so this dataset allows you to prime the model to do supervised fine-tuning.

  1121. 2:08:22

    So for example, the problem is, you know, what is the sum of all the real numbers, blah, blah, blah, blah, blah. And then you use DeepSeek-R1, so this is a trick, this is a hack.

  1122. 2:08:31

    You use DeepSeek-R1 to create some examples, and you shove this during the fine-tuning step. And essentially the model already learns how to do some reasoning, and that's the trick.

  1123. 2:08:40

    This data set is very small. It's only seven thousand rows, right? You don't need to have that much data for just this first step. You can have, like... I think I only used six hundred rows.

  1124. 2:08:48

    Very, very, very less data. So this is just data preparation step, not that important. Um, I need to skip to the reward function. This is the most important part of the model.

  1125. 2:08:58

    So this is the supervised fine-tuning step. So all of this is the supervise- supervised fine-tuning step. So this, this part, this part of the model, um, that's this part.

  1126. 2:09:08

    So not that important. Um, okay, we skip all of that. This is the most important part. The reward function creation is the most important part, and I feel like, you know, the majority of people, like, neglect this part.

  1127. 2:09:20

    It is the most hardest part to do. Um,

  1128. 2:09:25

    okay, let's see. Where is the reward function? Oh, here it is. Okay.

  1129. 2:09:29

    For example... Okay. No, no. Th- this one is a regular expression to match if your format is correct. For example, remember we have-- we asked the question, we asked the model to say, "Please put your working out between start working out and end working out."

  1130. 2:09:45

    This regular expression essentially, um, essentially rewards the model to see start w- working out and end working out. If it doesn't have this, you will actually penalize it, and so this is one reward function that I created.

  1131. 2:09:59

    For example, I give it an example. If you say, "Let me think, end working out,"

  1132. 2:10:05

    it extracts two. Yes, that's good, but remember, we force the model to say, we force the model, "You must generate the answer between this and this." And it successfully extracted two, so that's good.

  1133. 2:10:17

    But, you know, also sometimes a model might generate some random spaces. It's possible. You know, the model might not actually-- the model might not follow your exact format. We still try to match it, right?

  1134. 2:10:26

    Even if it generates extra spaces, we still successfully match the number two, so that's good.

  1135. 2:10:33

    This is a reward function, right? So essentially what we see is if it matches the format exactly, we add the score by three.

  1136. 2:10:41

    If not, we just put zero. And remember, this match format essentially matches regular expression. We have to create it by hand for matching the format.

  1137. 2:10:50

    This number does not have to be plus three. It can be plus three hundred, whatever you like. It can be plus one. I don't know. It can be anything that you like.

  1138. 2:10:57

    But I just found plus three to work fine. Um, so you can do whatever you like, um, yeah, anything. And remember, the score is zero if you don't see it.

  1139. 2:11:05

    You can also do minus one. For example, if it's else, right, if it's not good, you can also just score minus, right? Minus three. You can minus three points from it, so up to you.

  1140. 2:11:14

    You can design your reward function as whatever you like.

  1141. 2:11:18

    But remember, if the model, if the model output is not exactly following your format, we should still at least reward it a little bit, right? Otherwise, the reward would just be zero, zero, zero, zero, zero.

  1142. 2:11:30

    So the trick is, if we see a keyword, we plus one, and if we don't... Oh, sorry. Plus zero point five if you see the keyword. And if you don't see the keyword, right, if you see the keyword this, if you see this in the output, you should at least plus zero point five.

  1143. 2:11:45

    But if you don't see it, then you minus one, and so this essentially allows you to partially reward the model.

  1144. 2:11:55

    More reward functions, now it gets more complicated. This alle- this large reward function essentially allows you to calculate the distance-based scoring. For example, remember we said, um, over here, um...

  1145. 2:12:06

    Where is it? Um, this one, right? What is two plus two? Four is, is correct, but three is also a better answer than D, right? If you ri- if you output D, it's definitely wrong, right?

  1146. 2:12:18

    If you output five, it's okay, but it's wrong. And so this function essentially allows you to take the good answer... So, sorry. This is the guess divided by your true answer, and it's like a ratio, and this ratio essentially allows you to reward...

  1147. 2:12:34

    If your number is close to the actual answer, you give it higher reward. And if your, if your answer is very, very, very far off, then you penalize it by minusing reward, and so this essentially allows you to do that.

  1148. 2:12:48

    If it's exactly correct, you also add five points.

  1149. 2:12:53

    And so this, this is probably the mo- this is probably the most important reward function. Um, but this is only for maths. Um, for other, like, code and stuff like that, you have to create more reward functions.

  1150. 2:13:05

    Now we test if our reward functions actually work. Um, and yes, it extracts the numbers. Um, this is just format reward, sorry. This is just extracting the solution, and you can see that it extracted zero point three four.

  1151. 2:13:16

    It extracted this number, it extracted this, and extracted this. Um, if your reward function is not working very well, you probably did something wrong in the re- regular expression, so please, like, edit that.

  1152. 2:13:28

    And then this is helper functions. Oh, this is another, another reward function. If you see, if you see the number one two three comma four five six, we want to remove the comma because you can't actually convert this into Python.

  1153. 2:13:41

    So you wanna remove the comma. Um, and then if it's equal to the true answer, you plus three point five reward, and if it's not, then you minus one point five reward.

  1154. 2:13:53

    More data set preparation functions. Not that important. Um, and here is the meat of the code for training, right? We call vLLM. Top P is one point zero. Um, you don't-- It's not that im- you can probably-- It's probably not a good...

  1155. 2:14:07

    One point zero just means you're sampling the entire space, um, so that's good. You can set this to, like, zero point eight or something else, up to you, but I generally set it to be one point zero to be, like, full sampling of the entire space.

  1156. 2:14:18

    Min P is zero point one. I suggest people to use this because otherwise the model might, like, go into, like... It might do inference of, like, random outputs, so p-- you know, use zero point one.

  1157. 2:14:29

    And temperature. I did suggest people to increase temperature to zero point-- one point two, right? Or you can do one. Um, try to increase your temperature as much as possible.

  1158. 2:14:38

    The more temperature you increase, the model becomes very, very, um, creative, right? It, like, creates random outputs. If you increase the temperature too much, like, you know, two Your model will be like gibberish.

  1159. 2:14:49

    So probably don't do too large numbers. So I normally suggest one point zero, one point one, or one point two, or somewhere around there. You should utilize min P together, right?

  1160. 2:15:00

    You should utilize min P together with high temperature numbers. Um, there is a paper about using temperature one point five and zero point one min P. You should utilize that.

  1161. 2:15:11

    There are some o- some other things that we utilize. Um,

  1162. 2:15:15

    num generations is very important. I set this to be four. This number is, is this thing. Um, where is it?

  1163. 2:15:26

    This number is this. How many, how many, like, uh, how many roll outs do you wanna do? How many inference steps do you wanna do for GRPO? Right. We chose four, and so four will just means what is two plus two, it will create four, four options, and so that is this number.

  1164. 2:15:42

    If you increase this number too much, you will use much more memory. Um, you should increase this as mu- as much as possible, if you can. Um,

  1165. 2:15:50

    and there's like gradient-- uh, there's like batch size. We set this to be one. Um, the trick is the batch size times the gradient accumulation is equivalent most of the time.

  1166. 2:16:00

    Um, GRPO is not. But essentially, what this does is if we do one, it just means we're doing one... What is one, what is two plus two? If you set batch size to be three, then we shove all of these three examples together into one.

  1167. 2:16:16

    Generally, you should set batch size to be much larger. Um, the problem is if you set batch size to be too big, you're going to use more memory. So the trick is instead you do gradient accumulation, you set this to be sixteen.

  1168. 2:16:26

    That's the trick. Um, gradient accumulation, it essentially allows you to do addition of gradients over time, and you can skip using too much memory.

  1169. 2:16:36

    And then there's like evaluation. If you wanna do evaluation, there's some functions for that. Um, and then we shall see the training. Um,

  1170. 2:16:44

    you will get a large table of numbers during the training process. This took two hours and fifty-four minutes on a free Colab. Um, look at the reward column, right?

  1171. 2:16:53

    The reward column. Minus seven point five, minus five point five, minus five point five. All very bad. Oh, and then suddenly, plus thirteen. Just by chance, suddenly, it's plus thirteen.

  1172. 2:17:03

    Remember, GRPO, the trick is if you see this plus thirteen, let's maximize this even more.

  1173. 2:17:09

    And then it's... Oh, but then it didn't really work, so it goes back to minus seven point five, minus five point five, minus seven point five, and so on.

  1174. 2:17:16

    And then plus eleven. You see another good reward. We want to maximize this as well. And so so on, so on, so on, right? That, that's the trick of GRPO.

  1175. 2:17:25

    By luck, by chance, literally by luck, you will have good answers. By good answers, you wanna maximize this. And if you keep looking down, right, if you s- keep scrolling down, in the end-- Okay, I need to make a plot.

  1176. 2:17:38

    But in general, your reward will increase over time, right? Look, these are all positive numbers now, right? These are all positive numbers. Your minus numbers are getting less and less and less and less.

  1177. 2:17:46

    Um, I think if I can-- I don't know if I can plot this, but okay, I pr- I'll plot this later. But essentially, if you plot this over time, the reward will actually increase over time.

  1178. 2:17:56

    There is also other numbers, like completion length. So essentially, remember, when you use a reasoning model, the reasoning trace can be extremely long. So this column just tracks how long the reasoning process is.

  1179. 2:18:06

    Um, it's... Over time, in general, the reasoning length should get longer and longer, um, but sometimes not always the case. Um, yeah, not always the case. There is also another column called KL divergence.

  1180. 2:18:18

    This essentially tells you how far the model is, the, you know, the final model from the original model. And the larger the number, it means it's getting very, very, very far away from the original model.

  1181. 2:18:28

    Um, in general, this number should get bigger over time, in general. Um, sometimes it doesn't move, but you should make this number go as much, much higher as possible.

  1182. 2:18:36

    Um, and then we also f- we made separate reward functions. Each of those reward functions also has their own reward. Um, the most important one is the last column, right?

  1183. 2:18:48

    Or the second last column. These are the two numbers. So in, in RL, there is a problem. Most RL fr-- uh, most RL training runs just follow the format, and it doesn't actually learn.

  1184. 2:19:00

    So the format th- columns are not important. Do not look at the format columns. These are useless. You need to look at the last two columns. And if you look at the last two columns, right, rewards check, check numbers, this essentially checks if the output is good or bad.

  1185. 2:19:13

    You see that it's minus two point five, minus two point five, not very good, and then suddenly, three point five, right? Three point five is good. We want to maximize this.

  1186. 2:19:21

    If you keep looking, over time, in general, if you take the-- if you take like a rolling average, in general, the model gets better and better and better. Obviously, we only trained this for two hours and fifty minutes.

  1187. 2:19:33

    You know, if you train it for twenty days, it might actually do very well. Um, but remember, this is a free Colab GPU. Um, so in general, remember, so like the goal is, the goal of GRPO is suddenly we see a good answer with a good reward.

  1188. 2:19:50

    We want to maximize that. And that's the whole point of GRPO. And it, it's, it's nothing fancy. It's just like by luck we see it, and we just wanna maximize it.

  1189. 2:19:59

    We can also see some output from the model. Right, at the very beginning, right, at the very beginning of the model... Let's see. Um, where is it? An example.

  1190. 2:20:08

    Compute the number of positive integers that, that divide at least to a blah, blah, blah, some question. And then it does some reasoning trace. Remember, we already fine-tuned it a little bit.

  1191. 2:20:17

    So it does something, right? It does something. But the answer, it just goes on and blah, blah, blah. It just blah, blah, blah. It just keeps going on, blabbering on.

  1192. 2:20:25

    Um, but then if you look at the actual answer, um, where is it?

  1193. 2:20:29

    If you keep-- Okay, there's a lot. Um, we print out every single-- We print out a lot. Okay, it just keeps... Okay, whatever. It keeps going on and on and on.

  1194. 2:20:37

    Um, this is the output of the GRPO algorithm. And you will see over time, if you inspect this, you will see that the model actually gets better and better and better.

  1195. 2:20:45

    Um, for-- This is an example, right? Let's say we ask the model, what is the square root of one hundred and one, right? It's n- we don't just say, what is the square root of one hundred?

  1196. 2:20:53

    That's just ten. We say, what is the square root of one hundred and one?

  1197. 2:20:58

    If you do not train the model, this is what you get. It will say answers, education, math and arithmetic, what is the square root of one hundred and one, Wiki user.

  1198. 2:21:06

    Ooh, Wiki user. [laughs] This is what it actually will say, right? That's actually what it will say. Does, do... Where do you think this data comes from? Does anyone know?

  1199. 2:21:13

    Can you take a guess where do you think this data comes from? Probably Wikipedia, right? So if you ask the question, what is the square root of one hundred and one, it doesn't do anything, right?

  1200. 2:21:22

    Remember, this is the base model. The base model is useless, right? You're not gonna get... It's not gonna answer the question.

  1201. 2:21:29

    But after we do GRPO, right, what is the square root of one hundred and one? We ask the question again. It says, "Okay, so I need to find the square root of one hundred and one.

  1202. 2:21:39

    Hmm, let me think. I remember that the square roots of numbers between the perfect numbers are ra- rational, blah, blah, blah, blah, blah, blah, blah." Right? It just says blah, blah, whatever, some whatever, and it says solution, ten point...

  1203. 2:21:49

    Right here. Solution, 10.049875. I think that's correct. I don't know if that's correct, but probably it's very close. And so the goal of... So the whole point is GRPO allows you...

  1204. 2:22:01

    The GRPO algorithm produced all of this, right? This reasoning trace. In the olden days, you actually have to have a human write all of this, and then you have to fine-tune the model.

  1205. 2:22:14

    GRPO, you skip. You don't need to make this anymore. It's automatic. Right? That's the trick of GRPO and reinforcement learning. All of this reasoning phase is automatic, totally produced from nothing, and i- in the end, it gets a solution.

  1206. 2:22:28

    Question.

  1207. 2:22:28

    Yes.

  1208. 2:22:28

    You had that seven thousand rows-

  1209. 2:22:31

    Yes. That's the trick. If you do the base model... The trick of doing this is if you just do the base model, go into this phase, you will still get this, but it'll be too long.

  1210. 2:22:40

    I see.

  1211. 2:22:41

    Otherwise, you'll wait there for, like, twenty days. And for demonstration purposes in the Colab, you have to do the supervised fine-tuning step. That's the trick.

  1212. 2:22:47

    Okay.

  1213. 2:22:48

    Yeah. Y- yes.

  1214. 2:22:49

    What is the advantage of doing this seven thousand examples fine-tuning versus using a instruct model out of the box?

  1215. 2:22:56

    You can use an instruct. We actually have notebooks for that. So if you go to GRPO in general, um, we have notebooks for using instruct model. Okay, the internet's very slow.

  1216. 2:23:06

    Um, GRPO, we actually have other notebooks. For example, if you use Llama three point two three billion, that is using instruct model. You don't need to use a base model, but we showed that you can use a base model.

  1217. 2:23:17

    Um, yeah, you can use... I suggest people to use instruct. Um, you sh- probably shouldn't use base. It's all about efficiency as well. Um, yes.

  1218. 2:23:24

    Can you quickly touch on the KL divergence, so, like, using GRPO with beta equals zero?

  1219. 2:23:30

    Sorry, what? What did...

  1220. 2:23:32

    You mentioned the KL divergence-

  1221. 2:23:33

    Yes, yes

  1222. 2:23:34

    ... along with drifting away from the original.

  1223. 2:23:35

    Yes, correct.

  1224. 2:23:36

    Can you talk a little bit about, like, why using that or putting beta equal zero so you get, like, a non-KL divergence GRPO?

  1225. 2:23:44

    Yes. So the goal of KL divergence is you want the model not to stray too much away from the original model, right? KL divergence is... I shouldn't say distance, but KL divergence is like a distance between the, the current model, the current model that you're training, and the previous, very, very, very beginning of the model, right?

  1226. 2:24:02

    And so essentially, if the model is too far away, your KL divergence will be very large. Right? If you look at the plot... Uh, where is the table? Uh, I'll scroll up a bit.

  1227. 2:24:11

    Um, right. Which, which one is the KL divergence? Oh, here. This is the column, right? This column is a KL divergence, um, column. Over time, it should get larger and larger and larger over time, right?

  1228. 2:24:22

    The numbers should get larger and larger and larger because the model is straying away from the fine-tu- uh, the original model. If you set the beta to be zero, then you remove this term.

  1229. 2:24:32

    Yeah.

  1230. 2:24:33

    Maybe this might make the capabilities of the model more, maybe, because you're essentially f- not forcing the model to be as close as possible to the base model.

  1231. 2:24:42

    Active area of research. So yeah, some people might set it to zero, some people might not set it to zero. You know, I, I think it's zero point zero f- I think the default is zero point zero five or zero point zero three.

  1232. 2:24:52

    It's that.

  1233. 2:24:52

    Yeah.

  1234. 2:24:53

    Okay. Any other questions for... Yes.

  1235. 2:24:56

    You just took the, um, what you did, the s- the seven thousand rows or whatever. Um, where between the gibberish to, uh, the proper reasoning would it happen? What would the output look like for the same question if you did it right then before doing any GRPO?

  1236. 2:25:12

    So for fine-tuning... Actually, fine-tuning is actually very helpful already. Um, where is the loss? Where is the loss? The base model... So the base model already is very bad.

  1237. 2:25:20

    Um, if you do, uh... Here. There's a loss. There's a loss. This is using the fine-tuning step, the pre, the priming stage, right? You use, like, a data set to firstly prime the da- to prime the model.

  1238. 2:25:31

    The loss does decrease. Remember, if you see a loss of zero point six four, that's good. Um, if you see a loss higher than thirty, definitely something's wrong. Um, higher than three is very bad.

  1239. 2:25:40

    Um, you can see the loss definitely decreases over time.

  1240. 2:25:43

    Right.

  1241. 2:25:43

    So yes, doing the fine-tuning stage does teach the model a little bit to do reasoning, and it learns how to do some stuff. Also, a good, a very interesting, um, fact is we used DeepSeek-R1, some of the reasoning process, a- to do the fine-tuning step.

  1242. 2:26:01

    And interestingly, if you just call the model without doing GRPO, it kind of does reasoning already by doing seven thousand examples, right? It already says... Remember the question was, um, "What is two plus..."

  1243. 2:26:14

    Oh, wait. This is just a general question, right? It kind of learns how to do reasoning somewhat, but it's not perfect. And so the goal of GRPO is to forcibly make it perfect.

  1244. 2:26:24

    Okay, not perfect, but as much as possible. Um, so actually, the fine-tuning step already kind of learns a little bit.

  1245. 2:26:32

    Is the seven thousand the minimum examples that require the-

  1246. 2:26:37

    No, we... No. Actually, you don't need to use seven thousand. I think I only used...

  1247. 2:26:43

    I think I used one hundred and eighteen.

  1248. 2:26:45

    Oh.

  1249. 2:26:45

    It's... So it's, uh, uh, two training epochs. I think it's... Yeah, it's one hundred and eighteen. I only use one hundred and eighteen rows. You don't need to... You can use ten rows.

  1250. 2:26:55

    You can use... Yeah, use as less. You must use more than three rows, though, um, because when you do LoRA, the gradients become, are zero, so you must use more than three.

  1251. 2:27:03

    But l- but anything more than three is fine

  1252. 2:27:06

    Yeah, so even if you use 118, it does fine. Um, yeah.

  1253. 2:27:12

    Yes.

  1254. 2:27:12

    Are there any things to consider if you need to scale this training for-- to start with the bigger base models, like, uh, apart from, uh, taking a bigger GPU, are there any tricks or strategies?

  1255. 2:27:25

    Do you mean, like, a small model versus a big model? What's the difference?

  1256. 2:27:28

    Um, can you, uh, tell any tri-tricks if you want to do this training on a bigger base model?

  1257. 2:27:33

    Oh yeah, go ahead. You can, you can take the notebook. You will need a better GPU though. Take the notebook, edit,

  1258. 2:27:41

    edit this here. Not four, you can do 14 billion. Wait, I think there's a 14 billion, I think. Or is it 12 billion? I can't remember. You can do whatever you like.

  1259. 2:27:48

    You could even do, you know, Llama 3.3, 70 billion. I don't know, up to you. Um, do whatever you like, and... But the goal is, for a Colab demonstration, because it's a small GPU, I use, like, a small model.

  1260. 2:28:01

    Um, we actually have notebooks for free Colab which fits 14 billion. Um, so Phi-4 14 billion actually fits in a free Colab. Um, so you can do big models in a free Colab.

  1261. 2:28:11

    Um, Kaggle. Again, I said Kaggle, use Kaggle, free GPUs. Um, there are like no-- So essentially, this whole page has, like, notebooks. Uh, where's Kaggle? Uh, Kaggle. Kaggle has, like, notebooks for GPU as well.

  1262. 2:28:22

    Um, so you can do whatever you like for large models. Um, yeah.

  1263. 2:28:25

    Do you have examples on sampling and training, separating them on different machines?

  1264. 2:28:31

    Oh, do you mean, like, for vLLM rollouts, like... Oh, so the trick, what we do is you-- we co-locate. So you use the same machine for inference and fine-tuning, and the trick is you can reduce memory usage because you're sharing the vLLM weights.

  1265. 2:28:43

    So some other trainers, like VERL and TRL, you do have to put the inference on another server, and then your training is, like, a separate server, and they have to do communication.

  1266. 2:28:51

    We don't-- There is no communication for us. There is none. So we do s-- It's very close to asynchronous training. Nearly no delay in training.

  1267. 2:28:58

    But yes, we d-we don't support it yet, but we do plan to support, like, you know, larger training runs. Um, yeah.

  1268. 2:29:06

    Y-yes.

  1269. 2:29:07

    I have a quick question. [clears throat] What's the TPU, uh, roadmap look like for Unsloth?

  1270. 2:29:14

    Yes, people have asked that. No roadmap. [laughs] I don't know if we're gonna support it. It's a bit more complicated. You could use, I think, XLA. Like, they, they do have PyTorch converted down to TPU, so maybe it might work.

  1271. 2:29:24

    I don't know if it works. I've never tried it. Um,

  1272. 2:29:29

    maybe later. Um, yeah. Maybe later.

  1273. 2:29:32

    That's a valid answer.

  1274. 2:29:33

    Okay. Yeah. Yes.

  1275. 2:29:34

    Just a follow-up. Uh, what's the reason why we chose the base model instead of the instruct model?

  1276. 2:29:39

    Oh, you don't have to. You, you... The, the whole point, the, like, this thing, um,

  1277. 2:29:46

    uh, here, right? We chose the base model because you can show that you can do a base model going to the green dot. But then, unfortunately, in the Colab, we do have to do some supervised fine-tuning.

  1278. 2:29:57

    Otherwise, you'll wait there forever. The reward, again, will be zero, zero, zero, zero, zero. We just want to... Remember, all of AI is about efficiency and speed, so we just wanna showcase, okay, you do need to do the light blue step.

  1279. 2:30:09

    You still need to do the supervised fine-tuning step. Um, yeah, yeah.

  1280. 2:30:13

    So we need to construct model with APPs to do that?

  1281. 2:30:16

    Oh, no, you don't need to. You can take an instruct-- So the notebooks over here, so for example, if you go to the, uh,

  1282. 2:30:22

    the Llama 3.2 three billion notebook, we don't do any fine-tuning step at all. You skip directly because it's an instruct model already. It already learns how to do chat.

  1283. 2:30:31

    It already learns how to answer some questions. You can skip directly to GRPO, um, if it loads. Um, but yes, the notebooks... Oh, okay. You'll have to wait for it to load.

  1284. 2:30:41

    Um, whatever. [laughs] The internet's very bad. Um, it is loading. Um, yes. Any other questions?

  1285. 2:30:51

    Oh, over here.

  1286. 2:30:51

    Yes.

  1287. 2:30:52

    Over in the REINFORCE algorithm, you had the log probability of a state-- or no, of an action given a state.

  1288. 2:30:58

    Mm-hmm.

  1289. 2:30:58

    Is that happening inside the sampling model, or is that in the notebook happening?

  1290. 2:31:02

    Oh, the, the, the algorithm itself for GRPO? Oh, it's like behind the scenes. Like-

  1291. 2:31:06

    Yeah. Is that happening over in the G-GRPO trainer?

  1292. 2:31:10

    Yes, it's inside the training itself. Like, somewhere in the code, somewhere it does that.

  1293. 2:31:14

    It's figuring out, like, what is the probability of a given token versus all the tokens?

  1294. 2:31:18

    Oh, y- the calculation is inside the trainer. So, like, somewhere, you know, on the GPU, you're doing this calculation. But you do get the probabilities. Remember the language model, you get the probabilities already.

  1295. 2:31:27

    Yeah.

  1296. 2:31:27

    You just get the reward function, and you just wanna maximize it. So...

  1297. 2:31:31

    So do you, do you take the logits that come out of the large language model and turn them into pseudo probabilities, and then just assume that's the distribution?

  1298. 2:31:38

    I think so. Yes, that's correct.

  1299. 2:31:40

    Okay.

  1300. 2:31:40

    I think that's exponential of... Yes, I think that's correct. There is... If you go to, like, the code, there is, like, some derivation for it. But yes, you're correct.

  1301. 2:31:47

    Any... Oh, okay. The notebook loaded. But yes, there is another notebook which does the instruct here, right? There is the instruct model. And there is no fine-tuning step at all.

  1302. 2:31:57

    It just does the reward function, um, and stuff like that. Um, Will's notebook, for example, is also very good. So, like, if anyone wants to check other notebooks out, Will's notebook, um, also utilizes, I think, also the instruct model and then does GRPO.

  1303. 2:32:10

    Um, okay. The GR-- Okay, kind of time is running out, but okay, technically, the GRPO for-- portion is done. [laughs]

  1304. 2:32:20

    Oh, there's actually more portions. Um, I will have to breeze through them. There's only ten minutes left. Um, whoops. Um, any other... I, I will take questions at the very end.

  1305. 2:32:29

    Anyways, I'm gonna stay here anyways afterwards. Um,

  1306. 2:32:33

    quantization. We'll now sw-shift over to quantization. Um, so I don't know if you guys know about the DeepSeek-R1 1.58-bit quants that we did, um, but you can essentially download these models.

  1307. 2:32:44

    DeepSeek-R1 is seven hundred and thirty, I think seven hundred and thirty GB. You can quantize them down to one hundred and forty GB, um, without that much loss in accuracy.

  1308. 2:32:53

    Okay. There's obviously loss in accuracy, but the trick is you can quantize them down to be very small, and miraculously, they work.

  1309. 2:33:02

    Um, Llama 4 Scout, for example, right? You can't really see the accuracy plot. Um, the smallest number is eighty percent accuracy on MMLU five-shot, eighty percent. And the highest accuracy is eighty-one-point-something, right?

  1310. 2:33:15

    So it's actually only one percent difference. And the one on the left is a one-bit quant. It's very small. It's tiny in comparison to the full precision like float eight.

  1311. 2:33:27

    And so essentially you can make the model eight times smaller and you only decrease accuracy by one percent. Um, so that's very interesting. And so essentially we showcase that you can actually quantize the MoE layers, the mixture of experts layers, very heavily, but you must leave the attention layers, the shared experts, and other layers in higher precision,

  1312. 2:33:47

    and that's called-- that's what we call the dynamic quantization methodology.

  1313. 2:33:51

    Um, if you see there was like a benchmark of Llama 4 Scout, for example. Um, if you use a two-bit, a two-bit quant, it actually gets high accuracy than other, other full precision, um, providers, which is very interesting, right?

  1314. 2:34:07

    So like, for example, the two-bit quant gets, um, seventy-three percent accuracy, and then other inference providers get sixty-five percent accuracy, sixty-seven, right? There is a very large difference. And so the...

  1315. 2:34:18

    Okay, there is like some bugs in the models and there's maybe they quantize it incorrectly, but the goal is to show that if you quantize the model down to be very small bits, it still works.

  1316. 2:34:30

    We showcase this with an example. For example, if you take a vision model like, um, Qwen seven bill-- uh, Qwen two billion. Um, if you naively quantize all the layers to be four-bit, right?

  1317. 2:34:40

    You ask the model, "What does this image show?" It will say, "The image depicts a vibrant and colorful scene of a coastal sc- area," which is totally wrong, right?

  1318. 2:34:49

    The answer should be, "The image shows a train traveling on tracks," or something like that, right? If you quantize everything to four-bit, it's one point three six GB, but it's definitely bad.

  1319. 2:35:01

    So the trick is you must quantize some layers to be-- You must leave some layers to higher precision, and you only need to increase it by five hundred MB also to one point eight GB, and it works.

  1320. 2:35:12

    The image shows a train traveling on tracks. It suddenly works.

  1321. 2:35:16

    But the question is which layers do you not quantize? That's the question, right? Which layers? You could do an exhaustive search, right? You can check, oh, let's not quantize layer zero.

  1322. 2:35:25

    Let's check layer one, layer two. You're gonna check every single one, but it'll take forever. So definitely don't do that, right? You have seventy choose one or something. Seventy choose one plus seventy choose two.

  1323. 2:35:35

    Horrible, right? And remember, all of AI is about efficiency, so don't do that. The trick is you can check the activation quantization error and the weight quantization error, and you will see these large outliers.

  1324. 2:35:47

    For example, for Qwen, if you quantize the first few layers, it's extremely bad. So you must leave the first few layers not quantized. And also, this gigantic jump for the weight quantization error, this means you probably shouldn't quantize that layer as well.

  1325. 2:36:02

    There are some other plots that we show. For example, for Llama 3.2, it's interesting. All of these graphs are very different from each model. Um, you will notice Llama 3.2 has these weird, you know, continuous spikes.

  1326. 2:36:14

    Um, it's because they use attention, and then they put the attention back to the vision module. I think every single three-- I think it was every single three layers.

  1327. 2:36:21

    So every single three layers, it has these big jumps. Um, this means you should not quantize those layers.

  1328. 2:36:28

    Pixtral, for example, is also a different graph again. Um, Pixtral seems like you can't quantize many layers, unfortunately. Um, and so, like, the whole vision module must not be quantized.

  1329. 2:36:41

    There is a very popu-- uh, there is a very important paper talking about like, you know, why-- which layers you should quantize and which layers you should not quantize.

  1330. 2:36:47

    It's called the Super Weights paper. Um, you, you guys definitely should read that. Essentially, it says that in all language models, the first few layers of the down projection, there is a very, very, very important number in one of the numbers, one of the numbers in the models.

  1331. 2:37:02

    Very, very, very important. And you should never quantize it, ever.

  1332. 2:37:06

    But the trick is... The, the interesting finding is it's not actually a very large number. So there is a trend in the-- There is a trend in, um, language model space where people think that you should not quantize outliers.

  1333. 2:37:19

    The problem is, you know, these models have these big outliers. Like, suddenly in the model, there's like this big number, like three thousand, and if you quantize it, it essentially ruins the model.

  1334. 2:37:28

    But actually, this paper shows that it's not actually the outliers that are the problem. These numbers could be very small, and if you look at the plots, if you remove-- if you select these numbers, and if you make them zero,

  1335. 2:37:41

    the accuracy dr-- uh, the accuracy decreases dramatically. Um, and so, like, you should see, like, for example, if you remove one of the numbers, um, you know, the activation value totally decreases.

  1336. 2:37:52

    Very bad. They have very large activation values, and then if you remove them, it's very, very, very, very bad. Um, there is another trick that you can do. If you have a model that has seven billion parameters, make every single number go to zero.

  1337. 2:38:06

    The first parameter, make it go to zero. Check accuracy. The second number, make it go to zero. Check accuracy. The third number, go to zero. Check accuracy. You can do this seven billion times, and then you can see which number is the most important.

  1338. 2:38:17

    You could do that as well. Um, but remember, AI is about efficiency. Very not a good idea. Um,

  1339. 2:38:25

    more later research, for example, the new Blackwell chips, um, instead of doing a quantization to like one bit, two bit, three bit, four bit, NVIDIA chips also have this new architecture, this new format called FP4 or MXFP4.

  1340. 2:38:39

    Um, and essentially this is float four. Um, and float four is ver- it's, it's most likely going to be very used a lot in the future. Um, and then there's these like new formats for quantization as well, which, which essentially allows you to train models in very low precision.

  1341. 2:38:56

    I also made this plot, um, going from float thirty-two... So the question people always ask is, "Why is GPUs getting faster and faster and faster?" My take is actually this year is probably the last year you're gonna get a GPUs that is actually faster.

  1342. 2:39:08

    There is no more faster GPUs. Why? Because the majority of GPUs getting faster is because of numerical precision. From float thirty-two to float sixteen, you get five times faster.

  1343. 2:39:19

    Why? Why is it five times faster? Because, because, um, the calculation of faster is when you use transistors, it's the exponent plus the mantissa squared, right? In float 32, you have to use 23 numbers for the mantissa, and 23 squared is very large.

  1344. 2:39:36

    Float 16, you reduce the mantissa to 10, and that is why you get five times speed up from float 32 to float 16. It's because the number itself is getting smaller.

  1345. 2:39:44

    The representation inside the models for each of the weights is getting smaller. And then we moved from float 16 to be bfloat16. It is again maybe around two times faster than float 16.

  1346. 2:39:56

    We then have float 8. Um, float 8 is even more faster, um, you know, uses even less space. Um, but then there is a problem. Float 4, we get to float 4, and it's around two times faster than float 8, around.

  1347. 2:40:10

    Um, the problem is, what's next? Do we go to float 2, float 3, float 1? You know, I, I... There's-- You can't push anymore in terms of numerical precision.

  1348. 2:40:22

    There is not much more to go in terms of that space. And so, like, you can only get maybe 180 times faster than float 32, maybe 200 times faster.

  1349. 2:40:31

    But essentially, my take is float 4 might be the final float that's getting faster, you know, the final, um, precision, numerical precision. And in the future, GPUs are not gonna get faster.

  1350. 2:40:42

    Um, so maybe if, you know, people wanna buy Blackwell GPUs, that's probably, you should probably buy them. It's most likely not gonna get faster anymore. Um, that's kinda my take.

  1351. 2:40:51

    And also, for... Okay, I was gonna talk about kernels and stuff, but I don't think so I have enough time. You must use torch.compile. You know, every single function that you see, wrap it in torch.compile, try it out.

  1352. 2:41:02

    You know, like I, I always tell the PyTorch team, "Please make it by default." You know, definitely use torch.compile. Um, why? Because it makes your training faster sometimes, only sometimes, um, [chuckles] not all the time.

  1353. 2:41:14

    It reduces memory usage most of the time. If you see bugs, they'll probably fix it. Um,

  1354. 2:41:21

    but remember, torch.compile is not as easy as you think. You don't just do torch.compile the model. There is actually many options you can tune, right? I just listed a few options.

  1355. 2:41:31

    Um, this is, this is literally just a few. There is, like, ten more pages of options. I'm being serious, ten more pages you can tune. Imagine if you can, like, use torch com- torch.compile and tune every single one.

  1356. 2:41:43

    And that's why I highly suggest people to use torch.compile more effectively. Um, it's probably the biggest thing that can change your entire, you know, training run, um, make it more m- more memory efficient, make it faster.

  1357. 2:41:55

    So definitely look through this. Um, okay, so in general, yes, thank you. Definitely star us on GitHub. Um, join our Discord if you wanna have any questions on RL and stuff.

  1358. 2:42:05

    Um, we have a website as well. Um, and finally, we have stickers. Um, yes, there are some limited time stickers as well that we have somewhere, I think over there.

  1359. 2:42:15

    Um, and remember, if you have any questions, I'm still gonna stay around and ask. Um, yeah. Thanks a lot. [applause] [upbeat music]