英文原文
Inference scaling part 1.
Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by >2x)
00:00 Introduction and recap
00:31 Training-time and inference-time scaling
07:52 What we’ll implement
11:47 Notebook setup and model loading
17:43 Building a flexible text generation function
24:40 Chain-of-thought prompting
28:26 Sampling and output diversity
33:43 Next-token logits and greedy decoding
38:20 Temperature scaling step by step
42:46 Softmax and token probabilities
47:42 Multinomial sampling
54:51 Adding temperature sampling to text generation
59:31 Top-p filtering step by step
1:10:23 Adding top-p filtering to text generation
1:13:43 Sampling and LLM watermarking
1:16:01 Self-consistency and majority voting
1:20:36 Implementing self-consistency
1:29:02 MATH-500 results
1:35:01 Accuracy and compute tradeoffs
1:36:50 Next steps and self-refinement
And a link to the video on YouTube: https://youtu.be/t5y-kS9nNxU
中文翻译
推理时扩展,第一部分。
我们从一个经过改造的文本生成函数开始,加入温度缩放、Top-p 过滤和多项式采样,用它生成多样化输出,以实现自一致性和 Best-of-N(将答案准确率提升到 2 倍以上)。
00:00 介绍与回顾
00:31 训练时扩展与推理时扩展
07:52 我们将实现什么
11:47 Notebook 设置与模型加载
17:43 构建灵活的文本生成函数
24:40 思维链提示
28:26 采样与输出多样性
33:43 下一词元 logits 与贪心解码
38:20 逐步讲解温度缩放
42:46 Softmax 与词元概率
47:42 多项式采样
54:51 在文本生成中加入温度采样
59:31 逐步讲解 Top-p 过滤
1:10:23 在文本生成中加入 Top-p 过滤
1:13:43 采样与大语言模型水印
1:16:01 自一致性与多数投票
1:20:36 实现自一致性
1:29:02 MATH-500 结果
1:35:01 准确率与计算量之间的权衡
1:36:50 后续步骤与自我精炼
以及 YouTube 视频链接:https://youtu.be/t5y-kS9nNxU