英文原文

Inference scaling part 1.

Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by >2x)

00:00 Introduction and recap

00:31 Training-time and inference-time scaling

07:52 What we’ll implement

11:47 Notebook setup and model loading

17:43 Building a flexible text generation function

24:40 Chain-of-thought prompting

28:26 Sampling and output diversity

33:43 Next-token logits and greedy decoding

38:20 Temperature scaling step by step

42:46 Softmax and token probabilities

47:42 Multinomial sampling

54:51 Adding temperature sampling to text generation

59:31 Top-p filtering step by step

1:10:23 Adding top-p filtering to text generation

1:13:43 Sampling and LLM watermarking

1:16:01 Self-consistency and majority voting

1:20:36 Implementing self-consistency

1:29:02 MATH-500 results

1:35:01 Accuracy and compute tradeoffs

1:36:50 Next steps and self-refinement

And a link to the video on YouTube: https://youtu.be/t5y-kS9nNxU

中文翻译

推理时扩展,第一部分。

我们从一个经过改造的文本生成函数开始,加入温度缩放、Top-p 过滤和多项式采样,用它生成多样化输出,以实现自一致性和 Best-of-N(将答案准确率提升到 2 倍以上)。

00:00 介绍与回顾

00:31 训练时扩展与推理时扩展

07:52 我们将实现什么

11:47 Notebook 设置与模型加载

17:43 构建灵活的文本生成函数

24:40 思维链提示

28:26 采样与输出多样性

33:43 下一词元 logits 与贪心解码

38:20 逐步讲解温度缩放

42:46 Softmax 与词元概率

47:42 多项式采样

54:51 在文本生成中加入温度采样

59:31 逐步讲解 Top-p 过滤

1:10:23 在文本生成中加入 Top-p 过滤

1:13:43 采样与大语言模型水印

1:16:01 自一致性与多数投票

1:20:36 实现自一致性

1:29:02 MATH-500 结果

1:35:01 准确率与计算量之间的权衡

1:36:50 后续步骤与自我精炼

以及 YouTube 视频链接:https://youtu.be/t5y-kS9nNxU