英文原文节选与中文翻译

Can we allocate more compute to the harder problems?

我们能否把更多计算资源分配给更难的问题?

If your GRPO group has all wrong completions, sample more with probability P (~0.9).

如果一个 GRPO 分组中的回答全部错误,就以概率 P(约 0.9)继续采样。

两条原帖的完整中文翻译

扩展强化学习时有一个基本思路:我们能否把更多计算资源分配给更难的问题?

我们做了这样的实验:如果一个 GRPO 分组中的回答全部错误,就以概率 P(约 0.9)继续采样,目标是找到更多具有非零梯度的 GRPO 批次。结果有效。这个方法叫作 Never Give Up(NGU)。

很高兴这篇论文终于发布了,也很意外之前似乎没人继续推进这个想法。它在直觉上很容易扩展,但要让它在实践中真正运行起来并不简单。也祝贺 Michael——他是我的第一位实习生;配套博客也写得很好。

原帖引用线程的技术上下文

论文把一种现象称为 LLM 强化学习中的“马太效应”:标准 RL 的提升主要集中在初始模型已经较容易解决的问题上,困难问题改善较少。单纯把每个提示词的采样数从 k=4 增加到 k=32,按相同计算量比较时,在最难问题子集上反而更差。

NGU 从较小的 k 开始。如果首批回答全部错误,就以概率 p 重新加入该提示词并追加采样;若找到正确答案,再使用累计样本训练。引用线程给出的实验包括:

  • 数学任务:在 DeepScaler 与 Qwen 3 4B 设置中,NGU 优于标准 GRPO,提升主要出现在较难的 AIME 与 BRUMO 问题上,同时没有明显损害容易问题。
  • 代码任务:在 Manufactoria 中,标准 GRPO 会停留在反复解决和失败于中等难度测试;NGU 则继续推进更难测试,直至完整解决部分编程问题。
  • 已披露边界:如果任务几乎全部由极难问题组成,NGU 通过快速过滤容易问题来重新分配计算的优势可能有限;多轮异步采样也会带来旧样本和 off-policy 数据问题。

来源

Nathan Lambert 主帖:Nathan Lambert on X: "An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! Called "Never Give Up"" / X
Nathan Lambert 续帖:Nathan Lambert on X: "Excited this paper is out (and surprised I haven't seen someone pursue this idea)! It's very nice because it seems intuitively very scalable, but it's not trivial to get it working in practice. Congrats to Michael (my 1st intern). Great blog too: https://t.co/hFiLHucLua" / X
引用的原始线程:Michael Noukhovitch on X: "Is RL actually making your LLM better? Gains from RL are mostly on easy questions🤯 We're calling this the Matthew Effect for RL on LLMs. We then leverage async RL to solve harder problems by Never Giving Up! paper https://t.co/1C2xjunbWc blog https://t.co/X4H83wKgXl … / X
论文:[2609.13443] Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
技术博客:Michael Noukhovitch - Blog