英文原文

Very few will care about the actual reinforcement learning details here, but watching me grope around for understanding in the chat log might be of broader interest: https://chatgpt.com/share/6aad4252-f308-83ea-9783-6b4dc9cd3107

I have often noticed that our estimated Q values are higher than the observed returns, sometimes substantially so, and that can’t be good for performance.

For policy decisions, only the relative values matter, but an offset negatively impacts bootstrapping, so I was excited to see this paper on Relative Value Learning:
https://arxiv.org/pdf/1901.09732

After struggling a bit with “the Banach space of bounded antisymmetric pairwise functions“, I realized all that is essentially going on is subtracting two value functions. Given this framing, Chat was able to simplify the machinery into a particularly elegant form:

Subtracting the mean TD error from the individual sample TD errors gives all the benefits of relative value learning. Summarized as “relative Bellman regression is TD-error centering”.

Unfortunately, while you can create examples where this should be very valuable (someone should write a proper paper on it!), it didn’t actually improve performance on my tasks.

However, I think this has usefully narrowed down what is actually happening. Bellman iteration naturally corrects towards a correct absolute value, but only when the bootstrap values are also training targets. With Q-learning, you are often / mostly bootstrapping from a max-action that was not actually taken, so it never sees any downward Bellman pressure.

This is consistent with another result I have noted: learning state value functions offline from a frozen policy without actions doesn’t seem to suffer from any value overestimation.

中文翻译

这里真正关心强化学习细节的人可能很少,但看着我在聊天记录里摸索着理解这个问题,或许会让更多人感兴趣:https://chatgpt.com/share/6aad4252-f308-83ea-9783-6b4dc9cd3107

我经常注意到,我们估计出的 Q 值高于实际观察到的回报,有时甚至高出很多;这不可能有利于性能。

对于策略决策而言,真正重要的只有相对值,但一个整体偏移会对自举产生负面影响。因此,当我看到这篇关于相对价值学习的论文时非常兴奋:
https://arxiv.org/pdf/1901.09732

在“有界反对称成对函数的 Banach 空间”这个概念上挣扎了一阵之后,我意识到,这里发生的一切本质上只是两个价值函数相减。基于这种理解,Chat 把整套机制简化成了一种格外优雅的形式:

从每个样本的 TD 误差中减去平均 TD 误差,就能获得相对价值学习的全部好处。概括来说,就是“相对 Bellman 回归就是 TD 误差中心化”。

遗憾的是,虽然可以构造出一些例子,让这种方法看起来应该非常有价值(应该有人为此认真写一篇论文!),但它实际上并没有改善我这些任务的性能。

不过,我认为这次尝试确实帮助我缩小了问题范围,更清楚地看到了实际发生的事情。Bellman 迭代会自然地把估计修正到正确的绝对值,但前提是用于自举的值本身也被当作训练目标。使用 Q-learning 时,你经常——或者说大多数时候——是从一个并未真正执行的最大值动作上进行自举,因此它从来不会受到向下修正的 Bellman 压力。

这也和我观察到的另一个结果一致:如果从一个冻结的策略出发,在没有动作数据的情况下离线学习状态价值函数,似乎不会出现价值高估。