DCAI
← 返回全部动态
arXiv 机器学习规则精选09月24日 12:00

On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

arXiv:2609.26918v1 Announce Type: new Abstract: Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving sample-efficiency in Reinforcement Learning (RL). A natural extension to preference-conditioned multi-objective RL (MORL) relabels transitions with the preference direction the agent achieved rather than the one asked for. We show that this extension is frequently harmful: across four preference-conditioned off-policy algorithms spanning two critic backbones and two preference-sampling schemes on the continuous-control MO-Gymnasium suite, it degra

阅读 arXiv 机器学习 原文 ↗