DCAI
← 返回全部动态
arXiv 机器学习规则精选09月24日 12:00

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

arXiv:2609.27033v1 Announce Type: new Abstract: Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a determinist

阅读 arXiv 机器学习 原文 ↗