DCAI
← 返回全部动态
arXiv 机器学习规则精选09月25日 12:00

Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes

arXiv:2609.28792v1 Announce Type: new Abstract: Robust average-reward Markov decision processes provide a fundamental framework for long-term performance optimization under uncertainty, and can have optimal long-run rewards that depend on the initial state. This state dependence requires a vector Bellman theory that accounts for both recurrent-class rewards and transition uncertainty. We develop such a theory for finite models with compact, post-action $(s,a)$-rectangular ambiguity. A gain-first, bias-second optimization principle yields a coupled vector gain-bias system, and every finite solution identifies the optimal robust gain and suppli

阅读 arXiv 机器学习 原文 ↗