DCAI
← 返回全部动态
arXiv 机器学习规则精选09月25日 12:00

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

arXiv:2609.28737v1 Announce Type: new Abstract: Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal actio

阅读 arXiv 机器学习 原文 ↗