Foundations of Reinforcement Learning and Interactive Decision Making

这些讲义从统计学的角度介绍了强化学习和交互式决策制定的基础。我们提出了一个统一的框架来解决探索-利用困境,使用频率主义和贝叶斯方法,并以监督学习/估计和决策制定之间的联系和类比为主题。特别关注函数逼近和灵活的模型类,例如神经网络。涵盖的主题包括多臂赌博机和情境赌博机,结构化赌博机以及具有高维反馈的强化学习。
These lecture notes give a statistical perspective on the foundations of reinforcement learning and interactive decision making. We present a unifying framework for addressing the exploration-exploitation dilemma using frequentist and Bayesian approaches, with connections and parallels between supervised learning/estimation and decision making as an overarching theme. Special attention is paid to function approximation and flexible model classes such as neural networks. Topics covered include multi-armed and contextual bandits, structured bandits, and reinforcement learning with high-dimensional feedback.
许愿