Combinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal

dc.creatorOrtner, Ronald
dc.date2005-08-17
dc.date.accessioned2026-07-07T06:20:08Z
dc.date.available2026-07-07T06:20:08Z
dc.descriptionWe show that combinations of optimal (stationary) policies in unichain Markov decision processes are optimal. That is, let M be a unichain Markov decision process with state space S, action space A and policies π_j^*: S -> A (1\leq j\leq n) with optimal average infinite horizon reward. Then any combination πof these policies, where for each state i in S there is a j such that π(i)=π_j^*(i), is optimal as well. Furthermore, we prove that any mixture of optimal policies, where at each visit in a state i an arbitrary action π_j^*(i) of an optimal policy is chosen, yields optimal average reward, too.
dc.description9 pages
dc.identifierhttps://arxiv.org/abs/math/0508319
dc.identifierhttp://arxiv.org/abs/math/0508319
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/95250
dc.subjectCombinatorics
dc.subjectDiscrete Mathematics
dc.subjectMachine Learning
dc.subjectOptimization and Control
dc.subjectProbability
dc.subject90C40
dc.titleCombinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal
dc.typetext

Files

Collections