Combinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal
| dc.creator | Ortner, Ronald | |
| dc.date | 2005-08-17 | |
| dc.date.accessioned | 2026-07-07T06:20:08Z | |
| dc.date.available | 2026-07-07T06:20:08Z | |
| dc.description | We show that combinations of optimal (stationary) policies in unichain Markov decision processes are optimal. That is, let M be a unichain Markov decision process with state space S, action space A and policies π_j^*: S -> A (1\leq j\leq n) with optimal average infinite horizon reward. Then any combination πof these policies, where for each state i in S there is a j such that π(i)=π_j^*(i), is optimal as well. Furthermore, we prove that any mixture of optimal policies, where at each visit in a state i an arbitrary action π_j^*(i) of an optimal policy is chosen, yields optimal average reward, too. | |
| dc.description | 9 pages | |
| dc.identifier | https://arxiv.org/abs/math/0508319 | |
| dc.identifier | http://arxiv.org/abs/math/0508319 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/95250 | |
| dc.subject | Combinatorics | |
| dc.subject | Discrete Mathematics | |
| dc.subject | Machine Learning | |
| dc.subject | Optimization and Control | |
| dc.subject | Probability | |
| dc.subject | 90C40 | |
| dc.title | Combinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal | |
| dc.type | text |