Self-Optimizing and Pareto-Optimal Policies in General Environments based on Bayes-Mixtures
| dc.creator | Hutter, Marcus | |
| dc.date | 2002-04-17 | |
| dc.date.accessioned | 2026-07-07T03:18:20Z | |
| dc.date.available | 2026-07-07T03:18:20Z | |
| dc.description | The problem of making sequential decisions in unknown probabilistic environments is studied. In cycle $t$ action $y_t$ results in perception $x_t$ and reward $r_t$, where all quantities in general may depend on the complete history. The perception $x_t$ and reward $r_t$ are sampled from the (reactive) environmental probability distribution $μ$. This very general setting includes, but is not limited to, (partial observable, k-th order) Markov decision processes. Sequential decision theory tells us how to act in order to maximize the total expected reward, called value, if $μ$ is known. Reinforcement learning is usually used if $μ$ is unknown. In the Bayesian approach one defines a mixture distribution $ξ$ as a weighted sum of distributions $ν\in\M$, where $\M$ is any class of distributions including the true environment $μ$. We show that the Bayes-optimal policy $p^ξ$ based on the mixture $ξ$ is self-optimizing in the sense that the average value converges asymptotically for all $μ\in\M$ to the optimal value achieved by the (infeasible) Bayes-optimal policy $p^μ$ which knows $μ$ in advance. We show that the necessary condition that $\M$ admits self-optimizing policies at all, is also sufficient. No other structural assumptions are made on $\M$. As an example application, we discuss ergodic Markov decision processes, which allow for self-optimizing policies. Furthermore, we show that $p^ξ$ is Pareto-optimal in the sense that there is no other policy yielding higher or equal value in {\em all} environments $ν\in\M$ and a strictly higher value in at least one. | |
| dc.description | 15 pages | |
| dc.identifier | https://arxiv.org/abs/cs/0204040 | |
| dc.identifier | http://arxiv.org/abs/cs/0204040 | |
| dc.identifier | Proceedings of the 15th Annual Conference on Computational Learning Theory (COLT-2002) 364-379 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/31069 | |
| dc.subject | Artificial Intelligence | |
| dc.subject | Machine Learning | |
| dc.subject | Optimization and Control | |
| dc.subject | Probability | |
| dc.subject | I.2 | |
| dc.title | Self-Optimizing and Pareto-Optimal Policies in General Environments based on Bayes-Mixtures | |
| dc.type | text |