General Discounting versus Average Reward
| dc.creator | Hutter, Marcus | |
| dc.date | 2006-05-09 | |
| dc.date.accessioned | 2026-07-07T07:09:30Z | |
| dc.date.available | 2026-07-07T07:09:30Z | |
| dc.description | Consider an agent interacting with an environment in cycles. In every interaction cycle the agent is rewarded for its performance. We compare the average reward U from cycle 1 to m (average value) with the future discounted reward V from cycle k to infinity (discounted value). We consider essentially arbitrary (non-geometric) discount sequences and arbitrary reward sequences (non-MDP environments). We show that asymptotically U for m->infinity and V for k->infinity are equal, provided both limits exist. Further, if the effective horizon grows linearly with k or faster, then existence of the limit of U implies that the limit of V exists. Conversely, if the effective horizon grows linearly with k or slower, then existence of the limit of V implies that the limit of U exists. | |
| dc.description | 17 pages, 1 table | |
| dc.identifier | https://arxiv.org/abs/cs/0605040 | |
| dc.identifier | http://arxiv.org/abs/cs/0605040 | |
| dc.identifier | Proc. 17th International Conf. on Algorithmic Learning Theory (ALT 2006) pages 244-258 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/111086 | |
| dc.subject | Machine Learning | |
| dc.title | General Discounting versus Average Reward | |
| dc.type | text |