A Unified View of TD Algorithms; Introducing Full-Gradient TD and Equi-Gradient Descent TD
| dc.creator | Loth, Manuel | |
| dc.creator | Preux, Philippe | |
| dc.date | 2006-11-29 | |
| dc.date.accessioned | 2026-07-07T07:31:47Z | |
| dc.date.available | 2026-07-07T07:31:47Z | |
| dc.description | This paper addresses the issue of policy evaluation in Markov Decision Processes, using linear function approximation. It provides a unified view of algorithms such as TD(lambda), LSTD(lambda), iLSTD, residual-gradient TD. It is asserted that they all consist in minimizing a gradient function and differ by the form of this function and their means of minimizing it. Two new schemes are introduced in that framework: Full-gradient TD which uses a generalization of the principle introduced in iLSTD, and EGD TD, which reduces the gradient by successive equi-gradient descents. These three algorithms form a new intermediate family with the interesting property of making much better use of the samples than TD while keeping a gradient descent scheme, which is useful for complexity issues and optimistic policy iteration. | |
| dc.identifier | https://arxiv.org/abs/cs/0611145 | |
| dc.identifier | http://arxiv.org/abs/cs/0611145 | |
| dc.identifier | Dans European Symposium on Artificial Neural Networks (2006) | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/118881 | |
| dc.subject | Machine Learning | |
| dc.title | A Unified View of TD Algorithms; Introducing Full-Gradient TD and Equi-Gradient Descent TD | |
| dc.type | text |