A Dantzig Selector Approach to Temporal Difference Learning - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2012

A Dantzig Selector Approach to Temporal Difference Learning

Bruno Scherrer
Alessandro Lazaric
Mohammad Ghavamzadeh
  • Fonction : Auteur
  • PersonId : 868946

Résumé

LSTD is one of the most popular reinforcement learning algorithms for value function approximation. Whenever the number of samples is larger than the number of features, LSTD must be paired with some form of regularization. In particular, L1-regularization methods tends to perform feature selection by promoting sparsity and thus they are particularly suited in high-dimensional problems. Nonetheless, since LSTD is not a simple regression algorithm but it solves a fixed-point problem, the integration with L1-regularization is not straightforward and it might come with some drawbacks (see e.g., the P-matrix assumption for LASSO-TD). In this paper we introduce a novel algorithm obtained by integrating LSTD with the Dantzig Selector. In particular, we investigate the performance of the algorithm and its relationship with existing regularized approaches, showing how it overcomes some of the drawbacks of existing solutions.
Fichier non déposé

Dates et versions

hal-00749480 , version 1 (07-11-2012)

Identifiants

  • HAL Id : hal-00749480 , version 1

Citer

Matthieu Geist, Bruno Scherrer, Alessandro Lazaric, Mohammad Ghavamzadeh. A Dantzig Selector Approach to Temporal Difference Learning. ICML-12, Jun 2012, Edinburgh, United Kingdom. pp.1399-1406. ⟨hal-00749480⟩
348 Consultations
0 Téléchargements

Partager

Gmail Facebook X LinkedIn More