Sparse Q-learning with Mirror Descent

Sridhar Mahadevan; Bo Liu

arXiv:1210.4893·cs.LG·October 19, 2012·21 cites

Sparse Q-learning with Mirror Descent

Sridhar Mahadevan, Bo Liu

PDF

Open Access

TL;DR

This paper introduces a novel reinforcement learning framework using mirror descent, enabling efficient sparse solutions and improved convergence through Bregman divergences and proximal-gradient methods.

Contribution

It proposes a new class of sparse mirror-descent RL algorithms that find sparse fixed points efficiently, advancing the use of Bregman divergences in reinforcement learning.

Findings

01

Sparse mirror-descent methods outperform previous approaches in computational efficiency.

02

The proposed algorithms effectively find sparse solutions in high-dimensional RL problems.

03

Experimental results demonstrate success in discrete and continuous MDPs.

Abstract

This paper explores a new framework for reinforcement learning based on online convex optimization, in particular mirror descent and related algorithms. Mirror descent can be viewed as an enhanced gradient method, particularly suited to minimization of convex functions in highdimensional spaces. Unlike traditional gradient methods, mirror descent undertakes gradient updates of weights in both the dual space and primal space, which are linked together using a Legendre transform. Mirror descent can be viewed as a proximal algorithm where the distance generating function used is a Bregman divergence. A new class of proximal-gradient based temporal-difference (TD) methods are presented based on different Bregman divergences, which are more powerful than regular TD learning. Examples of Bregman divergences that are studied include p-norm functions, and Mahalanobis distance based on the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsModel Reduction and Neural Networks · Sparse and Compressive Sensing Techniques · Stochastic Gradient Optimization Techniques