Policy Iteration for Factored MDPs

Daphne Koller; Ron Parr

arXiv:1301.3869·cs.AI·January 18, 2013·152 cites

Policy Iteration for Factored MDPs

Daphne Koller, Ron Parr

PDF

Open Access

TL;DR

This paper introduces a new policy iteration method for factored Markov Decision Processes (MDPs) that efficiently computes approximate value functions and improves policies using a least-squares approach, with error bounds and complexity analysis.

Contribution

It presents a novel closed-form value determination algorithm for factored MDPs, enabling efficient policy iteration with error bounds and complexity analysis.

Findings

01

Efficient policy iteration for factored MDPs using least-squares value approximation.

02

Policies can be represented and manipulated compactly under certain restrictions.

03

Provides a method to compute error bounds for decomposed value functions.

Abstract

Many large MDPs can be represented compactly using a dynamic Bayesian network. Although the structure of the value function does not retain the structure of the process, recent work has shown that value functions in factored MDPs can often be approximated well using a decomposed value function: a linear combination of <I>restricted</I> basis functions, each of which refers only to a small subset of variables. An approximate value function for a particular policy can be computed using approximate dynamic programming, but this approach (and others) can only produce an approximation relative to a distance metric which is weighted by the stationary distribution of the current policy. This type of weighted projection is ill-suited to policy improvement. We present a new approach to value determination, that uses a simple closed-form computation to directly compute a least-squares decomposed…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsReinforcement Learning in Robotics · Machine Learning and Algorithms · Water resources management and optimization