On the expected total reward with unbounded returns for Markov decision   processes

Fran\c{c}ois Dufour; Alexandre Genadot

arXiv:1712.07874·math.PR·October 8, 2018

On the expected total reward with unbounded returns for Markov decision processes

Fran\c{c}ois Dufour, Alexandre Genadot

PDF

Open Access

TL;DR

This paper investigates optimal strategies in Markov decision processes with unbounded rewards, establishing existence results under broad conditions and introducing a new approach for weak convergence of probability measures.

Contribution

It provides the first general existence theorem for optimal strategies in MDPs with unbounded rewards and non-compact action sets, using a novel weak convergence characterization.

Findings

01

Existence of optimal strategies under broad conditions

02

New characterization for weak convergence of probability measures

03

Illustrative examples demonstrating theoretical results

Abstract

We consider a discrete-time Markov decision process with Borel state and action spaces. The performance criterion is to maximize a total expected {utility determined by unbounded return function. It is shown the existence of optimal strategies under general conditions allowing the reward function to be unbounded both from above and below and the action sets available at each step to the decision maker to be not necessarily compact. To deal with unbounded reward functions, a new characterization for the weak convergence of probability measures is derived. Our results are illustrated by examples.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsEconomic theories and models