Supervised Learning for Stochastic Optimal Control

Vince Kurtz; Joel W. Burdick

arXiv:2409.05792·eess.SY·September 10, 2024

Supervised Learning for Stochastic Optimal Control

Vince Kurtz, Joel W. Burdick

PDF

Open Access

TL;DR

This paper introduces a method to generate supervised learning data for continuous-time nonlinear stochastic optimal control problems using the Feynman-Kac theorem, enabling the application of supervised learning techniques to control tasks.

Contribution

It presents a novel approach to automatically generate training data for stochastic optimal control by leveraging the Feynman-Kac theorem and stochastic process simulation.

Findings

01

Enables supervised learning for stochastic control problems.

02

Uses Feynman-Kac theorem to sample value functions.

03

Facilitates rapid data generation with hardware accelerators.

Abstract

Supervised machine learning is powerful. In recent years, it has enabled massive breakthroughs in computer vision and natural language processing. But leveraging these advances for optimal control has proved difficult. Data is a key limiting factor. Without access to the optimal policy, value function, or demonstrations, how can we fit a policy? In this paper, we show how to automatically generate supervised learning data for a class of continuous-time nonlinear stochastic optimal control problems. In particular, applying the Feynman-Kac theorem to a linear reparameterization of the Hamilton-Jacobi-Bellman PDE allows us to sample the value function by simulating a stochastic process. Hardware accelerators like GPUs could rapidly generate a large amount of this training data. With this data in hand, stochastic optimal control becomes supervised learning.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsFault Detection and Control Systems · Advanced Control Systems Optimization · Neural Networks and Applications