Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control

SM Mazharul Islam; Manfred Huber

arXiv:2508.13922·cs.LG·August 20, 2025

Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control

SM Mazharul Islam, Manfred Huber

PDF

TL;DR

This paper introduces Categorical Policies, a novel approach for modeling multimodal behavior in deep reinforcement learning, enhancing exploration and convergence in continuous control tasks.

Contribution

It proposes a new multimodal policy framework using categorical distributions, enabling differentiable sampling and improved exploration in continuous control environments.

Findings

01

Faster convergence compared to Gaussian policies

02

Improved exploration leading to better performance

03

Effective modeling of multimodal behaviors

Abstract

A policy in deep reinforcement learning (RL), either deterministic or stochastic, is commonly parameterized as a Gaussian distribution alone, limiting the learned behavior to be unimodal. However, the nature of many practical decision-making problems favors a multimodal policy that facilitates robust exploration of the environment and thus to address learning challenges arising from sparse rewards, complex dynamics, or the need for strategic adaptation to varying contexts. This issue is exacerbated in continuous control domains where exploration usually takes place in the vicinity of the predicted optimal action, either through an additive Gaussian noise or the sampling process of a stochastic policy. In this paper, we introduce Categorical Policies to model multimodal behavior modes with an intermediate categorical distribution, and then generate output action that is conditioned on…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.