Dynamic Decision-Making under Model Misspecification: A Stochastic Stability Approach

Xinyu Dai; Daniel Chen; Yian Qian

arXiv:2602.17086·econ.TH·February 20, 2026

Dynamic Decision-Making under Model Misspecification: A Stochastic Stability Approach

Xinyu Dai, Daniel Chen, Yian Qian

PDF

Open Access

TL;DR

This paper analyzes how Bayesian reinforcement learning, specifically Thompson Sampling, behaves under model misspecification, providing a geometric classification of its long-term beliefs and actions in structured bandit problems.

Contribution

It offers the first qualitative and geometric classification of Thompson Sampling under model misspecification, introducing a stochastic stability framework for posterior dynamics.

Findings

01

Identifies regimes of correct and incorrect model concentration and persistent belief mixing.

02

Provides sharp predictions for limiting beliefs, action frequencies, and asymptotic regret.

03

Develops a unified stochastic stability framework for posterior evolution.

Abstract

Dynamic decision-making under model uncertainty is central to many economic environments, yet existing bandit and reinforcement learning algorithms rely on the assumption of correct model specification. This paper studies the behavior and performance of one of the most commonly used Bayesian reinforcement learning algorithms, Thompson Sampling (TS), when the model class is misspecified. We first provide a complete dynamic classification of posterior evolution in a misspecified two-armed Gaussian bandit, identifying distinct regimes: correct model concentration, incorrect model concentration, and persistent belief mixing, characterized by the direction of statistical evidence and the model-action mapping. These regimes yield sharp predictions for limiting beliefs, action frequencies, and asymptotic regret. We then extend the analysis to a general finite model class and develop a unified…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Bandit Algorithms Research · Reinforcement Learning in Robotics · Game Theory and Applications