UCB-type Algorithm for Budget-Constrained Expert Learning

Ilgam Latypov; Alexandra Suvorikova; Alexey Kroshnin; Alexander Gasnikov; Yuriy Dorn

arXiv:2510.22654·cs.LG·January 19, 2026

UCB-type Algorithm for Budget-Constrained Expert Learning

Ilgam Latypov, Alexandra Suvorikova, Alexey Kroshnin, Alexander Gasnikov, Yuriy Dorn

PDF

TL;DR

This paper introduces M-LCB, a UCB-style algorithm for selecting and updating multiple adaptive experts under fixed training budgets, providing regret guarantees in stochastic settings.

Contribution

It presents the first regret guarantees for training multiple adaptive experts simultaneously with per-round budget constraints.

Findings

01

M-LCB achieves regret bounds of O(\u221a{KT/M} + (K/M)^{1-ss}T^ss)

02

Applicable to parametric models and multi-armed bandit experts

03

Extends classical bandit paradigm to resource-limited, self-learning experts.

Abstract

In many modern applications, a system must dynamically choose between several adaptive learning algorithms that are trained online. Examples include model selection in streaming environments, switching between trading strategies in finance, and orchestrating multiple contextual bandit or reinforcement learning agents. At each round, a learner must select one predictor among $K$ adaptive experts to make a prediction, while being able to update at most $M \leq K$ of them under a fixed training budget. We address this problem in the \emph{stochastic setting} and introduce \algname{M-LCB}, a computationally efficient UCB-style meta-algorithm that provides \emph{anytime regret guarantees}. Its confidence intervals are built directly from realized losses, require no additional optimization, and seamlessly reflect the convergence properties of the underlying experts. If each expert achieves…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.