Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics

Aaron Wei; Milad Jalali; Danica J. Sutherland

arXiv:2512.13997·stat.ML·December 17, 2025

Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized U-Statistics

Aaron Wei, Milad Jalali, Danica J. Sutherland

PDF

Open Access

TL;DR

This paper extends the theory of MMD-based two-sample tests to handle unequal sample sizes by using generalized U-statistics, improving test power and data utilization in practical scenarios.

Contribution

It introduces a new theoretical framework for MMD with unequal samples, enabling more accurate and powerful two-sample testing without discarding data.

Findings

01

Derived asymptotic distributions for MMD with unequal samples

02

Provided a new criterion for optimizing MMD test power

03

Characterized variance of MMD estimators and their properties

Abstract

Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require discarding valuable data, unnecessarily reducing test power. We address this long-standing limitation by extending the theory of generalized U-statistics and applying it to the usual MMD estimator, resulting in new characterization of the asymptotic distributions of the MMD estimator with unequal sample sizes (particularly outside the proportional regimes required by previous partial results). This generalization also provides a new criterion for optimizing the power of an MMD test with unequal sample sizes. Our approach preserves all available data, enhancing test accuracy and applicability in realistic settings. Along the way, we give much cleaner…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsStatistical Methods and Inference · Mathematical Approximation and Integration · Statistical Methods in Clinical Trials