Dealing with overdispersion in multivariate count data

Noemi Corsini; Cinzia Viroli

arXiv:2107.00470·stat.ME·February 24, 2025·Comput. Stat. Data Anal.

Dealing with overdispersion in multivariate count data

Noemi Corsini, Cinzia Viroli

PDF

TL;DR

This paper reviews likelihood-based models for overdispersion in multivariate count data, introduces a new advanced model improving variability approximation, and demonstrates its superior performance through simulation studies.

Contribution

It proposes a deeper Dirichlet-Multinomial model with a new estimation method, enhancing overdispersion modeling in high-dimensional count data.

Findings

01

The new model better captures observed variability.

02

Simulation studies confirm superior performance.

03

Model applicable to high-dimensional data.

Abstract

The problem of overdispersion in multivariate count data is a challenging issue. Nowadays, it covers a central role mainly due to the relevance of modern technologies data, such as Next Generation Sequencing and textual data from the web or digital collections. This work presents a comprehensive analysis of the likelihood-based models for extra-variation data proposed in the scientific literature. Particular attention will be paid to the models feasible for high-dimensional data. A new approach together with its parametric-estimation procedure is proposed. It is a deeper version of the Dirichlet-Multinomial distribution and it leads to important results allowing to get a better approximation of the observed variability. A significative comparison of these models is made through two different simulation studies that both confirm that the new model considered in this work allows to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.