A Divide and Conquer Algorithm of Bayesian Density Estimation

Ya Su

arXiv:2002.07094·stat.ME·February 18, 2020·1 cites

A Divide and Conquer Algorithm of Bayesian Density Estimation

Ya Su

PDF

Open Access

TL;DR

This paper introduces a divide and conquer Bayesian density estimation method that efficiently handles large datasets by splitting data, with theoretical guarantees and demonstrated advantages over existing methods in simulations and real data.

Contribution

It proposes a novel Bayesian divide and conquer approach for density estimation that maintains statistical efficiency and reduces computation time.

Findings

01

Posterior contraction rate remains nearly the same as the full data analysis.

02

Simulation studies show superior performance over the WASP estimator.

03

Application to GWAS data demonstrates practical advantages.

Abstract

Data sets for statistical analysis become extremely large even with some difficulty of being stored on one single machine. Even when the data can be stored in one machine, the computational cost would still be intimidating. We propose a divide and conquer solution to density estimation using Bayesian mixture modeling including the infinite mixture case. The methodology can be generalized to other application problems where a Bayesian mixture model is adopted. The proposed prior on each machine or subsample modifies the original prior on both mixing probabilities as well as on the rest of parameters in the distributions being mixed. The ultimate estimator is obtained by taking the average of the posterior samples corresponding to the proposed prior on each subset. Despite the tremendous reduction in time thanks to data splitting, the posterior contraction rate of the proposed estimator…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBayesian Methods and Mixture Models · Statistical Methods and Inference · Gaussian Processes and Bayesian Inference