Survey Design and Estimating Equations when Combining Big Data with Probability Samples
Ryan Covey (1), Lucca Buonamano (1) ((1) Methodology, Data Science, Division, Australian Bureau of Statistics)

TL;DR
This paper develops a unified method to combine big data with probability samples, producing estimators that are consistent, unbiased, and more efficient, addressing selection bias in official statistics.
Contribution
It introduces a novel estimating equation framework that integrates big data with probability sampling, ensuring valid inference under weak assumptions.
Findings
Integrated estimators are consistent and asymptotically unbiased.
Variance estimators are derived for both design-based and superpopulation-based inference.
Combining big data with probability samples improves efficiency when dependence is low and population is large.
Abstract
The use of big data in official statistics and the applied sciences is accelerating, but statistics computed using only big data often suffer from substantial selection bias. This leads to inaccurate estimation and invalid statistical inference. We rectify the issue for a broad class of linear and nonlinear statistics by producing estimating equations that combine big data with a probability sample. Under weak assumptions about an unknown superpopulation, we show that our integrated estimator is consistent and asymptotically unbiased with an asymptotic normal distribution. Variance estimators with respect to both the sampling design alone and jointly with the superpopulation are obtained at once using a single, unified theoretical approach. A surprising corollary is that strategies minimising the design variance almost minimise the joint variance when the population and sample sizes are…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsStatistical Methods and Bayesian Inference · Statistical Methods and Inference
