Using a Power Law Distribution to describe Big Data

Vijay Gadepally; Jeremy Kepner

arXiv:1509.00504·cs.SI·January 25, 2017

Using a Power Law Distribution to describe Big Data

Vijay Gadepally, Jeremy Kepner

PDF

TL;DR

The paper proposes a method to model big data with a power law distribution, enabling efficient filtering of uninteresting data and improving data analysis scalability.

Contribution

It introduces a technique to derive a power law background model from arbitrary datasets, aiding in data filtering and scalability for big data applications.

Findings

01

Model accurately fits social sensor data distributions

02

Enables automatic identification of high degree nodes

03

Scales effectively with large datasets

Abstract

The gap between data production and user ability to access, compute and produce meaningful results calls for tools that address the challenges associated with big data volume, velocity and variety. One of the key hurdles is the inability to methodically remove expected or uninteresting elements from large data sets. This difficulty often wastes valuable researcher and computational time by expending resources on uninteresting parts of data. Social sensors, or sensors which produce data based on human activity, such as Wikipedia, Twitter, and Facebook have an underlying structure which can be thought of as having a Power Law distribution. Such a distribution implies that few nodes generate large amounts of data. In this article, we propose a technique to take an arbitrary dataset and compute a power law distributed background model that bases its parameters on observed statistics. This…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.