The Future of Search and Discovery in Big Data Analytics: Ultrametric   Information Spaces

Fionn Murtagh; Pedro Contreras

arXiv:1202.3451·cs.IR·February 17, 2012

The Future of Search and Discovery in Big Data Analytics: Ultrametric Information Spaces

Fionn Murtagh, Pedro Contreras

PDF

Open Access

TL;DR

This paper explores ultrametric topologies for structuring data in big data analytics, enabling constant-time nearest neighbor searches and efficient hierarchy induction for large, high-dimensional datasets.

Contribution

It introduces methods for embedding data into ultrametric spaces and inducing hierarchies in linear time, enhancing proximity search efficiency in big data contexts.

Findings

01

Nearest neighbor search achieved in O(1) time after embedding.

02

Hierarchy induction in linear O(n) time.

03

Applicable to large, high-dimensional datasets.

Abstract

Consider observation data, comprised of n observation vectors with values on a set of attributes. This gives us n points in attribute space. Having data structured as a tree, implied by having our observations embedded in an ultrametric topology, offers great advantage for proximity searching. If we have preprocessed data through such an embedding, then an observation's nearest neighbor is found in constant computational time, i.e. O(1) time. A further powerful approach is discussed in this work: the inducing of a hierarchy, and hence a tree, in linear computational time, i.e. O(n) time for n observations. It is with such a basis for proximity search and best match that we can address the burgeoning problems of processing very large, and possibly also very high dimensional, data sets.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

Topicsadvanced mathematical theories · Chaos-based Image/Signal Encryption · Computational Physics and Python Applications