A fast algorithm for computing distance correlation

Arin Chaudhuri; Wenhao Hu

arXiv:1810.11332·stat.CO·February 7, 2019

A fast algorithm for computing distance correlation

Arin Chaudhuri, Wenhao Hu

PDF

TL;DR

This paper introduces a new $ ext{O}(n ext{log } n)$ algorithm for computing distance correlation between univariate variables, significantly improving efficiency over previous quadratic-time methods, thus enabling analysis of large datasets.

Contribution

The paper presents a simple, exact, and fast sorting-based algorithm for calculating distance correlation, reducing computational complexity from $ ext{O}(n^2)$ to $ ext{O}(n ext{log } n)$.

Findings

01

The algorithm is significantly faster than existing methods.

02

Empirical results confirm the efficiency and practicality of the proposed approach.

03

The method facilitates analysis of large datasets for dependence structures.

Abstract

Classical dependence measures such as Pearson correlation, Spearman's $ρ$ , and Kendall's $τ$ can detect only monotonic or linear dependence. To overcome these limitations, Szekely et al.(2007) proposed distance covariance as a weighted $L_{2}$ distance between the joint characteristic function and the product of marginal distributions. The distance covariance is $0$ if and only if two random vectors $X$ and $Y$ are independent. This measure has the power to detect the presence of a dependence structure when the sample size is large enough. They further showed that the sample distance covariance can be calculated simply from modified Euclidean distances, which typically requires $O (n^{2})$ cost. The quadratic computing time greatly limits the application of distance covariance to large data. In this paper, we present a simple exact $O (n lo g (n))$ algorithm to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.