Preserving local densities in low-dimensional embeddings

Jonas Fischer; Rebekka Burkholz; Jilles Vreeken

arXiv:2301.13732·cs.LG·February 1, 2023·1 cites

Preserving local densities in low-dimensional embeddings

Jonas Fischer, Rebekka Burkholz, Jilles Vreeken

PDF

Open Access

TL;DR

This paper identifies limitations of current low-dimensional embedding methods like tSNE and UMAP in preserving local densities, introduces dtSNE to address this, and demonstrates its improved accuracy in representing local structures.

Contribution

The paper introduces dtSNE, a new embedding method that better preserves local densities compared to existing techniques.

Findings

01

dtSNE maintains local density relationships more accurately.

02

Existing methods can distort local densities and cluster sizes.

03

dtSNE achieves comparable global structure reconstruction.

Abstract

Low-dimensional embeddings and visualizations are an indispensable tool for analysis of high-dimensional data. State-of-the-art methods, such as tSNE and UMAP, excel in unveiling local structures hidden in high-dimensional data and are therefore routinely applied in standard analysis pipelines in biology. We show, however, that these methods fail to reconstruct local properties, such as relative differences in densities (Fig. 1) and that apparent differences in cluster size can arise from computational artifact caused by differing sample sizes (Fig. 2). Providing a theoretical analysis of this issue, we then suggest dtSNE, which approximately conserves local densities. In an extensive study on synthetic benchmark and real world data comparing against five state-of-the-art methods, we empirically show that dtSNE provides similar global reconstruction, but yields much more accurate…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBioinformatics and Genomic Networks · Gene expression and cancer classification · Gene Regulatory Network Analysis

Methodsfail