Co-Clustering via Information-Theoretic Markov Aggregation

Clemens Bloechl; Rana Ali Amjad; Bernhard C. Geiger

arXiv:1801.00584·cs.LG·June 18, 2018

Co-Clustering via Information-Theoretic Markov Aggregation

Clemens Bloechl, Rana Ali Amjad, Bernhard C. Geiger

PDF

TL;DR

This paper introduces an information-theoretic cost function for co-clustering, grounded in Markov chain aggregation, unifying and extending previous approaches, with demonstrated effectiveness on real-world datasets.

Contribution

It proposes a novel cost function for co-clustering based on Markov aggregation, linking it to existing methods and providing insights into parameter effects.

Findings

01

Cost function aligns with known approaches for certain parameters.

02

Effective on synthetic and real-world datasets.

03

Highlights strengths and weaknesses of the method.

Abstract

We present an information-theoretic cost function for co-clustering, i.e., for simultaneous clustering of two sets based on similarities between their elements. By constructing a simple random walk on the corresponding bipartite graph, our cost function is derived from a recently proposed generalized framework for information-theoretic Markov chain aggregation. The goal of our cost function is to minimize relevant information loss, hence it connects to the information bottleneck formalism. Moreover, via the connection to Markov aggregation, our cost function is not ad hoc, but inherits its justification from the operational qualities associated with the corresponding Markov aggregation problem. We furthermore show that, for appropriate parameter settings, our cost function is identical to well-known approaches from the literature, such as Information-Theoretic Co-Clustering of Dhillon…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.