Evolving Text Data Stream Mining

Jay Kumar

arXiv:2409.00010·cs.IR·September 4, 2024

Evolving Text Data Stream Mining

Jay Kumar

PDF

Open Access

TL;DR

This paper reviews challenges and proposes new models for mining evolving text data streams, addressing issues like high dimensionality, semantic representation, and label scarcity in real-time online social platform data.

Contribution

It introduces novel learning models for clustering and multi-label learning tailored to the unique properties of evolving text streams.

Findings

01

Improved clustering performance on high-dimensional text streams

02

Effective semantic representation capturing evolving topics

03

Models handle label scarcity in streaming data

Abstract

A text stream is an ordered sequence of text documents generated over time. A massive amount of such text data is generated by online social platforms every day. Designing an algorithm for such text streams to extract useful information is a challenging task due to unique properties of the stream such as infinite length, data sparsity, and evolution. Thereby, learning useful information from such streaming data under the constraint of limited time and memory has gained increasing attention. During the past decade, although many text stream mining algorithms have proposed, there still exists some potential issues. First, high-dimensional text data heavily degrades the learning performance until the model either works on subspace or reduces the global feature space. The second issue is to extract semantic text representation of documents and capture evolving topics over time. Moreover,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Mining Algorithms and Applications