Exploring space efficiency in a tree-based linear model for extreme   multi-label classification

He-Zhe Lin; Cheng-Hung Liu; Chih-Jen Lin

arXiv:2410.09554·cs.LG·October 15, 2024

Exploring space efficiency in a tree-based linear model for extreme multi-label classification

He-Zhe Lin, Cheng-Hung Liu, Chih-Jen Lin

PDF

Open Access 1 Video

TL;DR

This paper analyzes the space efficiency of tree-based linear models in extreme multi-label classification, showing that storing only non-zero weights can drastically reduce storage without performance loss.

Contribution

It provides a theoretical and empirical analysis of space complexity in tree models, introducing a method to estimate model size beforehand and demonstrating significant storage savings.

Findings

01

Up to 95% reduction in storage space compared to one-vs-rest methods.

02

Sparse data leads to many zero weights, enabling space-efficient storage.

03

A simple procedure to estimate model size before training.

Abstract

Extreme multi-label classification (XMC) aims to identify relevant subsets from numerous labels. Among the various approaches for XMC, tree-based linear models are effective due to their superior efficiency and simplicity. However, the space complexity of tree-based methods is not well-studied. Many past works assume that storing the model is not affordable and apply techniques such as pruning to save space, which may lead to performance loss. In this work, we conduct both theoretical and empirical analyses on the space to store a tree model under the assumption of sparse data, a condition frequently met in text data. We found that, some features may be unused when training binary classifiers in a tree method, resulting in zero values in the weight vectors. Hence, storing only non-zero elements can greatly save space. Our experimental results indicate that tree models can achieve up to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Exploring Space Efﬁciency in a Tree-based Linear Model for Extreme Multi-label Classiﬁcation· underline

Taxonomy

TopicsText and Document Classification Technologies

MethodsPruning