Document Classification using File Names

Zhijian Li; Stefan Larson; Kevin Leach

arXiv:2410.01166·cs.CL·March 7, 2025

Document Classification using File Names

Zhijian Li, Stefan Larson, Kevin Leach

PDF

Open Access

TL;DR

This paper introduces a lightweight, file name-based document classification method that achieves high accuracy and speed, significantly outperforming complex models in time-sensitive applications.

Contribution

The paper presents a novel approach combining TF-IDF tokenization with lightweight supervised models for fast, accurate document classification using only file names.

Findings

01

Achieves over 99% accuracy on two datasets.

02

Processes more than 90% of documents with high accuracy.

03

Is 442 times faster than complex deep learning models.

Abstract

Rapid document classification is critical in several time-sensitive applications like digital forensics and large-scale media classification. Traditional approaches that rely on heavy-duty deep learning models fall short due to high inference times over vast input datasets and computational resources associated with analyzing whole documents. In this paper, we present a method using lightweight supervised learning models, combined with a TF-IDF feature extraction-based tokenization method, to accurately and efficiently classify documents based solely on file names, that substantially reduces inference time. Our results indicate that file name classifiers can process more than 90% of in-scope documents with 99.63% and 96.57% accuracy when tested on two datasets, while being 442x faster than more complex models such as DiT. Our method offers a crucial solution to efficiently process vast…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDigital and Cyber Forensics