Optimizing forensic file classification: enhancing SFCS with βk hyperparameter tuning

D. Paul Joseph; Viswanathan Perumal

PMC · DOI:10.7717/peerj-cs.2608·March 5, 2025

Optimizing forensic file classification: enhancing SFCS with βk hyperparameter tuning

D. Paul Joseph, Viswanathan Perumal

PDF

Open Access

TL;DR

This paper introduces a new forensic file classification system that improves accuracy and efficiency by optimizing topic modeling parameters.

Contribution

The novel βk hyperparameter enhances seed word selection through semantic and contextual similarity evaluation.

Findings

01

The proposed SFCS system removed 278k irrelevant files and identified 5.6k suspicious files.

02

The model achieved 94.6% accuracy, 94.4% precision, and 96.8% recall.

03

The system operates within O(n log n) time complexity.

Abstract

In forensic topical modelling, the α parameter controls the distribution of topics in documents. However, low, high, or incorrect values of α lead to topic sparsity, model overfitting, and suboptimal topic distribution. To control the word distribution across topics, the β parameter is introduced. However, low, high, or inappropriate β values lead to sparse distribution, disjointed topics, and abundant highly probable words. The βj parameter, in conjunction with seed-guided words based on Term Frequency and Inverse Document Frequency, is introduced to address the issues. Nevertheless, the data often suffers from skewness or noise due to frequent co-occurrences of unrelated polysemic word pairs generated using Pointwise Mutual Information. By integrating α, β, and βj into file classification systems, classification models converge to local optima with O(n log n* |V|) time complexity. To…

Linked entities

Genes, proteins, chemicals, diseases, species, mutations and cell lines named across the full text — each resolved to its canonical identifier and authoritative record.

Diseases3

DF SFCS RDC

Figures50

Click any figure to enlarge with its caption.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHandwritten Text Recognition Techniques · Topic Modeling · Digital and Cyber Forensics