An Instance-based Plus Ensemble Learning Method for Classification of   Scientific Papers

Fang Zhang; Shengli Wu

arXiv:2409.14237·cs.DL·September 24, 2024

An Instance-based Plus Ensemble Learning Method for Classification of Scientific Papers

Fang Zhang, Shengli Wu

PDF

Open Access

TL;DR

This paper presents a novel ensemble learning approach that combines instance-based methods and content/citation features for accurately classifying scientific papers into research fields, addressing the challenge of exponential publication growth.

Contribution

It introduces a new classification method that integrates seed papers, content and citation features, and ensemble techniques for scientific paper categorization.

Findings

01

Effective in categorizing papers into research areas

02

Utilizes both content and citation features

03

Demonstrates efficiency on DBLP datasets

Abstract

The exponential growth of scientific publications in recent years has posed a significant challenge in effective and efficient categorization. This paper introduces a novel approach that combines instance-based learning and ensemble learning techniques for classifying scientific papers into relevant research fields. Working with a classification system with a group of research fields, first a number of typical seed papers are allocated to each of the fields manually. Then for each paper that needs to be classified, we compare it with all the seed papers in every field. Contents and citations are considered separately. An ensemble-based method is then employed to make the final decision. Experimenting with the datasets from DBLP, our experimental results demonstrate that the proposed classification method is effective and efficient in categorizing papers into various research areas. We…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Text Analysis Techniques · Text and Document Classification Technologies