Investigation of Dataset Features for Just-in-Time Defect Prediction

Giuseppe Ng; Charibeth Cheng

arXiv:2109.13634·cs.SE·September 29, 2021·1 cites

Investigation of Dataset Features for Just-in-Time Defect Prediction

Giuseppe Ng, Charibeth Cheng

PDF

Open Access

TL;DR

This paper revisits the Kamei dataset for JIT defect prediction, highlighting preprocessing challenges, proposing new features for model training, and discussing dataset limitations affecting unsupervised learning.

Contribution

It identifies preprocessing issues, introduces new features for defect prediction, and analyzes the dataset's limitations for improved model development.

Findings

01

Preprocessing difficulties in the Kamei dataset

02

Proposed new features for defect prediction models

03

Limitations of the dataset for unsupervised learning

Abstract

Just-in-time (JIT) defect prediction refers to the technique of predicting whether a code change is defective. Many contributions have been made in this area through the excellent dataset by Kamei. In this paper, we revisit the dataset and highlight preprocessing difficulties with the dataset and the limitations of the dataset on unsupervised learning. Secondly, we propose certain features in the Kamei dataset that can be used for training models. Lastly, we discuss the limitations of the dataset's features.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSoftware Engineering Research · Software Reliability and Analysis Research · Imbalanced Data Classification Techniques