RepoMiner: a Language-agnostic Python Framework to Mine Software   Repositories for Defect Prediction

Stefano Dalla Palma; Dario Di Nucci; Damian Tamburri

arXiv:2111.11807·cs.SE·November 24, 2021

RepoMiner: a Language-agnostic Python Framework to Mine Software Repositories for Defect Prediction

Stefano Dalla Palma, Dario Di Nucci, Damian Tamburri

PDF

TL;DR

RepoMiner is a versatile Python framework that automates data collection, labeling, and metric calculation from software repositories, facilitating defect prediction research across multiple programming languages.

Contribution

It introduces a language-agnostic tool that simplifies and automates dataset creation for defect prediction, reducing manual effort and errors.

Findings

01

Successfully collects failure data from repositories

02

Automatically labels failure-prone components

03

Calculates relevant metrics for defect prediction

Abstract

Data originating from open-source software projects provide valuable information to enhance software quality. In the scope of Software Defect Prediction, one of the most challenging parts is extracting valid data about failure-prone software components from these repositories, which can help develop more robust software. In particular, collecting data, calculating metrics, and synthesizing results from these repositories is a tedious and error-prone task, which often requires understanding the programming languages involved in the mined repositories, eventually leading to a proliferation of language-specific data-mining software. This paper presents RepoMiner, a language-agnostic tool developed to support software engineering researchers in creating datasets to support any study on defect prediction. RepoMiner automatically collects failure data from software components, labels them as…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.