An Automatically Created Novel Bug Dataset and its Validation in Bug   Prediction

Rudolf Ferenc; P\'eter Gyimesi; G\'abor Gyimesi; Zolt\'an T\'oth,; Tibor Gyim\'othy

arXiv:2006.10158·cs.SE·June 19, 2020

An Automatically Created Novel Bug Dataset and its Validation in Bug Prediction

Rudolf Ferenc, P\'eter Gyimesi, G\'abor Gyimesi, Zolt\'an T\'oth,, Tibor Gyim\'othy

PDF

TL;DR

This paper introduces BugHunter, an automatically generated bug dataset capturing buggy and fixed code states over narrow timeframes, and demonstrates its effectiveness in building high-accuracy bug prediction models.

Contribution

The paper presents a novel, automatically created bug dataset that captures bug states over narrow timeframes, improving bug prediction model training.

Findings

01

Achieved F-measure over 0.74 in bug prediction models

02

Introduced a new dataset capturing bug and fix states over narrow timeframes

03

Validated the dataset's usefulness in bug prediction tasks

Abstract

Bugs are inescapable during software development due to frequent code changes, tight deadlines, etc.; therefore, it is important to have tools to find these errors. One way of performing bug identification is to analyze the characteristics of buggy source code elements from the past and predict the present ones based on the same characteristics, using e.g. machine learning models. To support model building tasks, code elements and their characteristics are collected in so-called bug datasets which serve as the input for learning. We present the \emph{BugHunter Dataset}: a novel kind of automatically constructed and freely available bug dataset containing code elements (files, classes, methods) with a wide set of code metrics and bug information. Other available bug datasets follow the traditional approach of gathering the characteristics of all source code elements (buggy and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.