An Automatically Created Novel Bug Dataset and its Validation in Bug Prediction

Ferenc, Rudolf; Gyimesi, Péter; Gyimesi, Gábor; Tóth, Zoltán; Gyimóthy, Tibor

doi:10.1016/j.jss.2020.110691

Computer Science > Software Engineering

arXiv:2006.10158 (cs)

[Submitted on 17 Jun 2020]

Title:An Automatically Created Novel Bug Dataset and its Validation in Bug Prediction

Authors:Rudolf Ferenc, Péter Gyimesi, Gábor Gyimesi, Zoltán Tóth, Tibor Gyimóthy

View PDF

Abstract:Bugs are inescapable during software development due to frequent code changes, tight deadlines, etc.; therefore, it is important to have tools to find these errors. One way of performing bug identification is to analyze the characteristics of buggy source code elements from the past and predict the present ones based on the same characteristics, using e.g. machine learning models. To support model building tasks, code elements and their characteristics are collected in so-called bug datasets which serve as the input for learning.
We present the \emph{BugHunter Dataset}: a novel kind of automatically constructed and freely available bug dataset containing code elements (files, classes, methods) with a wide set of code metrics and bug information. Other available bug datasets follow the traditional approach of gathering the characteristics of all source code elements (buggy and non-buggy) at only one or more pre-selected release versions of the code. Our approach, on the other hand, captures the buggy and the fixed states of the same source code elements from the narrowest timeframe we can identify for a bug's presence, regardless of release versions. To show the usefulness of the new dataset, we built and evaluated bug prediction models and achieved F-measure values over 0.74.

Subjects:	Software Engineering (cs.SE)
Cite as:	arXiv:2006.10158 [cs.SE]
	(or arXiv:2006.10158v1 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2006.10158
Related DOI:	https://doi.org/10.1016/j.jss.2020.110691

Submission history

From: Zoltán Tóth [view email]
[v1] Wed, 17 Jun 2020 21:03:57 UTC (775 KB)

Computer Science > Software Engineering

Title:An Automatically Created Novel Bug Dataset and its Validation in Bug Prediction

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Software Engineering

Title:An Automatically Created Novel Bug Dataset and its Validation in Bug Prediction

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators