Evaluating a bot detection model on git commit messages

Mehdi Golzadeh; Alexandre Decan; Tom Mens

arXiv:2103.11779·cs.SE·March 23, 2021·5 cites

Evaluating a bot detection model on git commit messages

Mehdi Golzadeh, Alexandre Decan, Tom Mens

PDF

Open Access 1 Repo

TL;DR

This paper presents an improved bot detection model for git commit messages, achieving higher precision and implemented as an open-source tool to identify bots in git repositories.

Contribution

It generalizes previous models to commit messages, retrains on a large dataset, and provides an open-source detection tool.

Findings

01

Precision increased from 0.77 to 0.80 with new model

02

Model successfully detects bots in git commit messages

03

Open-source tool BoDeGiC implemented for practical use

Abstract

Detecting the presence of bots in distributed software development activity is very important in order to prevent bias in large-scale socio-technical empirical analyses. In previous work, we proposed a classification model to detect bots in GitHub repositories based on the pull request and issue comments of GitHub accounts. The current study generalises the approach to git contributors based on their commit messages. We train and evaluate the classification model on a large dataset of 6,922 git contributors. The original model based on pull request and issue comments obtained a precision of 0.77 on this dataset. Retraining the classification model on git commit messages increased the precision to 0.80. As a proof-of-concept, we implemented this model in BoDeGiC, an open source command-line tool to detect bots in git repositories.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

mehdigolzadeh/BoDeGiC
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Malware Detection Techniques · Spam and Phishing Detection · Hate Speech and Cyberbullying Detection