Exploring the Impact of Bayesian Optimization Towards Performance for Software Defect Prediction
Abstract
Hyperparameter tuning for software defect prediction (SDP) models via grid search or random search tends to be inefficient. This study applies Bayesian Optimization to SVM, Random Forest, Gradient Boosting, and Ensemble models, then measures its impact on accuracy and computation time across four public NASA PROMISE datasets.
01 / Problem
Grid search and random search for tuning SDP model hyperparameters require many trials without leveraging information from previous trials, making them inefficient and prone to missing the best hyperparameter combination.
02 / Method
Bayesian Optimization is applied to tune the hyperparameters of SVM, Random Forest, Gradient Boosting, and an Ensemble model, compared against their unoptimized counterparts, on four PROMISE datasets (CM1, KC1, KC2, PC1) with a 70/30 split and SMOTE class balancing.
03 / Experiment
Each model (with and without Bayesian Optimization) was trained and evaluated on all four datasets, measuring classification accuracy and computation time before and after optimization.
04 / Results
The optimized Ensemble achieved the highest accuracy (~94%), followed by optimized Random Forest (~92%); SVM saw the largest accuracy gain from optimization. However, optimization also increased computation time across all models (e.g. Gradient Boosting rose from 4ms to 6ms).
05 / Contribution
Shows that Bayesian Optimization consistently improves software defect prediction accuracy, at the cost of increased computation time that must be weighed against the application's needs.
Limitation
Evaluation used only NASA PROMISE datasets, so generalizing the results to codebases outside that context remains limited; the optimization-driven increase in computation time was also not evaluated at production scale.
Keywords
Read the paper
The full paper is rendered here — no download or external viewer needed.