Exploring the Impact of Ensemble Learning using Various AI Models for Software Defect Prediction
Abstract
This study proposes an RGC ensemble model (combining Random Forest, Gradient Boosting, and CNN) for software defect prediction, compared against single models such as RF, GB, SVM, CNN, LSTM, XGBoost, AdaBoost, and KNN across five public NASA PROMISE datasets.
01 / Problem
Single models for software defect prediction each have different weaknesses in capturing defect patterns in code, limiting their predictive accuracy compared to an approach that combines the strengths of multiple models.
02 / Method
The RGC ensemble model combines predictions from Random Forest, Gradient Boosting, and CNN, compared against eight single models (RF, GB, SVM, CNN, LSTM, XGBoost, AdaBoost, KNN) on five PROMISE datasets, measuring accuracy, precision, recall, and computation time.
03 / Experiment
Each model was trained and evaluated on all five datasets under the same accuracy, precision, recall, and computation-time metrics, for an apples-to-apples comparison across models.
04 / Results
RGC achieved 92.3% accuracy, 91.8% precision, and 90.7% recall, outperforming all compared single models — but with the highest computation time (median >12ms).
05 / Contribution
Shows that combining tree-based models (RF, GB) with a deep learning model (CNN) in one ensemble improves software defect prediction accuracy beyond any single model tested.
Limitation
The ensemble model's black-box nature limits the interpretability of its predictions, and evaluation relied only on static code metrics, not broader project context.
Keywords
Read the paper
The full paper is rendered here — no download or external viewer needed.