Exploring the Impact of Various Image Compression on Machine and Deep Learning Model Accuracy
Abstract
Large medical image datasets (ISIC 2016 skin lesions) slow down machine learning and deep learning model training. This study compares the impact of Huffman Encoding versus Discrete Cosine Transform (DCT) compression on preprocessing time, classification time, and accuracy across CNN, SVM, Random Forest, Logistic Regression, Gradient Boosting, and Ensemble models.
01 / Problem
High-resolution medical image datasets demand large storage and processing time, but it is unclear which compression method can significantly cut data size without degrading classification model accuracy.
02 / Method
The dataset is compressed separately using Huffman Encoding and DCT, then six classification models (CNN, SVM, RF, LR, GB, Ensemble) are trained on each compressed version, comparing preprocessing time, classification time, data size, and accuracy against the uncompressed data.
03 / Experiment
All six models were trained and evaluated on three dataset versions (original, Huffman-compressed, DCT-compressed), measuring the resulting data size, processing time, and each model's classification accuracy on each version.
04 / Results
Huffman Encoding cut data size by 83.16% (to ~131MB), more than DCT's 48.33% reduction (to ~402MB). Random Forest and Gradient Boosting kept stable accuracy around 81% under both compressions, while SVM's accuracy dropped (from 77% to 70% under Huffman compression).
05 / Contribution
Provides practical guidance for choosing a compression method in medical-imaging ML pipelines: Huffman Encoding excels at storage efficiency, but the choice of classification model (RF/GB vs. SVM) determines how well accuracy is preserved.
Limitation
Evaluation used only one dataset (ISIC 2016) and two compression methods, so generalizing to other medical image types or other compression methods still requires further research.
Keywords
Read the paper
The full paper is rendered here — no download or external viewer needed.