Enhancing Binary Image Classification Accuracy Using Low-Rank Adaptation (LoRA) for Deepfake Detection

Authors

  • Tam Thanh Thi Pham Ho Chi Minh City University of Education, Viet Nam
  • Thao Thanh Thi Nguyen Ho Chi Minh City University of Education, Viet Nam
  • Thai Hoang Le University of Science, Viet Nam
  • Hai Son Tran Ho Chi Minh City University of Education, Viet Nam

DOI:

https://doi.org/10.64539/sjcs.v2i2.2026.475

Keywords:

Low-Rank Adaptation, LoRA, Deepfake Detection, Vision Transformer, ResNet-50, Binary Image Classification, Parameter-Efficient Fine-Tuning, Transfer Learning

Abstract

Deepfake technology poses an increasingly serious threat to personal reputation and social trust, necessitating the development of accurate yet computationally efficient detection systems. Although large pre-trained vision models offer exceptional feature extraction capabilities, full fine-tuning them demands prohibitive computational resources and risks overfitting. This study investigates the application of Low-Rank Adaptation (LoRA) to enhance binary image classification accuracy for deepfake face detection, bridging the gap between parameter efficiency and high classification performance. We systematically integrate LoRA into two dominant architectural paradigms: the Vision Transformer (Swin-T) and ResNet-50. Computational evaluations are conducted on a 40K sub-dataset from the 140K Real and Fake Faces dataset, comparing LoRA against full fine-tuning baselines under identical environments. Experimental results demonstrate that Swin-T + LoRA achieves an outstanding test accuracy of 99.14% and an F1-score of 0.9913, outperforming its full fine-tuning baseline by 9.95 percentage points while training only 5.88% of the total parameters. Conversely, ResNet-50 + LoRA improves test accuracy by 14.11 percentage points over its full fine-tuning baseline, although its performance remains substantially below that of Swin-T + LoRA, indicating that LoRA effectiveness varies across architectural paradigms. These findings demonstrate that parameter-efficient fine-tuning, particularly when combined with Transformer attention layers, offers a promising approach for developing accurate and computationally efficient deepfake detection systems under resource constraints.

References

[1] Z. Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” arXiv preprint arXiv:2103.14030, 2021. https://doi.org/10.48550/arXiv.2103.14030.

[2] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770-778, 2016. https://doi.org/10.1109/CVPR.2016.90.

[3] E. J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv preprint arXiv:2106.09685, 2021. https://doi.org/10.48550/arXiv.2106.09685.

[4] A. Aghajanyan, S. Gupta, and L. Zettlemoyer, “Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 2020. https://aclanthology.org/2021.acl-long.568.pdf.

[5] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “ QLoRA: Efficient Finetuning of Quantized LLMs,” arXiv preprint arXiv:2305.14314, 2023. https://doi.org/10.48550/arXiv.2305.14314.

[6] J. Zhang, S. Chen, X. Liu, and B. Shi, “AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning,” arXiv preprint arXiv:2303.10512, 2023. https://doi.org/10.48550/arXiv.2303.10512.

[7] M. Trigka and E. Dritsas, “A Comprehensive Survey of Deep Learning Approaches in Image Processing,” Sensors, vol. 25, no. 2, p. 531, 2025. https://doi.org/10.3390/s25020531.

[8] K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” arXiv preprint arXiv:1409.1556, 2014. https://doi.org/10.48550/arXiv.1409.1556.

[9] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” Communications of the ACM, vol. 60, no. 6, pp. 84 - 90, 2012. https://doi.org/10.1145/3065386.

[10] J. Yu, C. Luo, and F. Chen, “Human Face Analysis,” in Multi-Modal Human Modeling, Analysis and Synthesis, CRC Press, 2025, pp. 43-100. https://doi.org/10.1201/9781003408291-3.

[11] A. Vaswani et al., “Attention Is All You Need,” arXiv preprint arXiv:1706.03762, 2017. https://doi.org/10.48550/arXiv.1706.03762.

[12] D. Cozzolino, J. Thies, A. Rössler, C. Riess, M. Nießner, and L. Verdoliva, “FaceForensics++: Learning to Detect Manipulated Facial Images,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1-11, 2019. https://doi.org/10.1109/ICCV.2019.00009.

[13] L. Li, J. Bao, T. Zhang, H. Yang, D. Chen, F. Wen, and B. Guo, “Face X-ray for More General Face Forgery Detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5001-5010, 2020. https://doi.org/10.1109/CVPR42600.2020.00505.

[14] Y. Li, X. Yang, P. Sun, H. Qi, and S. Lyu, “Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3207-3216, 2020. https://doi.org/10.1109/CVPR42600.2020.00327.

[15] T. Karras, S. Laine, and T. Aila, “A Style-Based Generator Architecture for Generative Adversarial Networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4401-4410, 2019. https://doi.org/10.1109/CVPR.2019.00453.

[16] A. Karathanasis, J. Violos, and I. Kompatsiaris, “ A Comparative Analysis of Compression and Transfer Learning Techniques in DeepFake Detection Models,” Mathematics, vol. 13, no. 5, art. No. 887, 2025. https://doi.org/10.3390/math13050887.

[17] R. Tolosana, R. Vera-Rodriguez, J. Fierrez, A. Morales, and J. Ortega-Garcia, “Deepfakes and Beyond: A Survey of Face Manipulation and Fake Detection,” Information Fusion, vol. 64, pp. 131-148, 2020. https://doi.org/10.1016/j.inffus.2020.06.014.

[18] Xhlulu, “140k Real and Fake Faces,” Kaggle, 2020. [Online]. Available: https://www.kaggle.com/datasets/xhlulu/140k-real-and-fake-faces.

[19] R. A. Bafghi, N. Harilal, C. Monteleoni, and M. Raissi, “Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024. https://doi.org/10.1109/CVPRW63382.2024.00371.

[20] W. Shi, J. Xu, and P. Gao, “SSformer: A Lightweight Transformer for Semantic Segmentation,” arXiv preprint arXiv:2208.02034, 2022. https://doi.org/10.48550/arXiv.2208.02034.

[21] J. Yang, J. Liu, N. Xu, and J. Huang, “TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation,” in 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023. https://doi.org/10.1109/WACV56688.2023.00059.

[22] L. Tatzel, B. Mucsányi, O. Hackel, and P. Hennig, “Debiasing Mini-Batch Quadratics for Applications in Deep Learning,” arXiv preprint arXiv:2410.14325, 2024. https://doi.org/10.48550/arXiv.2410.14325.

[23] P. M. Abhilash and A. Ahmed, “Convolutional neural network–based classification for improving the surface quality of metal additive manufactured components,” The International Journal of Advanced Manufacturing Technology, vol. 126, pp. 3873–3885, 2023. https://doi.org/10.1007/s00170-023-11388-z.

[24] M. Kanagarajan et al., “AIM-Net: A Resource-Efficient Self-Supervised Learning Model for Automated Red Spider Mite Severity Classification in Tea Cultivation,” AgriEngineering, vol. 7, no. 8, art. No. 247, 2025. https://doi.org/10.3390/agriengineering7080247.

[25] M. Akter et al., “When uncertainty guides learning: a highly effective approach to kidney disease classification in CT imaging,” Frontiers in Big Data, vol. 9, art. No. 1825213. https://doi.org/10.3389/fdata.2026.1825213.

[26] O. O. Abayomi, A. E. Kayode, O. Oluwasanya, A. E. Mesioye, “Comparative Analysis of Machine Learning Algorithms for Predictive Maintenance,” Methods in Science and Technology Studies, vol. 2, no. 2, pp. 131-144, 2026. https://doi.org/10.64539/msts.v2i2.2026.505.

[27] H. H. Rashidi, S. Albahra, S. Robertson, N. K. Tran, and B. Hu, “Common statistical concepts in the supervised Machine Learning arena” Frontiers in Oncology, vol. 13, 2023. https://doi.org/10.3389/fonc.2023.1130229.

[28] T. Sarfraz, T. Ling, and A. Ijaz, “BReMS-Net: Prediction-Guided Coarse-to-Fine Refinement with Boundary-Aware Multi-Scale Dilated Fusion for Robust Breast Mass Segmentation,” Scientific Journal of Engineering Research, vol. 2, no. 3, pp. 327–338, 2026. https://doi.org/10.64539/sjer.v2i3.2026.489.

[29] G. P. Oise et al., “Interpretable Academic Outcome Prediction Using Explainable Boosting Machines,” Methods in Science and Technology Studies, vol. 2, no. 1, pp. 45–56, ,2026. https://doi.org/10.64539/msts.v2i1.2026.441.

[30] M. Mostafa, A. S. Almogren, M. Al-Qurishi, and M. Alrubaian, “Modality Deep-learning Frameworks for Fake News Detection on Social Networks: A Systematic Literature Review,” ACM Computing Surveys, vol. 57, no. 3, pp. 1 – 50, 2024, https://doi.org/10.1145/3700748.

Downloads

Published

2026-08-09

How to Cite

Pham, T. T. T., Nguyen, T. T. T., Le, T. H., & Tran, H. S. (2026). Enhancing Binary Image Classification Accuracy Using Low-Rank Adaptation (LoRA) for Deepfake Detection. Scientific Journal of Computer Science, 2(2), 346–354. https://doi.org/10.64539/sjcs.v2i2.2026.475

Similar Articles

<< < 1 2 3 

You may also start an advanced similarity search for this article.