Deep fake detection in face images and video by tuning a pretrained CLIP model
Maomao Ling, Włodzimierz Kasprzak
DOI: http://dx.doi.org/10.15439/2026F4803
Citation: Maomao Ling, Włodzimierz Kasprzak (2026). Deep fake detection in face images and video by tuning a pretrained CLIP model. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS). ACSIS, Vol. 47, pages 573–578.
Abstract. This study develops a video deepfake detection system that addresses the critical challenge of cross-dataset generalization in real-world scenarios. It adopts CLIP (Contrastive Language-Image Pre-training) as the foundation, leveraging its strong vision-language representations for detecting subtle facial manipulations. A parameter-efficient fine-tuning approach is proposed that achieves superior cross-dataset performance while updating only a minimal fraction of model parameters. The CLIP ViT-B model, utilizing the LayerNorm tuning, achieves an average video-level AUC of 90.9\\% across seven benchmark datasets, while training only 0.0344\\% of the total parameters. The study integrates CLIP with LayerNorm tuning into the Deepfake Bench framework, enabling comprehensive cross-dataset generalization analysis. Notably, the implemented solution matches more complex CLIP variants in generalization performance underscoring the significance of high-quality pre-training over architectural complexity. The findings highlight the effectiveness of vision-language models for robust facial manipulation detection and the importance of parameter-efficient adaptation strategies.
References
- Shichuang Xie, Tong Qiao, Sheng Li, Xinpeng Zhanga, Jiantao Zhoud, Guorui Fenga, “DeepFake detection in the AIGC era: A survey, benchmarks, and future perspectives”. Information Fusion, 127 (2026) 103740, https://doi.org/10.1016/j.inffus.2025.103740
- “Stable Diffusion”. https://stability.ai/
- “DALL-E”. https://openai.com/index/dall-e/
- O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, D. Lischinski, “Styleclip: textdriven manipulation of stylegan imagery”, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 2065–2074 https://doi.org/10.1109/ICCV48922.2021.00209
- B. Dolhansky, J. Bitton, B. Pflaum, J. Lu, R. Howes, M. Wang, C. Canton Ferrer, “The deepfake detection challenge (dfdc) dataset”, https://arxiv.org/abs/ 2006.07397, 2020. https://doi.org/10.48550/arXiv.2006.07397
- Y.Li, X.Yang, P.Sun, H.Qi, and S.Lyu, “Celeb-DF: A large-scale challenging dataset for deepfake forensics”, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 3207—3216 https://doi.org/10.48550/arXiv.1909.12962
- J. Thies, M. Zollhöfer, M. Stamminger, C. Theobalt, and M. Nießner, “Face2face: Real-time face capture and reenactment of rgb videos”, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2387—2395 https://doi.org/10.1109/CVPR.2016. 262
- W. Zhong, C. Fang, Y. Cai, P. Wei, G. Zhao, L. Lin, G. Li, “Identitypreserving talking face generation with landmark and appearance priors”, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9729–9738. https://doi.org/10.48550/ arXiv.2305.08293
- C. Zhang, C. Wang, Y. Zhao, S. Cheng, L. Luo, X. Guo, “Dr2: disentangled recurrent representation learning for data-efficient speech video synthesis”, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 6192-6202. https://doi. org/10.1109/WACV57701.2024.00609
- I. Goodfellow, J. Pouget-Abadie, M. Mirza, et al., “Generative adversarial nets”, in Advances in Neural Information Processing Systems (NeurIPS), vol.27, 2014, pp.2672– 2680. https://proceedings.neurips.cc/paper_files/paper/2014/file/ f033ed80deb0234979a61f95710dbe25-Paper.pdf
- J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models”, in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 6840–6851. https://doi.org/10.48550/arXiv.2006. 11239
- D.P. Kingma and M.Welling, “Auto-encoding variational bayes”, in International Conference on Learning Representations (ICLR), 2014. https://doi.org/10.48550/arXiv.1312.6114
- R. Liu, B. Ma, W. Zhang, Z. Hu, C. Fan, T. Lv, Y. Ding, X. Cheng, “Towards a simultaneous and granular identity-expression control in personalized face generation”, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 2114–2123. https://doi.org/10.1109/CVPR52733.2024.00206
- S. Xu, G. Chen, Y.-X. Guo, J. Yang, C. Li, Z. Zang, Y. Zhang, X. Tong, B. Guo, Vasa-1: lifelike audio-driven talking faces generated in real time, arXiv: 2404.10667 (2024) https://doi.org/10.48550/arXiv. 2404.10667
- J. Sun, Q. Deng, Q. Li, M. Sun, Y. Liu, Z. Sun, “Anyface++: a unified framework for free-style text-to-face synthesis and manipulation”, IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 11, pp. 9438–9453, Nov. 2025, https://doi.org/10.1109/TPAMI.2023.3345866
- A. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision”, in: International Conference on Machine Learning, PMLR, vol. 139, 2021, pp. 8748–8763. https://proceedings.mlr.press/v139/radford21a.html
- A. Yermakov, J. Cech, and J. Matas, “Unlocking the hidden potential of clip in generalizable deepfake detection”, arXiv:2503.19683, 2025 https://doi.org/10.48550/arXiv.2503.19683
- Z. Yan, Y. Zhang, X. Yuan, S. Lyu, and B. Wu, “Deepfakebench: A comprehensive benchmark of deepfake detection”, in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 4534–4565. https://doi.org/10.48550/arXiv.2307.01426
- A.Paszke, S. Gross, F. Massa, et al., “Pytorch: An imperative style, highperformance deep learning library”, in Advances in Neural Information Processing Systems, 2019, pp. 8024–8035. https://doi.org/10.48550/ arXiv.1912.01703
- NVIDIA Corporation, “Cuda toolkit documentation”, https://docs.nvidia. com/cuda/
- TorchVision maintainers and contributors, “Torchvision: Pytorch’s computer vision library”, https://github.com/pytorch/vision
- – “OpenCV”, https://opencv.org/releases/
- – “Transformer 4.5.0”, https://huggingface.co/transformers/v4.5.0/index. html
- D. E. King, “Dlib-ml: A machine learning toolkit”, Journal of Machine Learning Research, vol. 10, pp. 1755–1758, 2009. https://www.dlib.net/ python/
- TensorFlow Team, “Tensorboard - Tensorflow’s visualization toolkit”. https://www.tensorflow.org/tensorboard
- – “Scikit-learn”. https://scikit-learn.org/stable/index.html
- A. Rössler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “Faceforensics++: Learning to detect manipulated facial images”, in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1–11. https://doi.org/10.1109/ICCV.2019. 00009
- B. Dolhansky, R. Howes, B. Pflaum, N. Baram, and C. C. Ferrer, “The deepfake detection challenge (dfdc) preview dataset”, arXiv:1910.08854, 2019. https://arxiv.org/pdf/1910.08854
- Y. Li and S. Lyu, “Exposing deepfake videos by detecting face warping artifacts”, in IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, 2018, pp. 46–52. https://doi.org/10. 48550/arXiv.1811.00656
- L. Li, J. Bao, H. Zhang, et al., “Faceshifter: Towards high fidelity and occlusion aware face swapping”, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7390–7398. https://doi.org/10.48550/arXiv.1912.13457
- Nick Dufour and Andrew Gully, “Contributing data to deepfake detection research”, [Online] 2019. https://research.google/blog/ contributing-data-to-deepfake-detection-research/
- X.Fu, Z.Yan, T.Yao, S. Chen, and X.Li, “Exploring unbiased deepfake detection via token-level shuffling and mixing”, arXiv:2501.04376, 2025. https://doi.org/10.48550/arXiv.2501.04376