An integrated explainable framework for multimodal deepfake detection across image, audio, and video data
Ain Shams Engineering Journal, cilt.17, sa.7, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 17 Sayı: 7
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.asej.2026.104233
- Dergi Adı: Ain Shams Engineering Journal
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals, Engineering Source (EBSCO)
- Anahtar Kelimeler: BiLSTM, Cross attention, DenseNet169, Grad-CAM, Hybrid model, InceptionV3, LIME, LRP, Multimodal deepfake detection, SHAP, XAI
- İstanbul Medipol Üniversitesi Adresli: Evet
Özet
The rapid advancement of deepfake generation techniques has introduced significant challenges for digital forensics, particularly due to the increasing realism and multimodal nature of manipulated content. Existing detection approaches are predominantly unimodal and lack interpretability, limiting their effectiveness in real-world scenarios. To address these limitations, this study proposes a unified and explainable deepfake detection framework that operates across image, video, audio, and multimodal data. The proposed framework integrates modality-specific deep learning architectures, InceptionV3 for images, DenseNet169-Bidirectional Long Short-Term Memory (BiLSTM) for video, and Convolutional Neural Network (CNN)-BiLSTM with Extreme Gradient Boosting (XGBoost) for audio, along with a cross-attention-based fusion model for multimodal analysis. A key contribution of this work is a modality-adaptive explainable artificial intelligence (XAI) strategy that systematically combines Shapley Additive Explanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), Layer-wise Relevance Propagation (LRP), and Gradient-weighted Class Activation Mapping (Grad-CAM) to generate consistent visual and textual explanations across different data types. Extensive experiments conducted on benchmark datasets, including 140 K-Faces, Celeb-DF-V2, In-The-Wild, and LAV-DF, demonstrate the effectiveness of the proposed approach, achieving accuracy rates of 99% for images, 96% for videos, 98% for audio, and 99% for multimodal data. Furthermore, the framework provides interpretable insights into model decisions through an interactive interface, enhancing transparency and usability in practical applications. The results indicate that the proposed system not only improves detection performance but also addresses the critical need for unified and interpretable deepfake analysis in multimodal environments.