A REVIEW OF MULTIMODAL DEEPFAKE FORGERY DETECTION IN AUDIO-VIDEO-TEXT
DOI:
https://doi.org/10.63878/jalt2444Abstract
Recent advancements in generative models, particularly those based on deep learning and artificial intelligence, have significantly enhanced the realism and accessibility of deepfake content. While these technologies offer numerous creative and industrial benefits, they also pose a serious threat to the authenticity and trustworthiness of digital media. The rapid evolution of deepfake techniques has made it increasingly difficult to distinguish between genuine and manipulated content, thereby raising concerns across multiple domains, including media, politics, and cybersecurity. As a result, there is a growing need for robust and advanced detection mechanisms that can effectively identify such manipulations and preserve the integrity of information in the digital ecosystem.
In this review, we present a comprehensive synthesis of existing deep learning frameworks designed for multimodal deepfake detection. Unlike traditional unimodal approaches that focus solely on visual or auditory cues, multimodal detection systems integrate auditory, visual, and textual data to provide a more holistic and accurate analysis. Special emphasis is placed on fusion techniques that combine features from these modalities, enabling models to detect subtle inconsistencies that may not be evident when analyzing a single modality in isolation. Additionally, this study evaluates widely used public datasets and benchmark evaluation metrics, highlighting their role in training and validating detection systems. The findings demonstrate that multimodal approaches significantly improve detection performance, especially in complex scenarios involving highly sophisticated manipulations.
Furthermore, the importance of deep fake detection extends beyond technical challenges, as it has critical implications for society. Effective detection systems play a vital role in applications such as social media content moderation, judicial forensic analysis, and fraud prevention. At the same time, the deepfake of deceptive content can lead to serious psychological and social consequences, including misinformation, reputational damage, and mental health issues such as anxiety and depression. Therefore, future research should prioritize the development of lightweight, scalable, and efficient architectures that can be deployed in real-world environments. Moreover, there is a strong need for standardized evaluation protocols and interdisciplinary collaboration to ensure consistent benchmarking and improved robustness. Such efforts will be essential in building reliable systems capable of safeguarding digital content authenticity in an increasingly complex and dynamic technological landscape.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

