Deepfake Audio Synthesis and Detection: A Review of Recent Advances and Open Challenges
ISAISS Logo
Publication Type
Conference Article
Publication Year
2026
Author(s)
Afifa Yasin, Uswah Fatima, Muhammad Zaid, Asjad Amin
Conference Name
International Symposium on AI and Secure Systems (ISAISS 2026)
Volume/ Article
01 / 01
Pages
01 - 05
Article Type
Paper
KeywordsDeepfake audio, voice cloning, zero-shot TTS, voice conversion, audio forensics, anti-forensic detection, selfsupervised learning, neural codec models
Attachment
Download ⇩

Abstract: There has been tremendous growth in the application of Deep Learning algorithms in recent years, and this has brought about significant advancements in voice synthesis, allowing the possibility of lifelike ‘deepfake audio’ to be generated with models such as Transformers, GANs, and hybrids. These developments have enriched applications in terms of customization and ease of use but simultaneously bring about high risks associated with cybersecurity, social engineering, and misinformation with respect to ‘deepfake audio’ because of its possibility due to such advancements. This article review has considered 50 recent studies on ‘deepfake voice’ creation and detection to discuss its full application in this paper, developed in two stages. Stage 1 introduces synthesis techniques such as ‘prosody with attention,’ ‘diffusion models for diversity,’ and ‘GAN-based voice synthesis’ to describe recent advancements in ‘multilingual, zero-shot, and expressive capabilities. Phase 2 is focused on detection techniques that address spectral, temporal, and multimodal artifacts in addition to traditional methods for robustness in low-resource environments. These techniques include Transformer encoders, CNN hybrids, and self-supervised frameworks. Despite significant advancements, challenges with adversarial attacks, multilingual deepfakes, real-world noise, and generalization to unknown models still exist. In order to protect voice-based technologies in an AI-driven era, the review identifies future directions, that are centered around multimodal integration, continuous learning, and interpretable forensics.