DF-TransFusion: Multimodal Deepfake Detection via Lip-Audio Cross-Attention and Facial Self-Attention
arXiv:2309.06511 · cs.CV, cs.MM · September 2023
A multi-modal framework that processes audio and video concurrently. A cross-attention mechanism scores lip synchronization against the input audio while a fine-tuned VGG-16 extracts visual cues; a transformer encoder then applies facial self-attention. The approach exceeds prior multimodal detectors on F-1 and per-video AUC.