IDEAS home Printed from https://ideas.repec.org/a/gam/jftint/v18y2026i7p379-d1995417.html

Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection

Author

Listed:
  • Shahad Mohammad Bn Dokiey

    (Department of Cybersecurity and Digital Forensics, College of Forensic & Investigative Sciences, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia)

  • Tariq M. Khan

    (Department of Cybersecurity and Digital Forensics, College of Forensic & Investigative Sciences, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia
    Center of Artificial Intelligence for Security, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia)

  • Qazi Emad Ul Haq

    (Department of Cybersecurity and Digital Forensics, College of Forensic & Investigative Sciences, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia
    Center of Artificial Intelligence for Security, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia)

Abstract

Audio-visual deepfake detection remains challenging under cross-dataset distribution shift, especially when the source-domain training data are severely imbalanced. Existing middle-fusion detectors often rely on softmax-based cross-attention, which can learn sharp source-domain token interactions and may transfer poorly to unseen datasets. This study proposes FocalNet, an audio-visual detector that replaces the cross-attention block of the 2D3MF framework with cross-modal focal modulation. The proposed module aggregates multi-scale temporal context before audio-visual interaction, enabling softmax-free contextual modulation between visual MARLIN features and audio EAT features. We evaluate the method under a strict FakeAVCeleb-to-DFDC protocol, where all training and validation is performed on FakeAVCeleb and the DFDC is used only as an unseen target-domain test set. Compared with the reproduced 2D3MF baseline, which collapses to a single-class prediction pattern on the DFDC, FocalNet achieves substantially stronger zero-shot score separation, with a DFDC ROC-AUC of 0.9324. Thresholded analysis further shows an improved balanced accuracy, macro F1 score, and MCC when the frozen source-domain operating point is applied. The model also preserves practical efficiency, requiring comparable FLOPs and lower per-sample inference time than the reproduced baseline. These findings suggest that cross-modal focal modulation is a promising alternative to attention-based middle fusion for audio-visual deepfake detection under dataset shifts, while broader validation across additional unseen datasets, multi-seed training, and deployment-oriented calibration remain important for future work.

Suggested Citation

  • Shahad Mohammad Bn Dokiey & Tariq M. Khan & Qazi Emad Ul Haq, 2026. "Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection," Future Internet, MDPI, vol. 18(7), pages 1-26, July.
  • Handle: RePEc:gam:jftint:v:18:y:2026:i:7:p:379-:d:1995417
    as

    Download full text from publisher

    File URL: https://www.mdpi.com/1999-5903/18/7/379/pdf
    Download Restriction: no

    File URL: https://www.mdpi.com/1999-5903/18/7/379/
    Download Restriction: no
    ---><---

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jftint:v:18:y:2026:i:7:p:379-:d:1995417. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address (email available below). General contact details of provider: https://www.mdpi.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.