Publication
Can SSL Frontend Generalize to All-Type Audio Spoofing?
Arnab Das; Yassine El Kheir; Fabian Ritter Guttierez; Tim Polzehl; Sebastian Möller (Hrsg.)
Odyssey Workshop on Speaker and Language Recognition (Odyssey-2026), The Speaker and Language Recognition Workshop, June 23-26, Lisbon, Portugal, Proc. Odyssey 2026, Proc. Odyssey 2026, 2026.
Abstract
Advances in generative AI have increased the exposure of automatic speaker verification systems to synthetic audio across diverse modalities. Recent spoofing detection approaches typically rely on self-supervised learning (SSL) frontends followed by neural classifiers. While such representations achieve strong performance within specific modalities, their ability to generalise across diverse deepfake types has not been systematically investigated. In this work, we benchmark SSL frontends for cross-modal audio deepfake detection spanning speech, environmental sounds, music, and singing voice. Our experiments reveal that no single SSL model generalises well across all modalities and that pre-training data distribution strongly influences detection performance. We show that a dual-frontend architecture combining a speech-specific and a general-audio SSL model yields the strongest cross-modal performance, reducing average equal error rate compared to the best single-frontend baseline.
