Skip to main content Skip to main navigation

Publikation

Fake Faces, Real Bias: A Multi-Ethnic Benchmark for Deepfake Detection

Mario Schweikert; Andrey Guzhov; Ko Watanabe; Andreas Dengel; Isao Echizen; Junichi Yamagishi; Tobias Wirth
In: Proceedings of the 2026 IEEE International Joint Conference on Biometrics (IJCB). IEEE/IAPR International Joint Conference on Biometrics (IJCB-2026), September 1-4, Rom, Italy, IEEE/IAPR, 2026.

Zusammenfassung

Deepfake detectors are gatekeepers of digital identity. Yet whether they guard all humans equally remains an open question. We introduce FairFaceBench, a demographically controlled video dataset built on the FHIBE dataset, spanning nine world regions, two genders, and seven generation methods including state-ofthe- art avatar synthesis systems (Gemini Veo 3.1, Sora 2), with verified nationality and gender metadata. We evaluate four detectors (DINOv2 by Synthetiq Vision, Xception, EfficientNetB4, and F3Net from DeepfakeBench) and analyze bias from two complementary perspectives: false-alarm disparities on real faces, which cause legitimate users to be wrongly denied service and miss-rate disparities on synthetic faces, which allow demographically targeted fraud to bypass detection. Our analysis reveals that (i) the newest generative avatar models evade all tested detectors, and (ii) all four tested detectors exhibit substantial demographic bias, with worst-group AUC below 52 % and mean subgroup AUC deviation up to 9.7,pp for zero-shot baselines. To address these disparities, we fine-tune all four detectors on FairFaceBench and show that overall AUC improves by 19-26,pp and worst-group AUC by 25-42,pp. Fine-tuning incurs an expected distribution-shift penalty on FF++ in-distribution data (7-24,pp AUC drop), which is partially mitigated by a worst-group loss objective that additionally improves fairness across generation categories.

Projekte