2026
|
Brodarič, Marko; Scheirer, Walter; Jain, Deepak Kumar; Peer, Peter; Štruc, Vitomir Learning Forgery Signals from Synthetic Depth Targets for Generalizable Deepfake Detection Proceedings Article In: Proceedings of the IEEE International Joint Conference on Biometrics (IJCB 2026), pp. 1–10, 2026. @inproceedings{BrodaricIJCB2026,
title = {Learning Forgery Signals from Synthetic Depth Targets for Generalizable Deepfake Detection},
author = {Marko Brodarič and Walter Scheirer and Deepak Kumar Jain and Peter Peer and Vitomir Štruc},
url = {https://lmi.fe.uni-lj.si/wp-content/uploads/2026/07/PaperID138_camera_ready_compressed.pdf
https://lmi.fe.uni-lj.si/wp-content/uploads/2026/07/PaperID138_supplementary_camera_ready_compressed.pdf},
year = {2026},
date = {2026-09-01},
booktitle = {Proceedings of the IEEE International Joint Conference on Biometrics (IJCB 2026)},
pages = {1--10},
abstract = {Deepfake detection remains challenging under distribution shifts caused by unseen generation pipelines, post-processing, and unconstrained capture conditions. While most existing detectors primarily reason in RGB appearance space, facial manipulations also disturb geometric structure and consistency, making depth an attractive complementary cue. Existing depth-aware methods, however, typically use depth as an auxiliary input or guidance signal and may therefore underutilize the representational richness of modern monocular depth foundation models. To address this limitation, we propose a mechanism that adapts a pretrained monocular depth foundation model into a forgery-aware representation extractor for deepfake detection. Based on the adapted model, we then propose a novel depth-driven detector, termed {FADepth}, that uses the resulting depth representation as the primary source of evidence and complements it with RGB appearance cues. Extensive experiments on six benchmark datasets and with 24 state-of-the-art detectors show that FADepth yields highly competitive performance, achieving an overall mAUC of 88.70. The source code of the model is available at https://github.com/markobrodaric/FADepth},
keywords = {deepfake, deepfake DAD, deepfake detection, deepfakes, information forensics, media forensics},
pubstate = {published},
tppubtype = {inproceedings}
}
Deepfake detection remains challenging under distribution shifts caused by unseen generation pipelines, post-processing, and unconstrained capture conditions. While most existing detectors primarily reason in RGB appearance space, facial manipulations also disturb geometric structure and consistency, making depth an attractive complementary cue. Existing depth-aware methods, however, typically use depth as an auxiliary input or guidance signal and may therefore underutilize the representational richness of modern monocular depth foundation models. To address this limitation, we propose a mechanism that adapts a pretrained monocular depth foundation model into a forgery-aware representation extractor for deepfake detection. Based on the adapted model, we then propose a novel depth-driven detector, termed {FADepth}, that uses the resulting depth representation as the primary source of evidence and complements it with RGB appearance cues. Extensive experiments on six benchmark datasets and with 24 state-of-the-art detectors show that FADepth yields highly competitive performance, achieving an overall mAUC of 88.70. The source code of the model is available at https://github.com/markobrodaric/FADepth |
Hrvatič, Anja; Peer, Peter; Štruc, Vitomir; Batagelj, Borut I-SPAI: A Human Perception Dataset for Deepfake Detection and Perception-Aware Auxiliary Supervision Proceedings Article In: Proceedings of the IEEE International Joint Conference on Biometrics (IJCB 2026), pp. 1–10, 2026. @inproceedings{AnjaIJCB2026,
title = {I-SPAI: A Human Perception Dataset for Deepfake Detection and Perception-Aware Auxiliary Supervision},
author = {Anja Hrvatič and Peter Peer and Vitomir Štruc and Borut Batagelj},
url = {https://lmi.fe.uni-lj.si/wp-content/uploads/2026/07/IJCB_2026_compressed.pdf},
year = {2026},
date = {2026-09-01},
booktitle = {Proceedings of the IEEE International Joint Conference on Biometrics (IJCB 2026)},
pages = {1--10},
abstract = {Deepfake detectors typically overfit to low-level artifacts of the generative
models seen during training, leading to poor generalization on unseen manipulation
techniques. This work explores human visual perception as an alternative and more
stable source of supervision. We introduce I-SPAI (textbf{I}nconsistency,
textbf{S}aliency, and textbf{P}erception textbf{A}nnotations for Deepfake
textbf{I}dentification), the first publicly available deepfake dataset with human
perception annotations, consisting of 2,880 SBI-generated image pairs labeled by
435 participants with inconsistency types, spatial saliency maps, and confidence
scores. Using I-SPAI, we propose a multitask detection framework and conduct a
systematic evaluation of three types of human perceptual supervision: categorical
inconsistency labels, spatial saliency map prediction, and indirect saliency-based
loss supervision via CYBORG. Results on four benchmark datasets show that not all
human perception signals are equally effective. Spatial saliency maps incorporated
via a consecutive decoder provide the most consistent cross-dataset improvement,
while categorical and indirect supervision offer limited gains. This suggests that
spatial localization is the most transferable form of human perceptual knowledge
for deepfake detection, and that the form of the supervision signal matters as
much as its source.},
keywords = {deepfake, deepfake DAD, deepfake detection, deepfakes, media forensics},
pubstate = {published},
tppubtype = {inproceedings}
}
Deepfake detectors typically overfit to low-level artifacts of the generative
models seen during training, leading to poor generalization on unseen manipulation
techniques. This work explores human visual perception as an alternative and more
stable source of supervision. We introduce I-SPAI (textbf{I}nconsistency,
textbf{S}aliency, and textbf{P}erception textbf{A}nnotations for Deepfake
textbf{I}dentification), the first publicly available deepfake dataset with human
perception annotations, consisting of 2,880 SBI-generated image pairs labeled by
435 participants with inconsistency types, spatial saliency maps, and confidence
scores. Using I-SPAI, we propose a multitask detection framework and conduct a
systematic evaluation of three types of human perceptual supervision: categorical
inconsistency labels, spatial saliency map prediction, and indirect saliency-based
loss supervision via CYBORG. Results on four benchmark datasets show that not all
human perception signals are equally effective. Spatial saliency maps incorporated
via a consecutive decoder provide the most consistent cross-dataset improvement,
while categorical and indirect supervision offer limited gains. This suggests that
spatial localization is the most transferable form of human perceptual knowledge
for deepfake detection, and that the form of the supervision signal matters as
much as its source. |
Larue, Nicolas; Štruc, Vitomir; Peer, Peter; Vu, Ngoc-Son Learning the Manifold of Authenticity: Hybrid-Curvature Representation Learning for Generalizable Deepfake Detection Journal Article In: IEEE Access, pp. 1–14, 2026, ISBN: 2169-3536. @article{AccessHyperbolic,
title = {Learning the Manifold of Authenticity: Hybrid-Curvature Representation Learning for Generalizable Deepfake Detection},
author = {Nicolas Larue and Vitomir Štruc and Peter Peer and Ngoc-Son Vu},
url = {https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11557307},
doi = {10.1109/ACCESS.2026.3702429},
isbn = {2169-3536},
year = {2026},
date = {2026-06-06},
journal = {IEEE Access},
pages = {1--14},
abstract = {The practical utility of deepfake detectors is crippled by a crisis of generalization: models that perform well on known manipulation techniques consistently fail when faced with unseen forgeries.We argue this failure stems from a fundamental geometric mismatch. Existing methods implicitly assume that the manifold of authentic faces can be modeled in a space of uniform curvature, typically Euclidean, which inade-quately captures the complex, multi-scale structure of facial features. This paper validates the hypothesis that authentic faces lie on a manifold whose geometry is inherently hybrid, requiring both angular compactness (a spherical property) and hierarchical organization (a hyperbolic property). To resolve this geometric mismatch, we introduce a novel detector, CTrue, that learns a unified, hybrid-curvature representation of facial authenticity. Trained exclusively on real faces via self-supervised learning, our method simultaneously projects facial embeddings onto two complementary manifolds: a hypersphere to enforce compactness and a hyperbolic space to model the natural feature hierarchy. A single set of mathematically-optimal prototypes acts as a ‘‘geometric bridge’’, unifying the learning objectives in both spaces. At inference, a composite score measures an embedding’s deviation from this learned manifold. On challenging cross-dataset and cross-manipulation benchmarks, our method achieves competitive generalization under a strictly pristine-only training setting, showing that hybrid-curvature representations provide an effective and data-efficient alternative for deepfake detection.},
keywords = {deep learning, deepfake, deepfake DAD, deepfake detection, hyperbolic learning, media forensics},
pubstate = {published},
tppubtype = {article}
}
The practical utility of deepfake detectors is crippled by a crisis of generalization: models that perform well on known manipulation techniques consistently fail when faced with unseen forgeries.We argue this failure stems from a fundamental geometric mismatch. Existing methods implicitly assume that the manifold of authentic faces can be modeled in a space of uniform curvature, typically Euclidean, which inade-quately captures the complex, multi-scale structure of facial features. This paper validates the hypothesis that authentic faces lie on a manifold whose geometry is inherently hybrid, requiring both angular compactness (a spherical property) and hierarchical organization (a hyperbolic property). To resolve this geometric mismatch, we introduce a novel detector, CTrue, that learns a unified, hybrid-curvature representation of facial authenticity. Trained exclusively on real faces via self-supervised learning, our method simultaneously projects facial embeddings onto two complementary manifolds: a hypersphere to enforce compactness and a hyperbolic space to model the natural feature hierarchy. A single set of mathematically-optimal prototypes acts as a ‘‘geometric bridge’’, unifying the learning objectives in both spaces. At inference, a composite score measures an embedding’s deviation from this learned manifold. On challenging cross-dataset and cross-manipulation benchmarks, our method achieves competitive generalization under a strictly pristine-only training setting, showing that hybrid-curvature representations provide an effective and data-efficient alternative for deepfake detection. |
2025
|
Batagelj, Borut; Kronovšek, Andrej; Štruc, Vitomir; Peer, Peter Robust cross-dataset deepfake detection with multitask self-supervised learning Journal Article In: ICT Express, pp. 1-5, 2025. @article{DeepFake2025,
title = {Robust cross-dataset deepfake detection with multitask self-supervised learning},
author = {Borut Batagelj and Andrej Kronovšek and Vitomir Štruc and Peter Peer},
url = {https://www.sciencedirect.com/science/article/pii/S240595952500027X?via%3Dihub},
doi = {https://doi.org/10.1016/j.icte.2025.02.011},
year = {2025},
date = {2025-03-02},
journal = {ICT Express},
pages = {1-5},
abstract = {Deepfake detection is increasingly critical due to the rise of manipulated media. Existing methods often require extensive datasets and struggle with interpretability issues. To address these issues, this study introduces a novel one-class approach for detecting and localizing deepfake artifacts in videos, using authentic images to generate manipulated data for training. By integrating segmentation and leveraging convolutional neural networks with visual transformers, the method predicts both the presence and location of the generated manipulations. Experiments on seven deepfake datasets and emerging diffusion-based manipulations show that our approach consistently outperforms existing methods, demonstrating superior accuracy and localization capabilities.},
keywords = {deepfake, deepfake DAD, deepfake detection, multi-task learning, segmentation},
pubstate = {published},
tppubtype = {article}
}
Deepfake detection is increasingly critical due to the rise of manipulated media. Existing methods often require extensive datasets and struggle with interpretability issues. To address these issues, this study introduces a novel one-class approach for detecting and localizing deepfake artifacts in videos, using authentic images to generate manipulated data for training. By integrating segmentation and leveraging convolutional neural networks with visual transformers, the method predicts both the presence and location of the generated manipulations. Experiments on seven deepfake datasets and emerging diffusion-based manipulations show that our approach consistently outperforms existing methods, demonstrating superior accuracy and localization capabilities. |
Soltandoost, Elahe; Plesh, Richard; Schuckers, Stephanie; Peer, Peter; Struc, Vitomir Extracting Local Information from Global Representations for Interpretable Deepfake Detection Proceedings Article In: Proceedings of IEEE/CFV Winter Conference on Applications in Computer Vision - Workshops (WACV-W) 2025, pp. 1-11, Tucson, USA, 2025. @inproceedings{Elahe_WACV2025,
title = {Extracting Local Information from Global Representations for Interpretable Deepfake Detection},
author = {Elahe Soltandoost and Richard Plesh and Stephanie Schuckers and Peter Peer and Vitomir Struc},
url = {https://lmi.fe.uni-lj.si/wp-content/uploads/2025/01/ElahePaperF.pdf},
year = {2025},
date = {2025-03-01},
booktitle = {Proceedings of IEEE/CFV Winter Conference on Applications in Computer Vision - Workshops (WACV-W) 2025},
pages = {1-11},
address = {Tucson, USA},
abstract = {The detection of deepfakes has become increasingly challenging due to the sophistication of manipulation techniques that produce highly convincing fake videos. Traditional detection methods often lack transparency and provide limited insight into their decision-making processes. To address these challenges, we propose in this paper a Locally-Explainable Self-Blended (LESB) DeepFake detector that in addition to the final fake-vs-real classification decision also provides information, on which local facial region (i.e., eyes, mouth or nose) contributed the most to the decision process.~At the heart of the detector is a novel Local Feature Discovery (LFD) technique that can be applied to the embedding space of pretrained DeepFake detectors and allows identifying embedding space directions that encode variations in the appearance of local facial features. We demonstrate the merits of the proposed LFD technique and LESB detector in comprehensive experiments on four popular datasets, i.e., Celeb-DF, DeepFake Detection Challenge, Face Forensics in the Wild and FaceForensics++, and show that the proposed detector is not only competitive in comparison to strong baselines, but also exhibits enhanced transparency in the decision-making process by providing insights on the contribution of local face parts in the final detection decision. },
keywords = {CNN, deepfake DAD, deepfakes, faceforensics++, media forensics, xai},
pubstate = {published},
tppubtype = {inproceedings}
}
The detection of deepfakes has become increasingly challenging due to the sophistication of manipulation techniques that produce highly convincing fake videos. Traditional detection methods often lack transparency and provide limited insight into their decision-making processes. To address these challenges, we propose in this paper a Locally-Explainable Self-Blended (LESB) DeepFake detector that in addition to the final fake-vs-real classification decision also provides information, on which local facial region (i.e., eyes, mouth or nose) contributed the most to the decision process.~At the heart of the detector is a novel Local Feature Discovery (LFD) technique that can be applied to the embedding space of pretrained DeepFake detectors and allows identifying embedding space directions that encode variations in the appearance of local facial features. We demonstrate the merits of the proposed LFD technique and LESB detector in comprehensive experiments on four popular datasets, i.e., Celeb-DF, DeepFake Detection Challenge, Face Forensics in the Wild and FaceForensics++, and show that the proposed detector is not only competitive in comparison to strong baselines, but also exhibits enhanced transparency in the decision-making process by providing insights on the contribution of local face parts in the final detection decision. |
2024
|
Dragar, Luka; Rot, Peter; Peer, Peter; Štruc, Vitomir; Batagelj, Borut W-TDL: Window-Based Temporal Deepfake Localization Proceedings Article In: Proceedings of the 2nd International Workshop on Multimodal and Responsible Affective Computing (MRAC ’24), Proceedings of the 32nd ACM International Conference on Multimedia (MM’24), ACM, 2024. @inproceedings{MRAC2024,
title = {W-TDL: Window-Based Temporal Deepfake Localization},
author = {Luka Dragar and Peter Rot and Peter Peer and Vitomir Štruc and Borut Batagelj},
url = {https://lmi.fe.uni-lj.si/wp-content/uploads/2024/09/ACM_1M_DeepFakes.pdf},
year = {2024},
date = {2024-11-01},
booktitle = {Proceedings of the 2nd International Workshop on Multimodal and Responsible Affective Computing (MRAC ’24), Proceedings of the 32nd ACM International Conference on Multimedia (MM’24)},
publisher = {ACM},
abstract = {The quality of synthetic data has advanced to such a degree of realism that distinguishing it from genuine data samples is increasingly challenging. Deepfake content, including images, videos, and audio, is often used maliciously, necessitating effective detection methods. While numerous competitions have propelled the development of deepfake detectors, a significant gap remains in accurately pinpointing the temporal boundaries of manipulations. Addressing this, we propose an approach for temporal deepfake localization (TDL) utilizing a window-based method for audio (W-TDL) and a complementary visual frame-based model. Our contributions include an effective method for detecting and localizing fake video and audio segments and addressing unbalanced training labels in spoofed audio datasets. Our approach leverages the EVA visual transformer for frame-level analysis and a modified TDL method for audio, achieving competitive results in the 1M-DeepFakes Detection Challenge. Comprehensive experiments on the AV-Deepfake1M dataset demonstrate the effectiveness of our method, providing an effective solution to detect and localize deepfake manipulations.},
keywords = {CNN, deepfake DAD, deepfakes, deeplearning, detection, localization},
pubstate = {published},
tppubtype = {inproceedings}
}
The quality of synthetic data has advanced to such a degree of realism that distinguishing it from genuine data samples is increasingly challenging. Deepfake content, including images, videos, and audio, is often used maliciously, necessitating effective detection methods. While numerous competitions have propelled the development of deepfake detectors, a significant gap remains in accurately pinpointing the temporal boundaries of manipulations. Addressing this, we propose an approach for temporal deepfake localization (TDL) utilizing a window-based method for audio (W-TDL) and a complementary visual frame-based model. Our contributions include an effective method for detecting and localizing fake video and audio segments and addressing unbalanced training labels in spoofed audio datasets. Our approach leverages the EVA visual transformer for frame-level analysis and a modified TDL method for audio, achieving competitive results in the 1M-DeepFakes Detection Challenge. Comprehensive experiments on the AV-Deepfake1M dataset demonstrate the effectiveness of our method, providing an effective solution to detect and localize deepfake manipulations. |