AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol. In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T12:12:29.144Z. This is not the publication date.