REVIEW 3 major objections 4 minor 128 references
Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version)
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A speaker can embed a signed semantic fingerprint into every video of a live speech using imperceptible modulated light, so later deepfaked copies fail a content check.
desk verdict VeriLight is a credible, well-engineered physical-signature system for speech video verification; the headline detection numbers are in-sample, but the concept and the AUC evidence are strong enough to warrant serious peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the locality-sensitive hashing framework used to compress high-dimensional semantic features into 150-bit descriptors. Cosine-similarity LSH maps similar feature vectors to similar hashes, so small pose- and capture-induced differences between legitimate recordings do not change the hash much, while a semantically different face or lip motion does; a derived closed-form expression shows this verification performance is independent of the input dimension, justifying hashing 512-dimensional identity embeddings and long dynamic feature vectors into the same small size. The second mechanism is the spatio-temporal optical modulation scheme: an amplitude-modulating spatial light modulator projects a cell grid onto a planar surface, each data cell carrying one bit per two bitmaps through phase-keyed intensity changes, with border cells acting as synchronization references and corner cells as localization beacons. An adaptive control loop adjusts each cell's color and intensity to balance bit error rate against perceptual indistinguishability, which is what lets the signature survive compression, transcoding, and diverse surfaces.
What would settle it
Run VeriLight on a corpus of reenactment deepfakes from a model not in the tested set, such as a newer lip-sync model, that changes a single word in less than 20 percent of a 4.5-second window; if even one such fake is verified as real, the claim that the selected facial-analysis features catch content-changing falsifications fails for that regime.
Extended reading notes
Core claim
The central discovery is that semantic speech-content features can be compressed cryptographically and delivered through light in real time. The signature creation module runs a real-time facial analysis model and a face-embedding model on a window of video, concatenates 16 dynamic signals (5 lip landmark distances and 11 blendshapes) plus a 512-dimensional identity embedding, and hashes each with cosine locality-sensitive hashing into 150 bits. These hashes, together with window metadata and an HMAC-SHA256 message authentication code truncated to 80 bits, form the signature. The embedding module drives a spatial light modulator with a 16x9 grid of cells using binary phase-shift keying at 3 Hz, plus synchronization and localization cells, achieving over 200 bits per second. Verification localizes the cells by a Fourier heatmap, recovers the coded data through Viterbi and Reed-Solomon decoding, validates the MAC, and compares hashes by Hamming distance; if identity or dynamic distance exceeds its threshold, the video is declared falsified and the type of falsification is indicated. The paper reports perfect recall on over 2,000 deepfaked videos, with identity swaps and reenactments separated by which hash trips the threshold.
Load-bearing premise
The load-bearing premise is that the 16 facial-analysis signals selected in Section 4.2 are rich and pose-invariant enough that any falsification changing what the speaker appears to say visibly alters at least one hashed signal beyond the decision threshold, across recording angles up to 60 degrees, distances up to 3 meters, and any future deepfake model.
Editorial extensions
If this is right
- A speech venue with a VeriLight core unit produces a canonical, self-authenticating record: every attendee's camera captures the signature without any app, specialized hardware, or cooperation from the camera vendor.
- Compression, transcoding, brightness and contrast edits, and aesthetic filters do not break verification, because the descriptor lives at a semantic level and the optical code is error-corrected.
- VeriLight can report what kind of falsification it detected: identity-swap fakes raise only the identity-hash distance, while reenactment fakes raise only the dynamic-hash distance.
- For fine-grained edits, VeriLight remains usable down to roughly 1.35 seconds of modification within a 4.5-second window, where it reports an AUC of 0.90; partial-window edits below about 20 percent of a window are the hardest case, with AUC dropping to 0.72.
- Replay and signature-injection attacks fail because the descriptor is speech-specific and the MAC is keyed; an attacker cannot mint a valid signature for altered content.
Reading between the lines
- An implicit consequence is that the choice of the 16 facial-analysis features bounds the system's guarantee: attacks that alter speech through cues outside that set, such as changing intonation or word order without changing lip geometry, would not be flagged.
- A testable extension is to re-run the forward feature selection on a larger, multi-language corpus; if different features are selected there, the current 16-signal set may be overfit to the authors' English-paragraph corpus.
- If a future deepfake model learns to modify lip motion while keeping the 16 facial analysis signals nearly unchanged, VeriLight's dynamic threshold would need to be tightened, at the cost of more false positives from facial-tracking noise.
- The still-camera requirement is a deployment constraint rather than a conceptual one; coupling verification with video stabilization could extend the scheme to handheld recordings, and projecting onto the speaker's face would remove the need for a visible projection region.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes VeriLight, a system that embeds cryptographically secured optical signatures into live speech events by projecting imperceptible modulated light onto a surface near the speaker. The embedded signature carries a compact, locality-sensitive-hashed descriptor of the speaker's identity and lip/face motion, recovered later from arbitrary downstream videos and compared with descriptors computed from the portrayed content. The authors report AUCs of at least 0.99 and a 100% true positive rate on over 2,000 deepfaked videos, robustness across recording angles, distances, devices, surfaces and post-processing, and resistance to two white-box adversarial attacks. The paper includes a hardware prototype, a large evaluation corpus, artifact links, and a theoretical analysis of LSH hash size. The main concern is that the headline recall and some robustness numbers are computed on the same data used for threshold setting and feature selection, so the reported operating point is not yet an out-of-sample estimate.
Significance. If the results hold, VeriLight is a meaningful contribution to media authentication: it shifts protection from recording-device cooperation to the event site, provides an end-to-end hardware and software prototype, and demonstrates broad empirical separation between real and falsified videos. The evaluation is unusually extensive for the area, covering 1,883 reenactment and 473 identity-swap deepfakes, multiple cameras, surfaces, lighting conditions, distances, post-processing operations, and adversarial training objectives. The LSH-based compact descriptor framework is a useful general tool, and the explicit release of code and hardware reproduction guides strengthens the work. The central limitation is that several performance claims are in-sample: the decision thresholds are tuned on the section 8 data and the dynamic feature subset is selected on the section 9.1 data, so the 100% recall figure should be understood as an optimistic, threshold-fit estimate rather than a validated deployment guarantee.
major comments (3)
- [§7 and §8.1 (Table 1)] The paper states in §7 that "We empirically set our descriptor comparison decision thresholds using data from §8," and Table 1 then reports a 100% true positive rate on the videos of §8.1. Since the thresholds are fit to the same videos whose recall is reported, the 100% recall is an in-sample operating point, not an out-of-sample detection guarantee. Please split the data at the participant/session level into calibration and evaluation sets, or use nested resampling, and report recall at the resulting held-out threshold. At minimum, report a precision-recall sweep over thresholds and qualify the abstract and §8.1 claims accordingly.
- [§4.2 and §9.1 (Figure 8, Table 3)] The 16 dynamic FaceMesh signals are selected via forward sequential feature selection on the multi-camera dataset described in §9.1, and the same dataset is then used to evaluate pose invariance, modification granularity, and hash-size effects. If the reported AUCs are computed on the same recordings that determined the feature subset, the robustness numbers are optimistically biased. Please use nested feature selection with held-out sessions or cameras, or otherwise state explicitly whether any test recording contributed to feature or hash-size selection.
- [Appendix A.3, Theorem 2] Theorem 2 is presented as a proof that hash-size effects are independent of input dimensionality, but the proof as written is not a well-defined probability statement. It treats a single hash comparison as a Bernoulli event and multiplies over a continuum of pairs, and the integral expression uses a constant summation bound k*theta_th/pi even though the conditioned angle varies inside the integral. Please either provide a rigorous concentration argument over a finite comparison set or clearly label the derivation as an approximation, since the current text overstates the mathematical support for the hash-size choice.
minor comments (4)
- [§8.1] Please state the exact denominator for the 100% recall after excluding the 36 videos in which the signature could not be localized, and clarify whether a verification failure on a fake video is counted as detection, abstention, or exclusion.
- [Appendix A.2] There is a typo: "FaceMech" should be "FaceMesh."
- [Figure 12 and §A.3] The text calls theta_th a "cosine similarity decision threshold," but theta_th is an angle, not a cosine similarity; please use consistent terminology.
- [§7] The sentence "We empirically set our descriptor comparison decision thresholds using data from §8" should be moved to the experimental section and accompanied by a description of the calibration protocol, including how many videos were used and how thresholds were chosen.
Circularity Check
Headline 100% true positive rate is in-sample: decision thresholds are tuned on the same §8 videos used to report recall, and the dynamic features are selected on the same §9.1 dataset used to report their AUCs.
-
fitted input called prediction
[Section 7, Prototype Implementation; evaluation in Section 8.1]
"We empirically set our descriptor comparison decision thresholds using data from §8."
The decision thresholds are fitted parameters, and Section 8.1 then reports the true positive rate and final binary decisions on the same §8 dataset. The 100% recall is therefore the recall at thresholds tuned on the very videos whose recall is being reported; it is an in-sample operating point, not an out-of-sample prediction. The threshold-independent AUCs remain informative, but the deployment-facing claim that VeriLight detects all 2,000+ deepfaked videos is statistically forced by the threshold-fitting protocol described here.
-
fitted input called prediction
[Section 4.2, Semantically-Meaningful Video Descriptors; evaluation in Section 9.1]
"We identify this set of features as optimal via forward sequential feature selection [47] on all blendshape and distance signals, using a comprehensive multi-camera dataset (§9.1)."
The 16 dynamic FaceMesh signals are chosen on the multi-camera dataset, and Section 9.1 reports dynamic-feature AUCs (Figure 8, Table 3) on that same dataset. The reported fine-grained and pose-invariance performance is thus an in-sample evaluation of features explicitly selected to maximize performance on those recordings; it does not measure generalization to new poses, speakers, or manipulation types. This inflates the dynamic-feature AUCs as estimates of out-of-sample detection capability.
full rationale
The core architecture is not circular: the LSH theory is taken from Charikar (an external reference), the identity and face features come from pretrained ArcFace and MediaPipe models, the MAC mechanism is standard cryptography, and the optical embedding robustness and perceptibility results (BER tables, LPIPS, user study) are self-contained evaluations not derived from the detection claim. There is also no load-bearing self-citation chain: the authors' own prior work is cited only for screen-camera communication context, not as the justification for VeriLight's central mechanism. The circularity is concentrated in the headline detection evaluation. Section 7 states that descriptor comparison decision thresholds are empirically set using data from §8, and §8.1 then reports a 100% true positive rate on that same data; the recall number is an in-sample fit rather than a validated operating point. Separately, §4.2 selects the 16 dynamic features by forward sequential feature selection on the §9.1 multi-camera dataset, and §9.1 reports AUCs on that same dataset, making those AUCs partly a product of feature selection on the test set. The threshold-independent AUCs in §8 (Table 1) retain independent content, and the optical-signature and imperceptibility evaluations are not affected, so the circularity is partial rather than total. Score 6 reflects that the central empirical claim (100% TPR and some fine-grained AUCs) reduces to in-sample fitting, while the system's technical derivation remains substantially independent.
Assumptions & free parameters
free parameters (4)
- Descriptor comparison decision thresholds (identity and dynamic hash Hamming distances) =
not stated, set using data from Section 8
- Dynamic feature subset (5 lip distances + 11 blendshapes) =
16 signals selected from all FaceMesh distances and blendshapes
- LSH hash size k =
150 bits
- Optical modulation parameters (fd=3 Hz, fl=6 Hz, window=4.5 s, beta_max=0, phi_max=5, delta=5) =
empirically set
assumptions (6)
- standard math The cosine-similarity LSH scheme H_cos produces k-bit hashes such that Hamming distance estimates the angle between input vectors (Charikar 2002).
- domain assumption HMAC-SHA256 with a 128-bit secret key held by the core unit and the verification cloud provides message authentication of the descriptor.
- domain assumption FaceMesh blendshape scores and lip landmark distances are pose-invariant and accurate up to 60 degrees off-axis and 3 m, and ArcFace embeddings are reliable for identity comparison at those poses.
- domain assumption Every video to be verified is recorded with a still camera that captures the full optical projection region at at least 1080p and 24 FPS.
- domain assumption The adversary does not have access to the secret key that generates MACs.
- standard math Four localization cells can be detected in the heatmap and define a valid homography between the SLM bitmap and the camera view.
Cite this review
Pith. "Pith review of Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version)." pith.science (2026). https://pith.science/paper/HEZVDKIJ
@misc{pith2026250421846,
author = {Pith},
title = {Pith review of: Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version)},
year = {2026},
howpublished = {\url{https://pith.science/paper/HEZVDKIJ}},
note = {Machine review of arXiv:2504.21846}
}
abstract
High-profile speech videos are prime targets for falsification, owing to their accessibility and influence. This work proposes VeriLight, a low-overhead and unobtrusive system for protecting speech videos from visual manipulations of speaker identity and lip and facial motion. Unlike the predominant purely digital falsification detection methods, VeriLight creates dynamic physical signatures at the event site and embeds them into all video recordings via imperceptible modulated light. These physical signatures encode semantically-meaningful features unique to the speech event, including the speaker's identity and facial motion, and are cryptographically-secured to prevent spoofing. The signatures can be extracted from any video downstream and validated against the portrayed speech content to check its integrity. Key elements of VeriLight include (1) a framework for generating extremely compact (i.e., 150-bit), pose-invariant speech video features, based on locality-sensitive hashing; and (2) an optical modulation scheme that embeds $>$200 bps into video while remaining imperceptible both in video and live. Experiments on extensive video datasets show VeriLight achieves AUCs $\geq$ 0.99 and a true positive rate of 100% in detecting falsified videos. Further, VeriLight is highly robust across recording conditions, video post-processing techniques, and white-box adversarial attacks on its feature extraction methods. A demonstration of VeriLight is available at https://mobilex.cs.columbia.edu/verilight.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
deepfake
2019. Doctored Nancy Pelosi video highlights threat of "deepfake" tech. CBS News
2019
-
[2]
Ultra-Light-Fast-Generic-Face-Detector-1MB
2022. Ultra-Light-Fast-Generic-Face-Detector-1MB. GitHub repository. https: //github.com/Linzaer/Ultra-Light-Fast-Generic-Face-Detector-1MB
2022
-
[3]
CVPR2022-DaGAN
2023. CVPR2022-DaGAN. https://github.com/harlanhong/CVPR2022- DaGAN/tree/master?tab=readme-ov-file
2023
-
[4]
Fact Check: Video of Joe Biden calling for a military draft was created with AI
2023. Fact Check: Video of Joe Biden calling for a military draft was created with AI. Reuters
2023
-
[5]
Apple 3DFaceScan
2024. Apple 3DFaceScan. https://apps.apple.com/us/app/3dfacescan-structure- sdk/id6473282888
2024
-
[6]
Content Authenticity Initiative
2024. Content Authenticity Initiative. https://contentauthenticity.org/
2024
-
[7]
Deepfake Kamala Harris slurs her lines
2024. Deepfake Kamala Harris slurs her lines. AI, Algorithmic and Automation Incidents and Controversies
2024
-
[8]
Face landmark detection guide
2024. Face landmark detection guide. https://developers.google.com/mediapip e/solutions/vision/face_landmarker
2024
Show all 128 references
-
[9]
InsightFace: 2D and 3D Face Analysis Project
2024. InsightFace: 2D and 3D Face Analysis Project. https://github.com/deepins ight/insightface/tree/master
2024
-
[10]
Old Kamala Harris footage manipulated to slow her speech
2024. Old Kamala Harris footage manipulated to slow her speech. AFP Factcheck
2024
-
[11]
2024. SynthID. https://deepmind.google/technologies/synthid/
2024
-
[12]
Learned Perceptual Image Patch Similarity (LPIPS)
2025. Learned Perceptual Image Patch Similarity (LPIPS). https://lightning.ai/doc s/torchmetrics/stable/image/learned_perceptual_image_patch_similarity.html
2025
-
[13]
Open Neural Network Exchange
2025. Open Neural Network Exchange. https://onnx.ai/
2025
-
[14]
2025. Truepic. https://truepicvision.com/
2025
-
[15]
VeriLight Github
2025. VeriLight Github. https://github.com/MobileX-CU/verilight
2025
-
[16]
VeriLight Zenodo
2025. VeriLight Zenodo. https://zenodo.org/records/17063747
2025
-
[17]
Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. 2018. Mesonet: a compact facial video forgery detection network. InIEEE International Workshop on Information Forensics and Security
2018
-
[18]
Anubhav Agarwal, CV Jawahar, and PJ Narayanan. 2005. A survey of planar homography estimation techniques.Centre for Visual Information Technology, Tech. Rep(2005)
2005
-
[19]
Shruti Agarwal and Hany Farid. 2021. Detecting Deep-Fake Videos from Aural and Oral Dynamics. InProc. of CVPR Workshops
2021
-
[20]
Shruti Agarwal, Hany Farid, Tarek El-Gaaly, and Ser-Nam Lim. 2020. Detect- ing Deep-Fake Videos from Appearance and Behavior. InIEEE International Workshop on Information Forensics and Security
2020
-
[21]
Shruti Agarwal, Hany Farid, Ohad Fried, and Maneesh Agrawala. 2020. De- tecting Deep-Fake Videos from Phoneme-Viseme Mismatches. InProc. of CVPR Workshops. 2814–2822
2020
-
[22]
Shruti Agarwal, Hany Farid, Yuming Gu, Mingming He, Koki Nagano, and Hao Li. 2019. Protecting World Leaders Against Deep Fakes. InProc. of CVPR Workshops
2019
-
[23]
Shruti Agarwal, Liwen Hu, Evonne Ng, Trevor Darrell, Hao Li, and Anna Rohrbach. 2023. Watch those words: Video falsification detection using word- conditioned facial motion. InIEEE/CVF Winter Conference on Applications of Computer Vision
2023
-
[24]
Prana Air. 2023. Illuminance Levels Indoors: Your Standard Lux Level Chart. https://www.pranaair.com/blog/illuminance-levels-indoors-the-standard- lux-levels
2023
-
[25]
Arsha Nagrani and Joon Son Chung and Weidi Xie and Andrew Zisserman
-
[26]
Nicolo Bonettini, Edoardo Daniele Cannas, Sara Mandelli, Luca Bondi, Paolo Bestagini, and Stefano Tubaro. 2021. Video Face Manipulation Detection Through Ensemble of CNNs. InInternational Conference on Pattern Recognition
2021
-
[27]
Zhixi Cai, Kalin Stefanov, Abhinav Dhall, and Munawar Hayat. 2022. Do You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multi- modal Method for Temporal Forgery Localization. InInternational Conference on Digital Image Computing: Techniques and Applications
2022
-
[28]
Junyi Cao, Chao Ma, Taiping Yao, Shen Chen, Shouhong Ding, and Xiaokang Yang. 2022. End-to-end reconstruction-classification learning for face forgery detection. InProc. of CVPR
2022
-
[29]
Maria Caria, Gabriele Sara, Giuseppe Todde, Marco Polese, and Antonio Pazzona
-
[30]
Nicholas Carlini and Hany Farid. 2020. Evading Deepfake-Image Detectors with White- and Black-Box Attacks.arXiv preprint arXiv:2004.00622(2020)
2020 arXiv
-
[31]
Exploring smart glasses for augmented reality: A valuable and integrative tool in precision livestock farming.Animals9, 11 (2019), 903
2019
-
[32]
Moses S Charikar. 2002. Similarity estimation techniques from rounding algo- rithms. InProc. of ACM Symposium on Theory of Computing. 380–388
2002
-
[33]
Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. InIEEE Symposium on Security and Privacy
2017
-
[34]
Mingliang Chen, Xin Liao, and Min Wu. 2022. PulseEdit: Editing Physiological Signals in Facial Videos for Privacy Protection.IEEE Transactions on Information Forensics and Security17 (2022), 457–471
2022
-
[35]
Hong-Shuo Chen, Mozhdeh Rouhsedaghat, Hamza Ghani, Shuowen Hu, Suya You, and C-C Jay Kuo. 2021. Defakehop: A light-weight high-performance deepfake detector. InIEEE International Conference on Multimedia and Expo
2021
-
[36]
Umur Aybars Ciftci, Ilke Demir, and Lijun Yin. 2020. Fakecatcher: Detection of synthetic portrait videos using biological signals.IEEE Transactions on Pattern Analysis and Machine Intelligence(2020)
2020
-
[37]
Komal Chugh, Parul Gupta, Abhinav Dhall, and Ramanathan Subramanian. 2020. Not made for each other-audio-visual dissonance-based deepfake detection and localization. InACM International Conference on Multimedia
2020
-
[38]
Davide Cozzolino, Andreas Rössler, Justus Thies, Matthias Nießner, and Luisa Verdoliva. 2021. ID-reveal: Identity-aware Deepfake Video Detection. InInter- national Conference on Computer Vision. 15108–15117
2021
-
[39]
Daniel Cotting, Martin Naef, Markus Gross, and Henry Fuchs. 2004. Embedding imperceptible patterns into projected images for simultaneous acquisition and display. InIEEE International Symposium on Mixed and Augmented Reality
2004
-
[40]
Hao Cui, Huanyu Bian, Weiming Zhang, and Nenghai Yu. 2019. UnseenCode: Invisible On-screen Barcode with Image-based Extraction. InProc. of INFOCOM
2019
-
[41]
Andrew Critch. 2022. WordSig: QR streams enabling platform-independent self-identification that’s impossible to deepfake.arXiv preprint arXiv:2207.10806 (2022)
2022 arXiv
-
[42]
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019. Arcface: Additive angular margin loss for deep face recognition. InProc. of CVPR
2019
-
[43]
Farrell S
S. Farrell S. Boeyen R. Housley W. Polk D. Cooper, S. Santesson. 2008. RFC 5208 - Internet X.509 Public Key Infrastructure Certificate and Certificate Revocation List (CRL) Profile. https://datatracker.ietf.org/doc/html/rfc5280
2008
-
[44]
Donovan and Ray Scherer
Robert J. Donovan and Ray Scherer. 2005.Unsilent Revolution: Television News and American Public Life, 1948-1991. Woodrow Wilson International Center for Scholars; Cambridge University Press
2005
-
[45]
Whitfield Diffie and Martin E Hellman. 1976. New directions in cryptography. In Democratizing Cryptography: The Work of Whitfield Diffie and Martin Hellman
1976
-
[46]
Habiba Farrukh, Reham Mohamed Aburas, Siyuan Cao, and He Wang. 2020. FaceRevelio: A Face Liveness Detection System for Smartphones with a Single Front Camera. InProc. of MobiCom
2020
-
[47]
Han Fang, Dongdong Chen, Feng Wang, Zehua Ma, Honggu Liu, Wenbo Zhou, Weiming Zhang, and Nenghai Yu. 2022. TERA: Screen-to-Camera Image Code With Transparency, Efficiency, Robustness and Adaptability.IEEE Transactions on Multimedia24 (2022), 955–967
2022
-
[48]
Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. 2020. Leveraging Frequency Analysis for Deep Fake Image Recognition. InInternational Conference on Machine Learning
2020
-
[49]
Ferri, P
F.J. Ferri, P. Pudil, M. Hatef, and J. Kittler. 1994. Comparative study of techniques for large-scale feature selection. InPattern Recognition in Practice IV. Vol. 16. 403–413
1994
-
[50]
Gerstner and Hany Farid
Candice R. Gerstner and Hany Farid. 2022. Detecting real-time deep-fake videos using active illumination. InProc. of CVPR. 53–60
2022
-
[51]
John S Garofolo. 1993. Timit Acoustic Phonetic Continuous Speech Corpus. Linguistic Data Consortium(1993)
1993
-
[52]
Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2021. Lips don’t lie: A generalisable and robust approach to face forgery detection. InProc. of CVPR
2021
-
[53]
Canetti H
R. Canetti H. Krawczyk, M. Bellare. 1997. HMAC: Keyed-Hashing for Message Authentication. https://datatracker.ietf.org/doc/html/rfc2104#section-5
1997
-
[54]
Shu Hu, Yuezun Li, and Siwei Lyu. 2021. Exposing GAN-generated faces using inconsistent corneal specular highlights. InIEEE International Conference on Acoustics, Speech and Signal Processing
2021
-
[55]
Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. 2022. Depth-Aware Generative Adversarial Network for Talking Head Video Generation.Proc. of CVPR. Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version) CCS ’25, October 13–17, 2025, Taipei, Taiwan
2022
-
[56]
Kensei Jo, Mohit Gupta, and Shree K. Nayar. 2016. DisCo: Display-Camera Communication Using Rolling Shutter Sensors.ACM Trans. Graph.35, 5 (2016)
2016
-
[57]
Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. 2008. Labeled faces in the wild: A database forstudying face recognition in uncon- strained environments. InWorkshop on faces in ’Real-Life’ Images: detection, alignment, and recognition
2008
-
[58]
Cassidy Laidlaw, Sahil Singla, and Soheil Feizi. 2020. Perceptual adversarial ro- bustness: Defense against unseen threat models.arXiv preprint arXiv:2006.12655 (2020)
2020 arXiv
-
[59]
Fouad Khelifi and Ahmed Bouridane. 2017. Perceptual video hashing for content identification and authentication.IEEE Transactions on Circuits and Systems for Video Technology29, 1 (2017), 50–67
2017
-
[60]
Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. 2019. Face X-Ray for More General Face Forgery Detection.Proc. of CVPR(2019)
2019
-
[61]
Hui-Yu Lee, Hao-Min Lin, Yu-Lin Wei, Hsin-I Wu, Hsin-Mu Tsai, and Kate Ching-Ju Lin. 2015. RollingLight: Enabling Line-of-Sight Light-to-Camera Com- munications. InProc. of MobiSys
2015
-
[62]
Yuezun Li, Ming-Ching Chang, and Siwei Lyu. 2018. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. InIEEE International Workshop on Information Forensics and Security
2018
-
[63]
Campbell, and Xia Zhou
Tianxing Li, Chuankai An, Xinran Xiao, Andrew T. Campbell, and Xia Zhou
-
[64]
Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weim- ing Zhang, and Nenghai Yu. 2021. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. InProc. of CVPR
2021
-
[65]
Hongbo Liu, Zhihua Li, Yucheng Xie, Ruizhe Jiang, Yan Wang, Xiaonan Guo, and Yingying Chen. 2020. LiveScreen: Video Chat Liveness Detection Leveraging Skin Reflection. InProc. of INFOCOM
2020
-
[66]
Feng Liu, Michael Gleicher, Jue Wang, Hailin Jin, and Aseem Agarwala. 2011. Subspace video stabilization.ACM Trans. on Graphics30, 1 (2011), 1–10
2011
-
[67]
Yuxin (Myles) Liu, Zhihao Yao, Mingyi Chen, Ardalan Amiri Sani, Sharad Agar- wal, and Gene Tsudik. 2024. ProvCam: A Camera Module with Self-Contained TCB for Producing Verifiable Videos. InProc. of MobiCom
2024
-
[68]
Florian Lugstein, Simon Baier, Gregor Bachinger, and Andreas Uhl. 2021. PRNU- based deepfake detection. InProc. of ACM Workshop on Information Hiding and Multimedia Security
2021
-
[69]
Yuxin Liu, Yoshimichi Nakatsuka, Ardalan Amiri Sani, Sharad Agarwal, and Gene Tsudik. 2022. Vronicle: verifiable provenance for videos from mobile devices. InProc. of MobiSys
2022
-
[70]
Santiago Lyon. 2023. Leica Launches World’s First Camera with Content Cre- dentials. https://contentauthenticity.org/blog/leica-launches-worlds-first- camera-with-content-credentials
2023
-
[71]
Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. 2020. Two-branch recurrent network for isolating deepfakes in videos. InEuropean Conference on Computer Vision
2020
-
[72]
Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. 2021. Generalizing face forgery detection with high-frequency features. InProc. of CVPR. 16317–16326
2021
-
[73]
Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff. 2017. On Detecting Adversarial Perturbations. InInternational Conference on Learning Representations
2017
-
[74]
Peter Michael, Zekun Hao, Serge Belongie, and Abe Davis. 2025. Noise-Coded Illumination for Forensic and Photometric Video Analysis.ACM Trans. on Graphics44, 5 (2025), 1–16
2025
-
[75]
Bill McCarthy. 2024. Video of Biden botching Ukraine history is a deepfake. https://factcheck.afp.com/doc.afp.com.34LY8TK
2024
-
[76]
Paarth Neekhara, Shehzeen Hussain, Xinqiao Zhang, Ke Huang, Julian McAuley, and Farinaz Koushanfar. 2022. FaceSigns: Semi-Fragile Neural Watermarks for Media Authentication and Countering Deepfakes. arXiv:2204.01960
2022 arXiv
-
[77]
Huy H Nguyen, Junichi Yamagishi, and Isao Echizen. 2019. Capsule-forensics: Using capsule networks to detect forged images and videos. InIEEE International Conference on Acoustics, Speech and Signal Processing
2019
-
[78]
Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020. Emotions don’t lie: An audio-visual deepfake detection method using affective cues. InACM International Conference on Multimedia
2020
-
[79]
Kenji Okuma, James J Little, and David G Lowe. 2004. Automatic Rectification of Long Image Sequences. InAsian Conference on Computer Vision
2004
-
[80]
Josephine Passananti, Stanley Wu, Shawn Shan, Haitao Zheng, and Ben Y Zhao. 2024. Disrupting style mimicry attacks on video imagery.arXiv preprint arXiv:2405.06865(2024)
2024 arXiv
-
[81]
Yuval Nirkin, Yosi Keller, and Tal Hassner. 2019. FSGAN: Subject agnostic face swapping and reenactment. InProc. of CVPR
2019
-
[82]
Amna Qureshi, David Megías, and Minoru Kuribayashi. 2021. Detecting deep- fake videos using digital watermarking. InAsia-Pacific Signal and Information Processing Association Annual Summit and Conference. 1786–1793
2021
-
[83]
Evani Radiya-Dixit, Sanghyun Hong, Nicholas Carlini, and Florian Tramèr
-
[84]
Hua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie, Lei Ma, Wei Feng, Yang Liu, and Jianjun Zhao. 2020. Deeprhythm: Exposing deepfakes with attentional visual heartbeat rhythms. InACM International Conference on Multimedia
2020
-
[85]
Aruna Sankaranarayanan, Matthew Groh, Rosalind Picard, and Andrew Lipp- man. 2021. The presidential deepfakes dataset. InCEUR Workshop Proceedings
2021
-
[86]
Is this my president speaking?
Irtaza Shahid and Nirupam Roy. 2023. " Is this my president speaking?" Tamper- proofing Speech in Live Recordings. InProc. of MobiSys
2023
-
[87]
Shawn Shan, Emily Wenger, Jiayun Zhang, Huiying Li, Haitao Zheng, and Ben Y Zhao. 2020. Fawkes: Protecting privacy against unauthorized deep learning models. In29th USENIX Security Symposium
2020
-
[88]
Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to Detect Manip- ulated Facial Images. InInternational Conference on Computer Vision
2019
-
[89]
Gaurav Sharma, Wencheng Wu, and Edul N Dalal. 2005. The CIEDE2000 color- difference formula: Implementation notes, supplementary test data, and mathe- matical observations.Color Research & Application30, 1 (2005), 21–30
2005
-
[90]
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. InAdvances in Neural Information Processing Systems
2019
-
[91]
Michael H. Siegel. 1965. Color Discrimination as a Function of Exposure Time. Journal of the Optical Society of America55, 5 (May 1965), 566–568
1965
-
[92]
Jiacheng Shang and Jie Wu. 2020. Protecting Real-time Video Chat against Fake Facial Videos Generated by Face Reenactment. InInternational Conference on Distributed Computing Systems
2020
-
[93]
Satoshi Suzuki. 1985. Topological structural analysis of digitized binary images by border following.Computer Vision, Graphics, and Image Processing30, 1 (1985), 32–46
1985
-
[94]
Danielle Szafir. 2017. Effects of Size and Shape on Perceived Color Differences. Journal of Vision17, 10 (2017), 1189
2017
-
[95]
Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. InInternational Conference on Machine Learning
2019
-
[96]
S Shyam Sundar, Maria D Molina, and Eugene Cho. 2021. Seeing Is Believing: Is Video Modality More Powerful in Spreading Fake News via Online Messaging Apps?Journal of Computer-Mediated Communication26, 6 (08 2021), 301–319
2021
-
[97]
Nick Vlku. 2002. The History of Television. https://www.cs.cornell.edu/~pjs54 /Teaching/AutomaticLifestyle-S02/Projects/Vlku/history.html
2002
-
[98]
Natalie Wade. 2023. Deepfake of Bella Hadid misrepresents her statements on Israel. AFP Fact Check
2023
-
[99]
Anran Wang, Zhuoran Li, Chunyi Peng, Guobin Shen, Gan Fang, and Bing Zeng
-
[100]
Hiroshi Unno and Kazutake Uehira. 2020. Lighting Technique for Attaching Invisible Information Onto Real Objects Using Temporally and Spatially Color- Intensity Modulated Light.IEEE Transactions on Industry Applications56, 6 (2020), 7202–7207
2020
-
[101]
Jiadong Wang, Xinyuan Qian, Malu Zhang, Robby T Tan, and Haizhou Li. 2023. Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert. InProc. of CVPR
2023
-
[102]
Qiongqiong Wang, Koji Okabe, Kong Aik Lee, Hitoshi Yamamoto, and Takafumi Koshinaka. 2018. Attention Mechanism in Speaker Recognition: What Does it Learn in Deep Speaker Embedding?. In2018 IEEE Spoken Language Technology Workshop. 1052–1059
2018
-
[103]
Run Wang, Felix Juefei-Xu, Meng Luo, Yang Liu, and Lina Wang. 2021. Faketag- ger: Robust safeguards against deepfake dissemination via provenance tracking. InACM International Conference on Multimedia
2021
-
[104]
InFrame++: Achieve Simultaneous Screen-Human Viewing and Hidden Screen-Camera Communication. InProc. of MobiSys
-
[105]
Bolun Wang, Yuanshun Yao, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao
-
[106]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Transac- tions on Image Processing13, 4 (2004), 600–612
2004
-
[107]
Wiles, A.S
O. Wiles, A.S. Koepke, and A. Zisserman. 2018. Self-supervised learning of a facial attribute embedding from video. InBritish Machine Vision Conference
2018
-
[108]
Grace Woo, Andy Lippman, and Ramesh Raskar. 2012. VRCodes: Unobtrusive and active visual codes for interaction by exploiting rolling shutter. InIEEE International Symposium on Mixed and Augmented Reality
2012
-
[109]
Xander Elliards. 2024. AI deepfake video of John Swinney on Sky News goes viral on Twitter. The National. CCS ’25, October 13–17, 2025, Taipei, Taiwan Hadleigh Schwartz, Xiaofeng Yan, Charles J. Carver, and Xia Zhou
2024
-
[110]
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. 2020. CNN-Generated Images Are Surprisingly Easy to Spot... for Now. InProc. of CVPR
2020
-
[111]
Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, and Houqiang Li
-
[112]
Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu. 2023. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. InAd- vances in Neural Information Processing Systems
2023
-
[113]
Zixiang Yang, Wei Tsang Ooi, and Qibin Sun. 2004. Hierarchical, non-uniform locality sensitive hashing and its application to video identification. InIEEE International Conference on Multimedia and Expo, Vol. 1. 743–746
2004
-
[114]
Kai Zhang, Chenshu Wu, Chaofan Yang, Yi Zhao, Kehong Huang, Chunyi Peng, Yunhao Liu, and Zheng Yang. 2018. ChromaCode: A Fully Imperceptible Screen- Camera Communication System. InProc. of MobiCom
2018
-
[115]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[116]
Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. 2023. SadTalker: Learning Realistic 3D Motion Coeffi- cients for Stylized Audio-Driven Single Image Talking Face Animation. InProc. of CVPR
2023
-
[117]
Yuting Xu, Jian Liang, Gengyun Jia, Ziming Yang, Yanhao Zhang, and Ran He
-
[118]
InInternational Conference on Computer Vision
TALL: Thumbnail Layout for Deepfake Video Detection. InInternational Conference on Computer Vision
-
[119]
Zhiyuan Yan, Yong Zhang, Yanbo Fan, and Baoyuan Wu. 2023. UCF: Uncov- ering Common Features for Generalizable Deepfake Detection.arXiv preprint arXiv:2304.13949(2023)
2023 arXiv
-
[124]
The unreasonable effectiveness of deep features as a perceptual metric. In Proc. of CVPR
-
[126]
Xu Zhang, Svebor Karaman, and Shih-Fu Chang. 2019. Detecting and Simulating Artifacts in GAN Fake Images. InIEEE International Workshop on Information Forensics and Security
2019
-
[127]
Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. 2021. Multi-attentional Deepfake Detection. InProc. of CVPR
2021
-
[128]
Yes" to Q1 when embedding was occurring, and FPR is the rate at which they incorrectly responded
Yipin Zhou and Ser-Nam Lim. 2021. Joint audio-visual deepfake detection. In Proc. of CVPR. Appendix A.1 Perceptibility Evaluation Details Here, we detail the setup of our perceptibility user study and LPIPS evaluation and further discuss the studies’ results. User studyWe invi...
2021
-
[2015]
Real-Time Screen-Camera Communication Behind Any Scene. InProc. of MobiSys
-
[2018]
In27th USENIX Security Symposium
With great training comes great vulnerability: Practical attacks against transfer learning. In27th USENIX Security Symposium
-
[2019]
VoxCeleb: Large-scale speaker verification in the wild.Computer Science and Language(2019)
2019
-
[2021]
InInternational Conference on Learning Representations
Data poisoning won’t save you from facial recognition. InInternational Conference on Learning Representations
-
[2023]
AltFreezing for More General Video Face Forgery Detection. InProc. of CVPR
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.