Pith. sign in

Paper Citation Record · LEDGER

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

As of 8 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.15233.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15233 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:27:00.996471Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T15:08:25.309094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.040118Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact2
  • verified fuzzy53
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb4cd9f1-cff9-435c-89b3-9c71e68d3766 · outbound

This paper cites Mesonet: a compact facial video forgery detection network.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Mesonet: a compact facial video forgery detection network

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.335259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.074983Z digest=sha256:bc74119d52ee2c37856235c2e821fe68e7baa28bc3774bbd1cde2acb9300c6ea

Observation c58e74d7-cdc7-43a4-8e66-492732a5fe8c · outbound

This paper cites A review of modern audio deepfake detection methods: challenges and future directions.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A review of modern audio deepfake detection methods: challenges and future directions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.316636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.150530Z digest=sha256:b73c767f1583fd69d21b9dc0746c3bcc0266179504a0cf1309c06cbfb8e58a00

Observation 1f187ee3-29bd-4605-9309-632285bf018e · outbound

This paper cites Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.292173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.322384Z digest=sha256:1aea32d381330b52f1aa090ae188c123d57831ca034c8698d640b078a5889ed5

Observation fe38a54f-d59a-4ce4-8404-a8aa568b0c69 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.433137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.433137Z digest=sha256:90e7bb37c6394db8cbcbe623947046f1461e1e7a55d1c0fa712de57c9aca5e1e

Observation d6ce7981-1a59-44ce-bea4-20bcf51e17cb · outbound

This paper cites an unresolved cited work.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:27:04.263567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.539628Z digest=sha256:01adae99175bb9cd8eb839a8859b767bc08990facbe59cc5e5f1012e32d0aee1

Observation 933e19fd-3ba1-4275-9383-5013726f029d · outbound

This paper cites Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.246938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.649874Z digest=sha256:9b0bb3aa6bea79400c2a451eb51367bc7fe113d9a1baa158818195246d2193fa

Observation 32a37ffa-7e7c-4397-b206-ab43e5cbd069 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A Simple Framework for Contrastive Learning of Visual Representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.756360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.756360Z digest=sha256:9d3d2dd9e0f94c27954bac1629de6f2913d6e914be3a02b7061b9a427a2e42e8

Observation 50ac14fd-fba5-4306-a83f-f50b062b16be · outbound

This paper cites Exploring simple siamese representation learning.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploring simple siamese representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.230249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.856339Z digest=sha256:ac2f7bd07d35d1f3dc9ad9839030d7ba65a4a103e16914bee6f550f75e247d2b

Observation a8f0832a-cbc0-436c-8860-320875445262 · outbound

This paper cites Sophia Koepke, Ying Shan, and Zeynep Akata.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Sophia Koepke, Ying Shan, and Zeynep Akata

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.213684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:54.977105Z digest=sha256:42644effa8131b50b1a3b3e9360cafe8f0ce40680521bbdbed714f657eaf22fc

Observation 64487f5f-8e60-4cc6-bc87-df0abcfee9a4 · outbound

This paper cites V oice-face homogeneity tells deepfake.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation V oice-face homogeneity tells deepfake

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.197708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.048303Z digest=sha256:f793b25e8681387fe630c9f43bf64639aab06e5cfde38769175e5a8162fa1d4c

Observation 5769a50b-9188-4a15-8163-a547698489f2 · outbound

This paper cites Can We Leave Deepfake Data Behind in Training Deepfake Detector?.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Can We Leave Deepfake Data Behind in Training Deepfake Detector?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:55.184286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:55.184286Z digest=sha256:c4c837b4abd61d3df968f2022372816625ce627e993d44188e97a5aaa9a38d0b

Observation 129e1880-4817-4737-ae70-a38b7652923e · outbound

This paper cites Xception: Deep learning with depthwise separable convolutions.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Xception: Deep learning with depthwise separable convolutions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.180876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.302595Z digest=sha256:f7676f97dcdfc36a5dc286594fa2b99487a6ee3824be0ac3e9d5cd3e2563463d

Observation 8284eb72-ad9e-472d-a7aa-63f5969c98d0 · outbound

This paper cites Not made for each other- audio-visual dissonance-based deepfake detection and localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Not made for each other- audio-visual dissonance-based deepfake detection and localization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.155779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.405611Z digest=sha256:35737e90a9c9aef6d8b1c0252cfa63cc610d017af1a3c2f5f98015e9af7992b1

Observation 56d9c05b-5efd-4049-9806-edda9d0c0262 · outbound

This paper cites Audio-visual person-of-interest deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Audio-visual person-of-interest deepfake detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.138476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.485300Z digest=sha256:63ac5564a9c963e872dd5d2be4d6cbcccadd8aafa4e75310aa2065e191888bcd

Observation 5eb7d261-3663-4989-872a-9a2da6120e54 · outbound

This paper cites Dong, Jin Wang, Renhe Ji, Jiajun Liang, Haoqiang Fan, and Zheng Ge.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Dong, Jin Wang, Renhe Ji, Jiajun Liang, Haoqiang Fan, and Zheng Ge

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.119110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.597172Z digest=sha256:07f02b122a95157f1b7eb3c5ee160cff5fdd23f6f06799277a6dd543cc861bc6

Observation 5b2b32de-9a52-488c-848a-68dffb596ebb · outbound

This paper cites www.github.com/MarekKowalski/FaceSwap Accessed 2021-04-24.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation www.github.com/MarekKowalski/FaceSwap Accessed 2021-04-24

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.099701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.697949Z digest=sha256:e2f4a04409c8a5d4ec21a08c3e93cb2cd5600cda85b7efea3020884fc37ce510

Observation 85c18813-1f05-4efb-ba7f-1188277e5ec9 · outbound

This paper cites Self-supervised video forensics by audio-visual anomaly detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Self-supervised video forensics by audio-visual anomaly detection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.079789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.778876Z digest=sha256:02c01cbf3f157a7e2af34a009d10388b9e996c286a8f05937e134d406a4da3d1

Observation 2a551dd0-faf3-411d-88b2-5fbe681f7e98 · outbound

This paper cites Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:27:01.797723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.856641Z digest=sha256:7f5943c1ebdca70a7911fa6689b5b1ec23159cc88407898b69213b694a4de929

Observation 910b8953-1a46-46c7-80d4-6a8afc3b5515 · outbound

This paper cites Imagebind one embedding space to bind them all.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Imagebind one embedding space to bind them all

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.058641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:55.927230Z digest=sha256:20baf89b43a441751c73e4c3039333536bde769fc9a13eff8156f4e2b393e96e

Observation 1d42e731-48c8-4c7c-adb2-ceddac4857b8 · outbound

This paper cites Leveraging real talking faces via self-supervision for robust forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Leveraging real talking faces via self-supervision for robust forgery detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.039846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.007650Z digest=sha256:c9e5b59902b137fc71743674366c0b42868d00cfd5e70a9328b2e28082b9c06f

Observation 242ef530-74f6-4b75-955b-c98b0fde5954 · outbound

This paper cites Lips don’t lie: A generalisable and robust approach to face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lips don’t lie: A generalisable and robust approach to face forgery detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:04.016350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.106380Z digest=sha256:d0dc5d1294ac050215d36ab7a09245c0302233d1753c55525a4595b30aa699b1

Observation 459ebd5f-63b3-4792-9178-d88286def08e · outbound

This paper cites Masked autoencoders are scalable vision learners.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Masked autoencoders are scalable vision learners

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.992759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.188215Z digest=sha256:c0435d2e0da26c81b58807e266337ab7f911882f9b3fba34f40b871b6023a299

Observation c7900e5d-7aea-49b9-9312-8e74246eb0b9 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.972701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.329894Z digest=sha256:bec98204625054bb27844272dd0c90b1981c09cd71617115f8b8af1700fbefcb

Observation 105c5295-9d6b-4502-bb53-817ee1220e08 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lora: Low-rank adaptation of large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.400893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.400893Z digest=sha256:3c9a9c6eec4eb43f8fd8cf8c2a4ac52e955f41e9a3ddccc80dc8c6dac51680d1

Observation 02a2a4a0-ad58-4f6e-846c-7aec0d9ea89c · outbound

This paper cites Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio-visual deepfakes detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio-visual deepfakes detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.944903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.513786Z digest=sha256:62d4adfec0cfad3206af8b1cf6018e555ffb91a382532a6f1236f13304f75923

Observation 4c94acf6-c13d-49e4-b28c-c07b57924b0c · outbound

This paper cites Information theory and statistical mechanics.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Information theory and statistical mechanics

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.596373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.596373Z digest=sha256:644de98e944a1d0a5eda11eb8ef1b775d803dc4d59d16261d614e5fd96bd5e87

Observation ceb33f51-45cd-4bcf-8c1c-5be15e06db80 · outbound

This paper cites an unresolved cited work.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:27:03.913566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.679484Z digest=sha256:fbae1d412c8d71106d8c7cba3b5ff96c33a35e3f3a2ed983e22a4d96379ddf78

Observation 6a14a426-5563-4cd0-8978-90fd0c27ee70 · outbound

This paper cites FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.756989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.756989Z digest=sha256:06fce1e877c366cef09ff100c07035b43d2b5aa636dd984101849ebe5c05f11e

Observation aed45079-acd6-41b8-a7b5-9021e513db61 · outbound

This paper cites Fast face-swap using convolu- tional neural networks.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Fast face-swap using convolu- tional neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.893130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.853865Z digest=sha256:5329568b71ae52cfa3f6a7419b01ebc4c3dcefb7651ee3dca8170e4d15c3ae02

Observation 9a0605a9-493f-4e9d-8144-64663846db63 · outbound

This paper cites Face x-ray for more general face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Face x-ray for more general face forgery detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.874270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:56.962771Z digest=sha256:0b7a762a96291372717fc7a4b8e3adb4bdab16c46ba37173f26615f99a023af8

Observation 15b3d7be-6001-44bf-b67b-a03298ab009d · outbound

This paper cites A Survey on Speech Deepfake Detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A Survey on Speech Deepfake Detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.093958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.093958Z digest=sha256:2ee5e630e9882964d0504aef8f8b3f25b8852fed75081b5353c786add915207c

Observation a63c7869-feb6-493e-b625-b0cb81abfd36 · outbound

This paper cites Celeb-df: A large-scale challenging dataset for deepfake forensics.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Celeb-df: A large-scale challenging dataset for deepfake forensics

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.853254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.170497Z digest=sha256:4c23383daecea823b3553e0f7bb05ad7194b2e09bef7b9e140fad82094540fd1

Observation 22291b3e-f02a-49b8-9700-1ad0cdd5a626 · outbound

This paper cites Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.836537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.249489Z digest=sha256:9e56cc351de2f4eb31d879177c9d21bca16c78432cad4c7fb0ac98acd11eb7f2

Observation b27d098b-b827-4513-b161-e801418e734b · outbound

This paper cites Exploiting visual artifacts to expose deepfakes and face manipulations.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploiting visual artifacts to expose deepfakes and face manipulations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.817396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.377989Z digest=sha256:e3abd981e2f60d3798d6f294a41179d0987b4b18c51f346698bfda71ef1f734b

Observation ef0db15d-c892-49c4-9f39-138ae7df8e3c · outbound

This paper cites Emo- tions don’t lie: An audio-visual deepfake detection method using affective cues.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Emo- tions don’t lie: An audio-visual deepfake detection method using affective cues

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.798323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.468124Z digest=sha256:4ea13ee4421dbc41d4841f8ff8d994965e41ed0efbd5d306e2b81d258a91c301

Observation 9bc000c9-199a-4dc6-81f0-9058d7085044 · outbound

This paper cites Does audio deepfake detection generalize? arXiv preprint arXiv:2203.16263, 2022.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Does audio deepfake detection generalize? arXiv preprint arXiv:2203.16263, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.564874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.564874Z digest=sha256:9c21fecdd21c9a166d80b059180a3750146e3e081a8704afec8e5fddf28fb2bb

Observation 1761e696-aaa8-4fd5-9b5c-63c1bf77c4c1 · outbound

This paper cites Nguyen, Junichi Yamagishi, and Isao Echizen.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Nguyen, Junichi Yamagishi, and Isao Echizen

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.780765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.664207Z digest=sha256:387cb499e5cf15b78548116cc94f8955384bd0de7d4ca91951c11a8b3db0b43d

Observation f62a7ef7-de00-4d25-b642-f086ef6c0a94 · outbound

This paper cites Towards universal fake image detectors that generalize across generative models.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Towards universal fake image detectors that generalize across generative models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.759973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.758259Z digest=sha256:30d52e77315b04cfad73ed92eb0458bbfe3231903b1edc42b884c4168febe77a

Observation b33bf91a-4459-486a-a57f-7206d11da1d0 · outbound

This paper cites Avff: Audio-visual feature fusion for video deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Avff: Audio-visual feature fusion for video deepfake detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.738263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.831456Z digest=sha256:746edfa6ef783abc72087a2643d5f41cf6b6caaaf04c876eeb718b8d6043a31c

Observation 2e33d572-82a6-49c0-a661-9a5ed21294cf · outbound

This paper cites Gpt-4: A large multimodal model.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Gpt-4: A large multimodal model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.709277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:57.907416Z digest=sha256:21bd8286c3e500948dce91a0ffda874756563502a2ffa2173ae835c36c568584

Observation a961ec6c-c6cd-4f8e-8810-3f959b1956e7 · outbound

This paper cites Deepfake generation and detection: A benchmark and survey.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Deepfake generation and detection: A benchmark and survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:57.992374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:57.992374Z digest=sha256:02f4affd7de692801cb173773d7676a5110643ee1d80470fd0e5fb12e76278fb

Observation 4939c70d-d9cf-4849-abff-6806947f0f82 · outbound

This paper cites Namboodiri, and C.V.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Namboodiri, and C.V

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.690146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.114333Z digest=sha256:af74fc227d85df0536f8b3dd945d880baf41e3e9b659b950afe9c59fe4c9c360

Observation ce7db574-1fc5-4d78-8ee6-13722ead2071 · outbound

This paper cites Audio-visual deep neural network for robust person verification.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Audio-visual deep neural network for robust person verification

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.667273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.192243Z digest=sha256:58b4d807d0684d82cddc2ade704ff4e35a19d28e02e36f232a7c4f1012fab862

Observation 998addff-3393-417d-b781-5e3d60ff255a · outbound

This paper cites Learning transferable visual models from natural language supervision.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Learning transferable visual models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.644870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.255476Z digest=sha256:d38c2d736403ace2e6109a335bdb623425301d84aa50d76328279f1b0cccde21

Observation ed3b185b-2c25-44f5-8a5d-80711d7b9155 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.327095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.327095Z digest=sha256:139bf1b799b9806e1ad38e3e8affa70955bd0b5e56bccf0b2424ec34b2ae9e4f

Observation 27b2ac9b-a357-416b-8236-577cc5b91d59 · outbound

This paper cites Faceforensics++: Learning to detect manipulated facial images.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Faceforensics++: Learning to detect manipulated facial images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.606717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.398378Z digest=sha256:8bc40587e3fada119412276be9cc89e607f9a786af7b66a5188bd233d66637aa

Observation 4221d3cd-8174-470c-804c-b320f7db35c7 · outbound

This paper cites A comprehensive overview of deepfake: Generation, detection, datasets, and opportunities.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation A comprehensive overview of deepfake: Generation, detection, datasets, and opportunities

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.586152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.461992Z digest=sha256:6d6750f27efef0c49f30af5615c89d1e7abaaf91f2261b6972837d675907d70f

Observation 5bdf06cd-1658-4f2d-a4ef-2585d7efe552 · outbound

This paper cites Detecting deepfakes with self-blended images.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Detecting deepfakes with self-blended images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.518035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.518035Z digest=sha256:480c1a94b1c6d272157d7f63fdfb32430ce6175a768dc2814322cf82339a9b13

Observation 1a3c8a52-071e-4a1b-9382-492116dc065e · outbound

This paper cites Representative forgery mining for fake face detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Representative forgery mining for fake face detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.548967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.570297Z digest=sha256:97ee90cd02ac9f11646545e05484cb34fc3c68be439e36ad339eeb814c47f029

Observation 97f193b6-0890-403e-8b58-e6db7fecf709 · outbound

This paper cites Exploring Depth Information for Detecting Manipulated Face Videos.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploring Depth Information for Detecting Manipulated Face Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:27:01.335494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.647254Z digest=sha256:274a9d1d606e6542ffa80268779fdc435db931230c75e85700ae34dc93aa3729

Observation ae051bda-29c8-4dc1-a08a-1e6535fcd5a8 · outbound

This paper cites Tan, and Haizhou Li.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Tan, and Haizhou Li

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.522275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.706727Z digest=sha256:5e912f8aca0e4bf42f4935de00356dd68836418025ca53569b4a401ca9868623

Observation 0ad396e1-9008-44d0-909e-bce6a7cfc442 · outbound

This paper cites Deep spatial gradient and temporal depth learning for face anti-spoofing.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Deep spatial gradient and temporal depth learning for face anti-spoofing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.497856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.768804Z digest=sha256:67944ee7cf702c20e33ca7cbdd1089325622dbf99fb7ad51fb9f06bf5cc4363e

Observation 8654a873-3fb6-43db-b586-f4a3d6f549a9 · outbound

This paper cites Altfreezing for more general video face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Altfreezing for more general video face forgery detection

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.470186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.844117Z digest=sha256:ebd1a6a0b4c267da5077b995c95bf664773277a8314019cc794026d4ddc180ad

Observation febc692d-c351-4c3a-85a6-300c68ce4ca3 · outbound

This paper cites Deepfake Video Detection Using Convolutional Vision Transformer.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Deepfake Video Detection Using Convolutional Vision Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:58.916792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:58.916792Z digest=sha256:24ff8a4fe7ada81d5ffe1ce47f326a307d61a4bdcf461f788174aa314f4958c2

Observation ddcbe21a-8038-433e-9d37-8d3b1f313f5f · outbound

This paper cites Binaural audio-visual localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Binaural audio-visual localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.445286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:58.988970Z digest=sha256:0dcc6337e050176ff3c434f45d14385c36a3e7831fb41d0635a22648f44d6535

Observation b4697dcc-b1a5-47a8-ad45-b2d8fffe439d · outbound

This paper cites Identity- driven multimedia forgery detection via reference assistance.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Identity- driven multimedia forgery detection via reference assistance

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.421904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.078297Z digest=sha256:bf0ed5e606c66a2fda830931848a8803a77ba5bc7ac64560043c2da5d30ad8db

Observation 15ce0a97-bc90-4e7e-ac37-97abf628d8bb · outbound

This paper cites Tall: Thumbnail layout for deepfake video detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Tall: Thumbnail layout for deepfake video detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.390976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.161502Z digest=sha256:b7602b40de9212a7f9bf82b5c5ebd5d6fcf08646537fb6a22f32b5f8a0a99f82

Observation c6811e39-49e0-4c8c-af3b-f003898121cd · outbound

This paper cites Transcending forgery specificity with latent space augmentation for generalizable deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Transcending forgery specificity with latent space augmentation for generalizable deepfake detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.362476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.229390Z digest=sha256:09b75eaf1866989a949cfcbd1b2ef3aebadf76ea2de8aae9328143350f891e43

Observation 971ad1d4-6480-44f9-a694-e2f7f915d7e2 · outbound

This paper cites DF40: Toward Next-Generation Deepfake Detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation DF40: Toward Next-Generation Deepfake Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.305225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.305225Z digest=sha256:23d0287b2033dfef49c5197f657258e0cfe3436470766f5da592d41c61f4c71c

Observation 1d877f55-ac55-496c-9d20-c55c0911f808 · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.364495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.364495Z digest=sha256:0a83edc8c7ec73b919c3c91de07e5494d8d9d7d363ab1df117a71bead77a2842

Observation d25b769b-48ca-4156-bb5f-d47bcb261003 · outbound

This paper cites Ucf: Uncovering common features for generalizable deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Ucf: Uncovering common features for generalizable deepfake detection

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.337860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.442188Z digest=sha256:7168b008877ab62ba908129b2a7379254a20eced8783b104dcc2d8d9cc89c03c

Observation 9c056045-81dc-401a-a5b1-e90a2e2f039f · outbound

This paper cites Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.498474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.498474Z digest=sha256:4ea2a23cde0b7210fb9078817eb4f774880cd193ef807146a8c27e2f16382461

Observation ad1cfbbd-6071-4c67-919d-5c751b1c3a31 · outbound

This paper cites Avoid-df: Audio-visual joint learning for detecting deepfake.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Avoid-df: Audio-visual joint learning for detecting deepfake

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:03.032256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.566845Z digest=sha256:15deb644b47cba19a6d629214c9f5e806bdb0b88e3f0550c204a30ef296612b9

Observation be58bad2-38dc-4944-941b-4c71e3c5c80e · outbound

This paper cites Exposing deep fakes using inconsistent head poses.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exposing deep fakes using inconsistent head poses

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.912242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.676740Z digest=sha256:e59a2d247f38a633808ef59c0ff31fa56ad96adbad5e8c4983cc00ac2731c289

Observation 4939e96a-4e65-4f9e-b493-fec38e132990 · outbound

This paper cites Audio Deepfake Detection: A Survey.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Audio Deepfake Detection: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.790961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.790961Z digest=sha256:07964c86e0141e3b91436fd772088dff41cd12a5ac15c67f352871213b1118ab

Observation f57e4aa7-b32c-4034-9cac-eafa94f7c76a · outbound

This paper cites Learning natural consistency representation for face forgery video detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Learning natural consistency representation for face forgery video detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.835435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:26:59.977050Z digest=sha256:b095df5f4c78d45a28b296fc5e805d21a6dbca800abb7cac3ccbb872eac69f5d

Observation 0a1ded16-7871-485e-a9f1-31875ef21b89 · outbound

This paper cites Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:00.152871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:00.152871Z digest=sha256:dfc71b496a6a02c96f4cc6e7e8debfb57281cb71d0fa9f850bc63d452a89e203

Observation 7b594adb-5bb3-4472-b87f-5e8c0bae3f49 · outbound

This paper cites Multi-attentional deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Multi-attentional deepfake detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.717608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.352741Z digest=sha256:69bace41a2d7e5a8af2695e87ef9be8b85fb9a2536e6c4536bd85cf8174a23af

Observation e5f35ecf-5e29-4f4a-aa9d-05a896ab8679 · outbound

This paper cites Attention-based spatial-temporal multi-scale network for face anti-spoofing.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Attention-based spatial-temporal multi-scale network for face anti-spoofing

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.577494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.559647Z digest=sha256:f32593ebe73ece40a570637a979bfa5ed1f2984d771777a8831386351ca4fda1

Observation 98e6c56d-f524-4cca-8c56-5edb2bfa5436 · outbound

This paper cites Exploring temporal coherence for more general video face forgery detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Exploring temporal coherence for more general video face forgery detection

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.445239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.736070Z digest=sha256:05a333d23e07ddaf572d1955c5384cfaf383940144cf95b504c1713bd1a26772

Observation ddf7a907-6417-4d79-bc02-88e3e060afbe · outbound

This paper cites Learning deep features for discriminative localization.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Learning deep features for discriminative localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.349469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.875898Z digest=sha256:d04bd3031a880f80204a9090b4708e3049f15e6d1b2613f74f776de9250c52fa

Observation 80b7e914-261f-4337-9107-8595853eb3e8 · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Makelttalk: speaker-aware talking-head animation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.294628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.979116Z digest=sha256:27fe008d4b31c102b981e901695232b812c65b8ef5dc677b64dacae230cdc6fc

Observation 80b73bbd-895d-4d40-88b8-38d16dff6ea2 · outbound

This paper cites Joint audio-visual deepfake detection.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Joint audio-visual deepfake detection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.134020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.986411Z digest=sha256:5ca3eef867e73d9c335b9066b793c2f18440dc29f0182cd36bf60cc4410698a9

Observation ec5cef83-b4fd-464f-a38d-b37ac7d8645e · outbound

This paper cites Lan- guagebind: Extending video-language pretraining to n-modality by language-based semantic alignment.

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation Lan- guagebind: Extending video-language pretraining to n-modality by language-based semantic alignment

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:27:02.020529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:27:00.996471Z digest=sha256:24ed8ffa3365056a3769b0622fa3b3eb40c8cf637ee2a8bf4b817a8703e658f6

Pith citing papers

Observation 4929d710-214b-4f3a-b5a0-17abcb6b57a5 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.042025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:aace61ffd386cd6399990fd3dc6c96b57b01a93e15b4c60e502385c0f140c5cb