Pith. sign in

Paper Citation Record · LEDGER

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

As of 23 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 11 inbound Pith citation observations for arXiv:2411.10193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10193 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:57:43.285908Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:30.005614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:07:02.288478Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact1
  • verified fuzzy67
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba01d280-6c6d-4288-8824-d622934a0dc5 · outbound

This paper cites Mesonet: a compact facial video forgery detection network.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Mesonet: a compact facial video forgery detection network

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.926251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.926251Z digest=sha256:f5c41abd97bfe2ef148e85d33c73ae0640e33cbaeb0f9c38f63c04adba573ed1

Observation f1151922-57ea-4923-b5e7-fa222dbbfae5 · outbound

This paper cites Deep audio-visual speech recognition.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deep audio-visual speech recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.931568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.931568Z digest=sha256:1e8779aca86b18d725374fb7deeac0bf4ded4e760bf85b96267c1d6a77e78cf7

Observation 828b5171-0cd1-4e72-aec0-54453fecb232 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization LRS3-TED: a large-scale dataset for visual speech recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.935850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.935850Z digest=sha256:54917930d4702a8d8d44ce1cb64cf59c751004b95579d2703d8848f6974b61eb

Observation c509f8c8-f2bd-4b6f-85c7-63c96cfbf0fd · outbound

This paper cites Transitions in neural oscillations reflect pre- diction errors generated in audiovisual speech.Na- ture neuroscience, 14(6):797–801, 2011.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Transitions in neural oscillations reflect pre- diction errors generated in audiovisual speech.Na- ture neuroscience, 14(6):797–801, 2011

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.940463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.940463Z digest=sha256:f0d245e561d170696f659930d4104a1d1e2a058368c74c0d063868995503bb19

Observation ec047c76-d001-42ee-b4db-eb30974fa3e7 · outbound

This paper cites Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.944640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.944640Z digest=sha256:d6664aad74f23544354934ac84cdc66faac807424efab3f3e9ce2649f7edf6a2

Observation ccd8201c-efa2-444d-bfad-649975a7270d · outbound

This paper cites Phoneme-to- viseme mappings: the good, the bad, and the ugly.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Phoneme-to- viseme mappings: the good, the bad, and the ugly

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.948696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.948696Z digest=sha256:3849777242b3ecc8bb4683039fa371c9a373dec913ce08fdaee410bf2adc846c

Observation 1477da99-97c7-4dfb-9fb4-a43a6fd92262 · outbound

This paper cites Lost in trans- lation: Lip-sync deepfake detection from audio- video mismatch.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lost in trans- lation: Lip-sync deepfake detection from audio- video mismatch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.952997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.952997Z digest=sha256:80922765e338b9d530f063d611f6cc7cfcff20549030830ebbfe51ef0c75578d

Observation e1221b4f-f42b-4d64-b23b-02301d3d5f55 · outbound

This paper cites Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.957178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.957178Z digest=sha256:a7178ec5fd904058fa2560a44a4f24014ccd6a45b38d65bd02a8a8a61289f555

Observation e696f7b7-445d-47e5-9cb6-678e484b7745 · outbound

This paper cites Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.961279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.961279Z digest=sha256:0c40443cf3deacdbaa1258ed85f8265345a2a68f1de57581704ca8750e953b0b

Observation 821fc75e-e240-4207-962b-3a2555be233c · outbound

This paper cites Do you really mean that? con- tent driven audio-visual deepfake dataset and mul- timodal method for temporal forgery localiza- tion.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Do you really mean that? con- tent driven audio-visual deepfake dataset and mul- timodal method for temporal forgery localiza- tion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.965129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.965129Z digest=sha256:5d12a04a2d68ed25e488059fee87132d24fab70b2a82e4c85fdd843fc7eddc00

Observation 4d4cbacb-e155-40f2-ad9e-501ea836ce7b · outbound

This paper cites End- to-end reconstruction-classification learning for face forgery detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization End- to-end reconstruction-classification learning for face forgery detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.968775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.968775Z digest=sha256:dcadaed9d40f40e6402cf221bf96a59632a1c4ce8d52082cb4f879682c587342

Observation 9dd11b83-9da6-41c9-b8ea-f61fc5d1582d · outbound

This paper cites Chandrasekaran and A.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Chandrasekaran and A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.972427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.972427Z digest=sha256:59b77affad631bdbf46b8b165a661eb17816531f8c068f6b1d83285a80c39b7b

Observation ea3579cf-100f-4f19-b476-b0e23e09caba · outbound

This paper cites What you see depends on what you hear: Temporal averaging and crossmodal in- tegration.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization What you see depends on what you hear: Temporal averaging and crossmodal in- tegration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.976224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.976224Z digest=sha256:412227021080cb2f6034f4d8a49bdc11ae42c720f80a026a5daa11d895a9836e

Observation 7fb2ee97-c599-4f31-bf5d-81442ed91bad · outbound

This paper cites Voice-face ho- mogeneity tells deepfake.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Voice-face ho- mogeneity tells deepfake

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.980020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.980020Z digest=sha256:de434ba21642871afb6420953dd6a81d9f1346974b435e61305e5941de7bd580

Observation c18a556a-c7e7-4820-a2b9-2aa1177a9adb · outbound

This paper cites Not made for each other-audio-visual dissonance-based deepfake de- tection and localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Not made for each other-audio-visual dissonance-based deepfake de- tection and localization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:42.983840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:42.983840Z digest=sha256:de9b6e8a361f1fbf4e07bad7bf3ef60718eaadf944b316b8bce4ced29d81422a

Observation f699a059-1b0d-4958-929d-2cb7bd7f8bdc · outbound

This paper cites Vox- Celeb2: Deep speaker recognition.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Vox- Celeb2: Deep speaker recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.492884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:42.987741Z digest=sha256:19fda4df36a52c30220a8793f4ec5e1b22fa57b6678ba646229d6bda9a91e238

Observation a1af069b-7321-490b-95e3-84634ff16a2c · outbound

This paper cites Lip reading in the wild.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lip reading in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.482034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:42.991659Z digest=sha256:6d6ef055189bc5f6504fd6b3992ca2da912727b27f392e65146de3ed7732316a

Observation 6c21c2bd-d294-4363-8173-59661ba0919c · outbound

This paper cites Combin- ing efficientnet and vision transformers for video deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Combin- ing efficientnet and vision transformers for video deepfake detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.469704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:42.995406Z digest=sha256:d188fef5891bef41f0c4b65dc6d49537e02320aca0bc6ae948563c43c3b4f810

Observation 69f40b8d-c460-40c3-ada1-7031d3d862b7 · outbound

This paper cites Audio-visual person-of-interest deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Audio-visual person-of-interest deepfake detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.457998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:42.998582Z digest=sha256:accc280a45c2689d507af3fc612a702bb0919a1b413d21f24d17b8be675e88f8

Observation 0f0cdda5-73a8-4213-aa38-9ee2ff227120 · outbound

This paper cites Imagenet: A large-scale hi- erarchical image database.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Imagenet: A large-scale hi- erarchical image database

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.447053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.001811Z digest=sha256:ac6a3b0add821b4a1defed8a717dd14f4a4ea132875be22ac431ab144d93af94

Observation 00b02011-e2ab-4a8f-9e73-b72be070ad96 · outbound

This paper cites The DeepFake Detection Challenge (DFDC) Dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization The DeepFake Detection Challenge (DFDC) Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.005309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.005309Z digest=sha256:a0cd482ded6c0fc8bed50f4c3c7da352c319580ba83db7d9a7ee62a682bd0497

Observation 22db688a-6f2d-411c-97be-e116aad56a93 · outbound

This paper cites Taming transformers for high-resolution im- age synthesis.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Taming transformers for high-resolution im- age synthesis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.434065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.008860Z digest=sha256:46e58468ac51878873c59da17a2039e64410ff69c1282fba91cc0ad6aaea2e5c

Observation baed580b-9dc3-48d1-a5c5-39454dc8f02a · outbound

This paper cites Self-supervised video forensics by audio-visual anomaly detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Self-supervised video forensics by audio-visual anomaly detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.419991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.012814Z digest=sha256:9db48c4a85c6d7d59670a6d278a26580e55de6b9c2aadff51b9ee2e9a4de8c58

Observation 9c06d2ac-5a8e-43b1-9045-1e8eb5568f08 · outbound

This paper cites Fast r-cnn.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Fast r-cnn

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.406731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.016582Z digest=sha256:0ba7143b538f96db12a42da198b42c393709c703f38ce30a49ba37b0ee5f03bd

Observation dab7ee5a-300a-458e-88a4-3dee2f931577 · outbound

This paper cites Rational decisions.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Rational decisions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.392238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.020842Z digest=sha256:672965908e63282f6b18b04fa3cd05459f4327f264dd831cf6ff3eccf3b67c4f

Observation 1d1e77ac-7331-4cb8-96bf-a9322fe940b6 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Bootstrap your own latent-a new approach to self-supervised learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.379577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.025215Z digest=sha256:d5edb00e94582b5bd223f179ab462daa02559fc2079bac39e07f24cf93286cbd

Observation e0f1f88b-835e-48b7-92fa-4c93a9e0f5d9 · outbound

This paper cites Deepfake video detection using audio- visual consistency.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deepfake video detection using audio- visual consistency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.368609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.029858Z digest=sha256:728445e78a612cea2a917d101f4c84ac3c73d625362929b842189c1516c7261d

Observation 5ec760ee-13e0-4616-b656-f919a33436ee · outbound

This paper cites Delving into the local: Dynamic inconsistency learning for deep- fake video detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Delving into the local: Dynamic inconsistency learning for deep- fake video detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.357235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.033876Z digest=sha256:217a34a66b529d9f6e710162056308811e1961f48016693edc089406b7a4dbef

Observation 07767f8c-4def-4733-b97b-b8f0562a3030 · outbound

This paper cites Leveraging real talking faces via self-supervision for robust forgery detec- tion.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Leveraging real talking faces via self-supervision for robust forgery detec- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.347064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.037960Z digest=sha256:4d1a1acf8d3cd4bfc05ee6469c372670dd819d3f24005a4e659d7d02159d0636

Observation 3caaede2-915e-46f2-a7fb-292e27946345 · outbound

This paper cites Lips don’t lie: A generalisable and robust approach to face forgery detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lips don’t lie: A generalisable and robust approach to face forgery detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.333711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.041863Z digest=sha256:1b7cf541fea406932365df286e0b45cb525dfa951239879404aa2ed95ffcb2bb

Observation 23ead194-4691-43ea-8cea-a87b0e024c27 · outbound

This paper cites De- tection of fake images via the ensemble of deep representations from multi color spaces.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization De- tection of fake images via the ensemble of deep representations from multi color spaces

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.319906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.046264Z digest=sha256:ce6ead254a272e5b4f2deb86287e3dd725a32d25788dc8e994eea2b28c12868f

Observation 123db3e9-4e4d-48a2-a0aa-32c47a58ebb7 · outbound

This paper cites Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio- visual deepfakes detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio- visual deepfakes detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.307141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.051021Z digest=sha256:5f3a8926c863d0376c90c06ac297548edf3e7ca06bbc9e5db5c5f5aa74bc5a15

Observation 55b1aaac-a140-4697-b961-17b7222667d7 · outbound

This paper cites Deeperforensics-1.0: A large-scale dataset for real-world face forgery de- tection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deeperforensics-1.0: A large-scale dataset for real-world face forgery de- tection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.294608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.055087Z digest=sha256:7923084223b5a511bf9d7e86c48e8bd2ef3249e98837a0f6fed668313c580a66

Observation 72e2ad54-abcd-42d0-84d3-daf329e5ea0f · outbound

This paper cites Con- textual cross-modal attention for audio-visual deepfake detection and localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Con- textual cross-modal attention for audio-visual deepfake detection and localization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.281106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.059518Z digest=sha256:2fecb732ebb24965105c956868cbc0e25f6f2fe1fa1314183d20811045bd4410

Observation 34f77d1b-062f-4df0-a60f-8c76ce62a0c8 · outbound

This paper cites The Kinetics Human Action Video Dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization The Kinetics Human Action Video Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.063363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.063363Z digest=sha256:286c4e5a77ecc48e3b8c38088445f4104b0fe9080348d729d5cf7b03695df619

Observation 33a3881c-034c-4b63-a1a7-85ad53ce3fbd · outbound

This paper cites Fakeavceleb: A novel audio-video multimodal deepfake dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Fakeavceleb: A novel audio-video multimodal deepfake dataset

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.269258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.067809Z digest=sha256:3a0a3043b8b59cff0cdd9f7da0387e23350f55e1c68175f6689216c416c5ff08

Observation 128aad35-bf40-4b2f-8a22-d537538d6ff6 · outbound

This paper cites Deep- fakes: a new threat to face recognition? assess- ment and detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deep- fakes: a new threat to face recognition? assess- ment and detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.253639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.071622Z digest=sha256:a1432a5ed9043d0ba56ad317b7aa85463f1445fa260bdcb77d5cedf9d03baafe

Observation 9446f9ec-fdb8-4fac-bc67-01faaf4750f7 · outbound

This paper cites Kodf: A large-scale korean deepfake detection dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Kodf: A large-scale korean deepfake detection dataset

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.240094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.075368Z digest=sha256:f8dd807edf000c49290f55a586bdd0ed885b476630a09c7a7f091d70b44d6abb

Observation 9100865c-9700-49b6-bad7-8fe3530a42c8 · outbound

This paper cites Layer normalization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Layer normalization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.228890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.078998Z digest=sha256:b70d16904a6dd5eec55b7d5310f83efb6f3e5141e162baa2d0e5a2aa0224830c

Observation 30998388-2085-4b43-bf51-7e0ca0e0bb0a · outbound

This paper cites Face x-ray for more general face forgery detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Face x-ray for more general face forgery detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.216358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.082609Z digest=sha256:7754792500228965141089ecb82c7c7ed5cf2cce23df537ecc686783b450ac03

Observation 2795419c-c1ad-4f57-9679-787517bac0a5 · outbound

This paper cites Spatio-temporal catcher: A self-supervised transformer for deepfake video detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Spatio-temporal catcher: A self-supervised transformer for deepfake video detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.203291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.086486Z digest=sha256:d4bc025c5e69a0cfc907deaf6f916fe4454843b577efab7fb54903f7a6ab6ad8

Observation bc8a5313-d2de-4a7f-98a3-7455c11e82ce · outbound

This paper cites Zero-shot fake video de- tection by audio-visual consistency.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Zero-shot fake video de- tection by audio-visual consistency

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.190739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.090412Z digest=sha256:d374d54d764e9a254fb5ddb4a03767dab117f0b5951142c802262857483f9b24

Observation 76094af0-f5b3-41ae-b72e-c166b7fd1d7b · outbound

This paper cites Celeb-df: A large-scale challenging dataset for deepfake forensics.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Celeb-df: A large-scale challenging dataset for deepfake forensics

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.176445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.093763Z digest=sha256:c20c1178cd21762e52a73b1e0b5f1e437d095d194b6d35a1b24b56d7061d2248

Observation d51e9110-c6bd-495a-b85b-027388a9ad0b · outbound

This paper cites Bmn: Boundary-matching network for temporal action proposal generation.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Bmn: Boundary-matching network for temporal action proposal generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.163658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.097499Z digest=sha256:5f7842ba2b49a6a1e99caa48a77cedcd5d95ffed2c9cebea324010dcb1d87f46

Observation a17a86e6-b252-433a-9ef0-df14ea693e57 · outbound

This paper cites Focal loss for dense object detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Focal loss for dense object detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.151984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.101090Z digest=sha256:3232f3077aa02a005cdc42eea5c40ffda4d7d6139f5b8347e52ca84a1193cd0d

Observation 7a519223-7193-4e09-97c5-753c7f2c48fe · outbound

This paper cites Visual speech recognition for multiple languages in the wild.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Visual speech recognition for multiple languages in the wild

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.135058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.104974Z digest=sha256:4e52720b8dd8cd00b8a548af8a16fbb44d7d62e97f5365fb4f5e7c743cccadd3

Observation b974cab0-f249-4722-9cf9-bcbcf2d2ac50 · outbound

This paper cites Deepfakes generation and detection: State- of-the-art, open challenges, countermeasures, and way forward.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deepfakes generation and detection: State- of-the-art, open challenges, countermeasures, and way forward

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.121575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.108472Z digest=sha256:231267a7726181ad712c4ac65e98faee6589a498a968ab433716aad7fdfc044d

Observation a9bb3df7-8242-421b-afab-390bb0433827 · outbound

This paper cites Emotions don’t lie: An audio-visual deepfake de- tection method using affective cues.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Emotions don’t lie: An audio-visual deepfake de- tection method using affective cues

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.105024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.111885Z digest=sha256:21b462f388bdf3a6c0838e1a092f91cf4f7a00ff3a70929e088971d0e600cb23

Observation e1baa66b-138a-4536-b6a5-ea3083ca0c7a · outbound

This paper cites Df-platter: Multi-face heterogeneous deepfake dataset.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Df-platter: Multi-face heterogeneous deepfake dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.092989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.115608Z digest=sha256:b4ce0bd062237759ca9e2e6428a3319d0d56327ca30b1c9910726e1fb4c3c53f

Observation dfc75e8e-dbc3-4116-8f63-43df994f805f · outbound

This paper cites A neural basis for interindividual differences in the mcgurk effect, a multisensory speech illusion.Neu- roimage, 59(1):781–787, 2012.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization A neural basis for interindividual differences in the mcgurk effect, a multisensory speech illusion.Neu- roimage, 59(1):781–787, 2012

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.079001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.119372Z digest=sha256:09f84aaea6f187aae76f4e9654abee9f57fc639cda4196556f9d9fbc6c844644

Observation b1696474-fa3f-42fe-877b-1097f2b4397b · outbound

This paper cites Activity Graph Transformer for Temporal Action Localization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Activity Graph Transformer for Temporal Action Localization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.123104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.123104Z digest=sha256:377d33f66fecd666fdb16f32f2f3c3328521764e58b1a9c01785f6f475f9552d

Observation cf3fd579-3604-42dc-9051-37c14cf3a53a · outbound

This paper cites Frade: Forgery-aware audio- distilled multimodal learning for deepfake detec- tion.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Frade: Forgery-aware audio- distilled multimodal learning for deepfake detec- tion

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.065982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.127280Z digest=sha256:ed17cf174c6d616c388178896622ebad17277472e228e606c290c91b19c5c66f

Observation 15b9793b-a54f-47f6-857b-5463f88c8646 · outbound

This paper cites Seeing what you hear: Cross- modal illusions and perception.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Seeing what you hear: Cross- modal illusions and perception

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.053027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.131051Z digest=sha256:ef10de8881605b921e15b8d4691150130d2d9990053a62a9b6def904e2e36c29

Observation 6f2caed7-7260-4e84-9fd9-dbe25f435ce2 · outbound

This paper cites Avff: Audio-visual feature fusion for video deepfake de- tection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Avff: Audio-visual feature fusion for video deepfake de- tection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.040514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.134766Z digest=sha256:1ca199d75b3dda429b618697ebcb830edd924ed26f9effed5671de480bac8a7f

Observation f3ecac87-f8dd-4142-ba71-d8a010b2c416 · outbound

This paper cites Deep- fake generation and detection: A benchmark and survey.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Deep- fake generation and detection: A benchmark and survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.138759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.138759Z digest=sha256:95a2cfa1360508e4bb917142411a75f05cf74619f153f2373be2e7a55913fa4e

Observation ba1211c5-7536-46e7-a7e9-9dea4da2337e · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Powerset multi-class cross entropy loss for neural speaker diarization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.142519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.142519Z digest=sha256:ae20f62710cf0f0be4be47559c934d5d8675ff0e63e2174960b49ddcc5be1110

Observation 80cb1353-d837-4424-95ee-54c02dfeea93 · outbound

This paper cites Thinking in frequency: Face forgery detection by mining frequency-aware clues.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Thinking in frequency: Face forgery detection by mining frequency-aware clues

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.027082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.146788Z digest=sha256:a7ee96dd55c56f7fc89192474285c58b0b035ff7c8c7c98742e8d11eab3bda33

Observation f939d77a-9b46-493b-a82a-9045a855e781 · outbound

This paper cites Multimodaltrace: Deepfake detection us- ing audiovisual representation learning.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Multimodaltrace: Deepfake detection us- ing audiovisual representation learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:44.012935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.151257Z digest=sha256:b394183fe75dd26be10828ee50c0dd2d944131328a5bfdbdaf46d739305018cd

Observation 329e15fd-4d41-4a0e-86fe-881e18cf9f75 · outbound

This paper cites Detecting Deepfakes Without Seeing Any.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Detecting Deepfakes Without Seeing Any

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.155342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.155342Z digest=sha256:a3611b61f4805f546cbca3474c59e8e87ec1f80619532e3faf3d84c545fdc65e

Observation 91ddaaa7-7b49-44a8-8f89-047904d854d2 · outbound

This paper cites Faceforensics++: Learning to detect manipulated facial images.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Faceforensics++: Learning to detect manipulated facial images

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.998593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.159335Z digest=sha256:ceefad5f009959c54d6bdaf5b9818f9069214021a8784758f4a7c2029480e8f0

Observation f6185d66-bcc0-4bbc-88dd-8a96f22f67ab · outbound

This paper cites Lip sync matters: A novel multimodal forgery detector.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Lip sync matters: A novel multimodal forgery detector

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.982817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.163041Z digest=sha256:76bffd332ea1643615add6c0349f04e66ccc3c70f82c1013b3447bf500d34b30

Observation 9d1cf413-d3d9-4cb3-b389-8ca124c88ef2 · outbound

This paper cites Visual illusion induced by sound.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Visual illusion induced by sound

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.969553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.166706Z digest=sha256:d3ad7a4f35192cb8852a722c951114b7c381b12f2385e41aceafceb59688aebd

Observation 303df92f-315a-472c-8734-3c2767c40494 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal clus- ter prediction.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Learning audio-visual speech representation by masked multimodal clus- ter prediction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.953594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.170440Z digest=sha256:f1db1ca01e2b9ebbcdb16eaf2d75311e3e89a74dbca2ca4ec9c36803b6ad7254

Observation e037fbe6-c3dc-4e02-b3a3-0e3912b881d6 · outbound

This paper cites Tridet: Temporal ac- tion detection with relative boundary modeling.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Tridet: Temporal ac- tion detection with relative boundary modeling

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.939120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.174270Z digest=sha256:e8a4d663669c92598a96f839a57fa2a2fa9fa862aec849b0af027d2e77ae41b7

Observation e8caeb53-edcf-4e59-8dee-008f1461e8c7 · outbound

This paper cites De- tecting deepfakes with self-blended images.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization De- tecting deepfakes with self-blended images

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.923424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.178194Z digest=sha256:a5bc9f194c393a93adff385c67bba903376bd7a69b2763a88dfa97b41dee03ff

Observation 75c2d7f8-85ad-4b4a-920c-e4dde18c964a · outbound

This paper cites Locate and ver- ify: A two-stream network for improved deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Locate and ver- ify: A two-stream network for improved deepfake detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.910567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.182180Z digest=sha256:21c59b8dee6f51c40f6ec9ebf437a8f459048d3bb902a9f390f6952284e52d0a

Observation 874ad5b2-a553-4c21-85a1-bea3412b0123 · outbound

This paper cites The contribution of visual information to the perception of speech in noise with and without informative temporal fine structure.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization The contribution of visual information to the perception of speech in noise with and without informative temporal fine structure

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.898396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.185881Z digest=sha256:4cdee76f3aa6d1de6611f503443fc646c5ef9e8164e499c00ac8f444b9dd47b4

Observation 5fe41125-7665-4625-939e-331c12026a91 · outbound

This paper cites Visual contribution to speech intelligibility in noise.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Visual contribution to speech intelligibility in noise

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.885985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.189720Z digest=sha256:ce78cc1eaebcbaf244fa54df03eac2ca7af59efdc57d36ea8b0fd13d1eb4f58c

Observation 143b8ad7-dc2a-42f9-8d1e-2beda29d4862 · outbound

This paper cites Learning on gradients: Generalized artifacts representation for gan-generated images detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Learning on gradients: Generalized artifacts representation for gan-generated images detection

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.873276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.193519Z digest=sha256:579411e8a75e4f11b695a88ae664492a10bb962226d9d7cf385a3b71244a69c2

Observation 0ed2dd83-ab49-4660-aec6-7281ee692def · outbound

This paper cites Attention is all you need.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Attention is all you need

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:57:43.197257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:57:43.197257Z digest=sha256:0162b875a5276cdb3c0d012fedbe451562eda75e4af9153460d9628e4a332814

Observation 86cfc1e0-69a2-4970-803c-3a463ef1ed7b · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Videomae v2: Scaling video masked autoencoders with dual masking

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.852680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.200837Z digest=sha256:f430731231e922bf3938f0aba58470d07da6acb2efdc6e1605a4bc103c5a0726

Observation 81309a39-bd8b-44ff-b3e4-0ae69a8adf35 · outbound

This paper cites Building robust video-level deepfake detection via audio- visual local-global interactions.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Building robust video-level deepfake detection via audio- visual local-global interactions

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.835415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.204546Z digest=sha256:3dc0bf75d8237fca36b787ffbb0949f2dd75b4b5d21803f74f61d2e123d4ed24

Observation 6cb7def3-e3dd-465a-9962-4b2a978190b2 · outbound

This paper cites Audio-visual deep- fake detection using articulatory representation learning.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Audio-visual deep- fake detection using articulatory representation learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.822420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.208375Z digest=sha256:1f5885a04445d1fdaba894f568c582c135c5409af811d0decc9d180458188e76

Observation c3b3b261-43bb-46e4-8398-ff632abee857 · outbound

This paper cites What you see is what you hear: sounds alter the contents of visual per- ception.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization What you see is what you hear: sounds alter the contents of visual per- ception

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.810469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.212662Z digest=sha256:261da878ea67f86dc63d6e0ce1352b66a401038aec7740d83c92f9044cd764b8

Observation acd175f2-4e6a-4f3a-bb1a-e9e2b52ac72a · outbound

This paper cites Avoid-df: Audio-visual joint learning for detecting deepfake.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Avoid-df: Audio-visual joint learning for detecting deepfake

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.799054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.216685Z digest=sha256:5e8ffb6c1d9c289ae22b8b88b4c2ddca773eadeb3a93abcb6f0d3237520dcc09

Observation 6097e91f-0892-40c8-bf82-6e61e843dd01 · outbound

This paper cites Expos- ing deep fakes using inconsistent head poses.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Expos- ing deep fakes using inconsistent head poses

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.787515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.221186Z digest=sha256:397b7ab65b271a185681b5734966d8c4bfc8e662d7beb26c5177b13c6618605a

Observation 91778712-caf1-48d4-9c23-be4d264ebf1b · outbound

This paper cites Pvass-mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Pvass-mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.773917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.224964Z digest=sha256:36bd549c4cab1bca2ce01597801bc29d3319603eec25829155e936185319edde

Observation 4c63d14e-9419-4b77-bf3b-2e33525c25e8 · outbound

This paper cites Ac- tionformer: Localizing moments of actions with transformers.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Ac- tionformer: Localizing moments of actions with transformers

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.760264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.228298Z digest=sha256:799c69a3e3e481971b70dd4f1365259303a69d5a9331011ad43513aed43c9da1

Observation dca4b936-a2e0-4a10-b38c-09c9f991b6c0 · outbound

This paper cites Video- llama: An instruction-tuned audio-visual lan- guage model for video understanding.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Video- llama: An instruction-tuned audio-visual lan- guage model for video understanding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.746605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.231704Z digest=sha256:627ac5aa240f802f12ba32f36c3968c0518c4c4a798b7c7ad94a9275f2e751bc

Observation 140244de-1acb-40f0-af06-5e4ebbacff0b · outbound

This paper cites Um- maformer: A universal multimodal-adaptive transformer framework for temporal forgery local- ization.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Um- maformer: A universal multimodal-adaptive transformer framework for temporal forgery local- ization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.735455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.235779Z digest=sha256:d129e2144b2ea76f8507e6baece5daca2644e0476b577e783f8596fdf75d7a3c

Observation 1b9a6618-3e17-480e-940e-b685275af55b · outbound

This paper cites Joint audio-visual attention with contrastive learning for more general deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Joint audio-visual attention with contrastive learning for more general deepfake detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.723760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.239365Z digest=sha256:475dd22f0bcc157bdf1048226d9acec93610d985a7ab74f353efb89cd89001d7

Observation 556ae367-c25b-4bba-a206-dd7c2c67cf69 · outbound

This paper cites Exploring temporal co- herence for more general video face forgery de- tection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Exploring temporal co- herence for more general video face forgery de- tection

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.710125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.243054Z digest=sha256:fa13c1588873161e9af125abb43c92c9b49d379ae0c20daf9ce80fd5f27f6eb8

Observation 186c60e4-8edd-49de-a06f-0e8b077fa5e3 · outbound

This paper cites Distance-iou loss: Faster and better learning for bounding box regression.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Distance-iou loss: Faster and better learning for bounding box regression

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.697029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.247172Z digest=sha256:5208262f4a14106622408cb4d757efe0a0010490f70023715b023dc449625251

Observation 2ccd7527-90c6-4436-8597-75225c08094f · outbound

This paper cites Joint audio- visual deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Joint audio- visual deepfake detection

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.683747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.250958Z digest=sha256:d9ab094b546ffab5b030a8931d8dfa9a65dc03fbb74353021d39336677bf6d1e

Observation 970e70eb-b674-44f6-ab45-932773b56d88 · outbound

This paper cites Wilddeepfake: A chal- lenging real-world dataset for deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Wilddeepfake: A chal- lenging real-world dataset for deepfake detection

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.672699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.254182Z digest=sha256:486fade19e088d41fc60ef80b42e615ff3914fb4e3306aa918644f2f6ffcd66a

Observation 1d2e6818-3e17-4b91-945f-d50179b9694f · outbound

This paper cites Cross- modality and within-modality regularization for audio-visual deepfake detection.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Cross- modality and within-modality regularization for audio-visual deepfake detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.661608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.257554Z digest=sha256:f5e7ad843dd6919182369d63b73c710335b71488c20709ef9ea23e744e2f2cab

Observation 40f4ca2d-d003-428c-8292-3bb3bfa32533 · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:57:43.649278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.261344Z digest=sha256:71db489d4998f91c5727d8621867e3fa6fbceb6668f8b5adb353fea5cea5117e

Observation a51072a7-9b24-4a70-b4d8-53248fabc08a · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:57:43.633490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.265660Z digest=sha256:6c7ef554268b02ef76aaac489551b6785d6261320a88d8b9d9eaffeac9908756

Observation 310d7c0c-8b90-4791-a1d4-b8dfcf45e59c · outbound

This paper cites Since we apply the manual screening process on synthesized videos, the final video count is more than 20,000.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Since we apply the manual screening process on synthesized videos, the final video count is more than 20,000

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.619113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.269730Z digest=sha256:3500f4c48b48a845f6552ae155782349cdb4b2a7d0bd87a4b757dcfbf5baf87e

Observation c71b1ba0-f73b-4698-b244-9d98710a070e · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:57:43.604141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.274100Z digest=sha256:961b94b621c6b5a59ccf0a363ecdc68b0c2b004281a30b237b4404ec83425eaf

Observation 222fe795-9d65-45a4-85ec-9a3f8a14d386 · outbound

This paper cites In addition, DiMoDif outperforms A VFF un- der all perturbation scenarios and at all intensity levels.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization In addition, DiMoDif outperforms A VFF un- der all perturbation scenarios and at all intensity levels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.590814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.277915Z digest=sha256:b185417bc23557b367fa2625b136c76b660c6e10c448607baeb01f310422e393

Observation 9a41d238-2c84-4614-b317-998cf3c071cf · outbound

This paper cites Table 12 presents the corresponding results.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Table 12 presents the corresponding results

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:57:43.578772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.281841Z digest=sha256:06e32f6cfb5132b69f5ac63b62b99c9c60e521b18e05202a45b008542414aa65

Observation bcb532c3-57a7-4c9a-a5e0-490816961776 · outbound

This paper cites an unresolved cited work.

DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization Unresolved cited work

Reference 93

Resolution
verified exact
raw_fallback, observed 2026-08-12T19:57:43.390782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T19:57:43.285908Z digest=sha256:49a30c2db437e990563592a00904df8fa27887fbdc6af19b3a3aba2d6ea1303f

Pith citing papers

Observation 2f3969ed-538d-4505-85f6-cf0859757533 · inbound

Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning cites this paper.

Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:44:30.005614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:44:30.005614Z digest=sha256:14a2227dcfe74008339cc73fa6f9d639dd3a793fa473430dad3a16254eea20a7

Observation e2840c89-2355-444f-bdac-136f94fd0a5d · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 163

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:23.115682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:23.115682Z digest=sha256:1e23124039de044489c71515e1c2804f7f773330bd3911e4218cfed0e801e9d0

Observation 96394362-d0dc-478a-bf7d-d7b9b080d66c · inbound

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection cites this paper.

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:37.830717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:37.830717Z digest=sha256:d0bf79539e9d341e08f15a396e22e7d246a504877b01d31f66275ce8abf0dc03

Observation a89fe335-7651-4d5e-a693-ba01739559a8 · inbound

Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization cites this paper.

Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:36.277635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:36.277635Z digest=sha256:72f32022ffdc9a8849ce741ff08f80de00a7fd481e1e0f348cddc08753ac0bda

Observation cb3528dc-8b71-41c0-b153-3295687cd0cb · inbound

Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems cites this paper.

Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 180

Resolution
unresolved
no resolver link, observed 2026-08-06T14:34:11.191723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:34:11.191723Z digest=sha256:a11f27953665f0781d53f3ed995cfadcb8dd1be154962f5db68747b1a13e96c8

Observation 4c29502d-5ae7-445d-8642-809ed8fb6505 · inbound

Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes cites this paper.

Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:25:59.992133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:39:31.212977Z digest=sha256:5238d0fec77290d9bb90b05ad77bf76897f62f0bd00b619225d5a95e464b5d36

Observation b034ae34-662b-41a5-ac29-7e8303e9d928 · inbound

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization cites this paper.

Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:24.495845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T05:19:38.661190Z digest=sha256:eb27577ec84ba565e77916ff2477abe08d1a1cc1e6d4f84768622c46d6fe399a

Observation 739272c8-217b-465e-8123-fb9e5cf0c772 · inbound

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization cites this paper.

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:07:02.289977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T14:03:46.585001Z digest=sha256:94d717c48ef18f1ccb50da4285f3040de9a5ef607e3ad272787c2591ab57c337

Observation 37442a88-3105-480f-9aa7-a64e3f0ae0ed · inbound

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration cites this paper.

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T18:55:21.867598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:55:21.867598Z digest=sha256:475827cc1d7899900bc8a3c3c4fc1865bd0884a3c08d8633b1db01cfed212a59

Observation ed579bdb-c456-4701-acfd-a3a3eded52d6 · inbound

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization cites this paper.

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T18:34:29.549966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:34:29.549966Z digest=sha256:3a6641fa38df4b67f83d7d381e4ba229d9a246dc8b9ae755c3c15cc8691cd1f4

Observation bd5aa6ed-9377-4bee-8c26-4f1972756867 · inbound

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization cites this paper.

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.084084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:52.084084Z digest=sha256:1c6180387b48d75af8eada9fabdf027a2f7f94f1fc93dd324cc5cb37ea3bf5bb