Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2201.02184.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:53:50.391701Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:49:41.728084Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation fd50aea4-84af-4bc9-8022-d12eb1547d04 · inbound
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fcc39cd-9f1e-49b8-8764-0adcdd1345be · inbound
MuteSwap: Visual-informed Silent Video Identity Conversion Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120056b9-0ef1-4f67-994f-127feb7cfcbe · inbound
Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 808bfc5e-bf43-4fc8-9399-8e9f2e1dc2de · inbound
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9117a31-9416-46fa-9714-2941e7c61f58 · inbound
HumanOmni-Speaker: Identifying Who said What and When Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16998ea3-2045-49cb-80c7-b9342ea01be9 · inbound
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8176aad-ec4b-4cb7-b5af-e25c1a36bbb5 · inbound
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 709bdce4-0dfc-4a2d-b7a0-eb7505c97883 · inbound
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation caa96176-d73b-41ad-9fca-42357ca0e1d2 · inbound
Inconsistency-aware Multimodal Schr\"odinger Bridge for Deepfake Localization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1ac8eb4-36f1-4f64-a1c0-65dcd4a3abf9 · inbound
Your Multimodal Speech Model Says I Have a Face for Radio Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 994b35a5-0354-4e72-b202-f8ccd82040fe · inbound
Evaluating Multimodal Steganalysis for Split-Payload Audiovisual Steganography Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 641792ed-feaf-437a-8aec-8387dd5eddc8 · inbound
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68ffcd41-0334-43b0-9014-0b48bc71b45c · inbound
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e54e3b54-3554-46c1-b3f0-f9ec27863513 · inbound
Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 524a6f4f-0fc6-4ba4-a4ca-47e4a843d284 · inbound
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47d8d709-c8a7-4fa2-9f76-942b31511c2b · inbound
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00cdfd15-4ba4-4c9b-bfe7-2a6d6267fb59 · inbound
Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.