Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:14.805352Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2505.03186.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:14.805352Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3ac9d1eb-96a8-49c5-b3b3-662a4e4126b6 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Robust speech recognition via large-scale weak supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 626beae2-331e-40c3-8d07-f7e07c6f1686 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf19573-56c2-41df-8141-8454e2d90440 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ee59dd-884d-426b-9b54-0dce5284d186 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Qwen2-audio technical report, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ae44e0-4685-4a19-88e0-25d0f1ebae40 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Assessment for automatic speech recognition: Ii
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e863d1-c7b0-4231-a73e-1fa10e3e797f · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Auto-avsr: Audio-visual speech recognition with automatic labels
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31658b4e-5454-4c42-a0c0-c56795cce0fb · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69e8cff1-ce87-47ee-8f48-4bf620f51cb2 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5ea468d-4a1f-4257-a40e-c9b49796ebc2 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d9da77-5ffd-4467-b0f9-98f8463fde04 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Unified speech recognition: A single model for auditory, visual, and audiovisual inputs, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4ee4a0e4-4941-4724-95f9-13f75958055a · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Speech recognition models are strong lip- readers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14646b13-963a-4262-9c08-d723ac268e36 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ea2047c-6fb3-401f-8059-8d4a7f85e7d0 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization van de Ven, Nicholas Soures, and Dhireesha Kudithipudi
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d79bc66-5e5f-44d1-b6b7-8df976f3894d · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lip reading sentences in the wild
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f2bd975-770e-4681-b5fb-98a978667072 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Maas: Multi-modal assignation for active speaker detection
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4e685917-7581-4da2-b044-28ecc04cadc3 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Out of time: automated lip sync in the wild
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a0b99c-cd91-4b60-a5a3-64a0b827c765 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization A lip sync expert is all you need for speech to lip generation in the wild
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faab5e2-a5a3-4722-84e2-dac9fc254347 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae52f234-c1ae-4b5f-a74d-31bc48f3a648 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Asr is all you need: cross-modal distillation for lip reading, 2020
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1108b9bb-53f2-4992-a221-18a83f7e9e60 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Schuller, and Maja Pantic
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 92fc2400-5d70-4ad9-91dc-86d271b22ca8 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality, 2022
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56b24ef1-c4cc-49f2-b99a-859acfff1e0a · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9ea9bd24-eaa7-4d29-802c-09159655c654 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 38f8b3c4-a53a-4c7c-bd27-8c4eca718d53 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Conformer: Convolution-augmented transformer for speech recognition, 2020
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87103140-2d4c-47aa-a9fa-5dfd522edce5 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Multilingual audio-visual speech recognition with hybrid ctc/rnn-t fast conformer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f2c10cca-c52b-4b20-8197-d95c478c2fa3 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lrs3-ted: a large-scale dataset for visual speech recognition, 2018
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b12aa4-1e3a-49f9-98c3-f0fe336b05b8 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f401350-0f5e-4d95-9617-f9bb28436777 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Braven: Improving self-supervised pre-training for visual and auditory speech recognition, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7e0f453d-368b-4502-bfa2-bcdfc91b8238 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Large language models are strong audio-visual speech recognition learners, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c61101-e1e3-4c65-bfe4-7ef2d5ad22b7 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Visualvoice: Audio-visual speech separation with cross-modal consistency, 2021
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9740d49-4e6c-4702-b84c-f0d4c86339d1 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Ctcnet: A cnn- transformer cooperation network for face image super-resolution
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5127e6c6-1d23-4421-ad8e-df41f411260e · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Muse: Multi-modal target speaker extraction with visual cues, 2021
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4de8b763-bce7-4c53-9f63-e2b381b24b73 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86797946-3f20-476d-9058-dc0a29b65927 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Attention is all you need in speech separation, 2021
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2306f3dc-d4e5-4afd-b244-0d46e6daa9c1 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Separate in the speech chain: Cross-modal conditional audio-visual target speech extraction, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 42cc8e8e-7598-4613-b881-5b0a41f0a44a · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization A light weight model for active speaker detection, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation efdd07a1-4504-4b6e-9538-2dfe98b7b255 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization End-to-end active speaker detection, 2022
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4bdcfad-cf29-4f35-bd6d-5257cdf90f1f · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Loconet: Long-short context network for active speaker detection, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 961a5d6d-51b1-41b1-af05-564ebfbbf10b · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lr-asd: Lightweight and robust network for active speaker detection
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ea3bdbd8-7c68-4200-a748-65a4d799d393 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Tdn: Temporal difference networks for efficient action recognition, 2021
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f3be7b5b-adc9-468b-9668-f40f52bc606d · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Dauphin, Angela Fan, Michael Auli, and David Grangier
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23630693-a3ce-4533-b4ae-8efd72aa2bd1 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Musan: A music, speech, and noise corpus, 2015
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d071a2-fb16-4b7b-bb6d-b6a310bdfb8d · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio- visual speech recognition with a hybrid ctc/attention architecture, 2018
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07be99f-6f8c-4ca5-b7a2-81137f33c9d7 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Es3: Evolving self-supervised learning of robust audio-visual speech representations
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73c5d417-7494-4daf-940f-cf1880addee0 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Syncvsr: Data-efficient visual speech recognition with end-to-end crossmodal audio token synchronization, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d3df58b-f8ac-4d6b-a411-aca47a226325 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Sub-word level lip reading with visual attention, 2021
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d5b44847-8856-4c5c-98f2-8588f060171a · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Deep audio-visual speech recognition
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 88998e02-a846-4a4b-80e2-f2dce41aa850 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition, 2022
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d474d59e-b104-4927-8778-875557e3b091 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio-visual efficient conformer for robust speech recognition, 2023
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8750052-a914-4197-980a-68f7e040c139 · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio-visual speech enhancement and separation by utilizing multi-modal self-supervised embeddings, 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1412804f-e3a4-46f8-be06-6b245d5192ae · outbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Time domain audio visual speech separation, 2019
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.