Pith. sign in

Paper Citation Record · LEDGER

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2505.03186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03186 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:14.805352Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ac9d1eb-96a8-49c5-b3b3-662a4e4126b6 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Robust speech recognition via large-scale weak supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.575052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.575052Z digest=sha256:ce9568f7febb85e817bc0c450ed5bd2c1c155b2801f11c4cabd4683c87c49250

Observation 626beae2-331e-40c3-8d07-f7e07c6f1686 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.581161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.581161Z digest=sha256:828cd862610c0f1b2bb46053ef7d2bcbc9ee35b8e61713f7fa14c6af8a8a3d0e

Observation bbf19573-56c2-41df-8141-8454e2d90440 · outbound

This paper cites Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.586465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.586465Z digest=sha256:3c5b892c243401f2c67d66b1cf051edc496ce7b461285b4342cb4718b9caab4a

Observation 89ee59dd-884d-426b-9b54-0dce5284d186 · outbound

This paper cites Qwen2-audio technical report, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Qwen2-audio technical report, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.592073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.592073Z digest=sha256:c5669ccd51ad4fe09289b98ff25f92c3a6fbce229ae34f761ce1a0d066848351

Observation b8ae44e0-4685-4a19-88e0-25d0f1ebae40 · outbound

This paper cites Assessment for automatic speech recognition: Ii.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Assessment for automatic speech recognition: Ii

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.597519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.597519Z digest=sha256:29f0e3e86a609e5bdddb0883d0b8561f7ad090da8478dedb7ae708cc75a3323f

Observation 18e863d1-c7b0-4231-a73e-1fa10e3e797f · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.495059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.602336Z digest=sha256:384650ce19b9486624364b7c531eae2135d41753195f1825110a9c04f965aa5d

Observation 31658b4e-5454-4c42-a0c0-c56795cce0fb · outbound

This paper cites mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:02:14.909462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.608210Z digest=sha256:fe5d47282279650f5e480c0b2a59f3767e9ce060b9941fbc40c52c327d411556

Observation 69e8cff1-ce87-47ee-8f48-4bf620f51cb2 · outbound

This paper cites Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.480421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.613614Z digest=sha256:83fbc3ddf5e70c176ae4bbd8144cb8a20966078c48393738899bf2de2ca9763a

Observation a5ea468d-4a1f-4257-a40e-c9b49796ebc2 · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.618734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.618734Z digest=sha256:c447900519e6cdf2f3d95de8198cbc3e85242f2f6a093a7abaa7a1fdb7ae3b96

Observation 99d9da77-5ffd-4467-b0f9-98f8463fde04 · outbound

This paper cites Unified speech recognition: A single model for auditory, visual, and audiovisual inputs, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Unified speech recognition: A single model for auditory, visual, and audiovisual inputs, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.464167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.623425Z digest=sha256:43de6fcdcd8639561b00aea7505dbdbd5900c4052194fd8de60653249cd0e1c6

Observation 4ee4a0e4-4941-4724-95f9-13f75958055a · outbound

This paper cites Speech recognition models are strong lip- readers.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Speech recognition models are strong lip- readers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.449473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.627623Z digest=sha256:5d16c638d11f39f8843d6c8db36310533b3c40839e0dca4965ec289339e26328

Observation 14646b13-963a-4262-9c08-d723ac268e36 · outbound

This paper cites Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.631987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.631987Z digest=sha256:ba8efdd794871f931af874dc996f2b217b1325c4ef2ce8a093cfc612685810fa

Observation 5ea2047c-6fb3-401f-8059-8d4a7f85e7d0 · outbound

This paper cites van de Ven, Nicholas Soures, and Dhireesha Kudithipudi.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization van de Ven, Nicholas Soures, and Dhireesha Kudithipudi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.434451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.636993Z digest=sha256:d4b03f88248e814b37f40831759df6e6c778503cb55cf18fbb9179ce5c7f723c

Observation 6d79bc66-5e5f-44d1-b6b7-8df976f3894d · outbound

This paper cites Lip reading sentences in the wild.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lip reading sentences in the wild

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.420078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.641960Z digest=sha256:246856c3422e1b55ac78f2e168dcac31e669d2d6bdbe32712ccdcb329834b825

Observation 9f2bd975-770e-4681-b5fb-98a978667072 · outbound

This paper cites Maas: Multi-modal assignation for active speaker detection.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Maas: Multi-modal assignation for active speaker detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.405666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.646482Z digest=sha256:5ae988b4eb7169b3c9786deac09cd43f72ea66a7c50afe90a86ae51caf055e08

Observation 4e685917-7581-4da2-b044-28ecc04cadc3 · outbound

This paper cites Out of time: automated lip sync in the wild.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Out of time: automated lip sync in the wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.650794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.650794Z digest=sha256:1b11ee4182387c71f8e8d1385a5382e913b6334ed78321af432d7bea2e99108e

Observation 40a0b99c-cd91-4b60-a5a3-64a0b827c765 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization A lip sync expert is all you need for speech to lip generation in the wild

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.655093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.655093Z digest=sha256:54e706125b006d06473938f110f288926b59ea309d66abeff3518c23b9202698

Observation 9faab5e2-a5a3-4722-84e2-dac9fc254347 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.659428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.659428Z digest=sha256:98b5a7305dcf735fff03effe2fc388e28b32b529796f94f9de2dde14c4d83c75

Observation ae52f234-c1ae-4b5f-a74d-31bc48f3a648 · outbound

This paper cites Asr is all you need: cross-modal distillation for lip reading, 2020.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Asr is all you need: cross-modal distillation for lip reading, 2020

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.371184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.663974Z digest=sha256:df04c0b3d38525812b3c45a415813d595905dbdbf917f7ae8c3420d21ba13937

Observation 1108b9bb-53f2-4992-a221-18a83f7e9e60 · outbound

This paper cites Schuller, and Maja Pantic.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Schuller, and Maja Pantic

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.356605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.668210Z digest=sha256:d73b04c0c5998230d8cfbba3043ac7a12df9e9645051bb96275d1508d3c02e8c

Observation 92fc2400-5d70-4ad9-91dc-86d271b22ca8 · outbound

This paper cites u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality, 2022.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.342794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.672686Z digest=sha256:33a1d6bafe43aa576813271c6433d207506abd0c78659347f4cdead4bc8238e7

Observation 56b24ef1-c4cc-49f2-b99a-859acfff1e0a · outbound

This paper cites Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Av-data2vec: Self-supervised learning of audio-visual speech representations with contextualized target representations, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.328435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.676961Z digest=sha256:10a57781c1dc90b326c09d1de8793c328ae3d887c98e6134fceb71b81dc3fe23

Observation 9ea9bd24-eaa7-4d29-802c-09159655c654 · outbound

This paper cites Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.314319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.681001Z digest=sha256:2b1c374ff2081baee2dfbd3ab455e3523fc2f663a89d58df9db74ccc8a2f7900

Observation 38f8b3c4-a53a-4c7c-bd27-8c4eca718d53 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition, 2020.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Conformer: Convolution-augmented transformer for speech recognition, 2020

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.685115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.685115Z digest=sha256:711d74cbbff11162bef75c805a7ba6bff8f3910859d995c5f8d847356268cbd9

Observation 87103140-2d4c-47aa-a9fa-5dfd522edce5 · outbound

This paper cites Multilingual audio-visual speech recognition with hybrid ctc/rnn-t fast conformer.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Multilingual audio-visual speech recognition with hybrid ctc/rnn-t fast conformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.288628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.690215Z digest=sha256:00445918b4140cf84f77dac88dd197ec5a9eb6b9ebb2a3a495935b1cf9a94827

Observation f2c10cca-c52b-4b20-8197-d95c478c2fa3 · outbound

This paper cites Lrs3-ted: a large-scale dataset for visual speech recognition, 2018.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lrs3-ted: a large-scale dataset for visual speech recognition, 2018

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.695123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.695123Z digest=sha256:002845bb3437d22ccfa279a20744813eaecd2a7ecb184f76b3966ea1f786cb14

Observation c6b12aa4-1e3a-49f9-98c3-f0fe336b05b8 · outbound

This paper cites MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.699358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.699358Z digest=sha256:3aaa6ee2619c6daaa571636cf7b5b9af3f8a8faa57e4d5d0e5c1efe9ab52ebcc

Observation 9f401350-0f5e-4d95-9617-f9bb28436777 · outbound

This paper cites Braven: Improving self-supervised pre-training for visual and auditory speech recognition, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Braven: Improving self-supervised pre-training for visual and auditory speech recognition, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.263968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.703772Z digest=sha256:6a7b9d2abd5ea41ce42ba994ea1f412de82232c86d418b7372ae250cba955f75

Observation 7e0f453d-368b-4502-bfa2-bcdfc91b8238 · outbound

This paper cites Large language models are strong audio-visual speech recognition learners, 2025.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Large language models are strong audio-visual speech recognition learners, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.708017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.708017Z digest=sha256:e3723aa4beafb4bdfd982dc348987280cef6523bfd931a3b2a6efe0399c73a07

Observation a4c61101-e1e3-4c65-bfe4-7ef2d5ad22b7 · outbound

This paper cites Visualvoice: Audio-visual speech separation with cross-modal consistency, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Visualvoice: Audio-visual speech separation with cross-modal consistency, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.238909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.712253Z digest=sha256:7a60b846105668b879ef8d2f1bfc9fe4749cce6ea47c6c85df151adca8e7abb4

Observation b9740d49-4e6c-4702-b84c-f0d4c86339d1 · outbound

This paper cites Ctcnet: A cnn- transformer cooperation network for face image super-resolution.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Ctcnet: A cnn- transformer cooperation network for face image super-resolution

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.224269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.716572Z digest=sha256:1d9c8631a3a2d25ebbbac94183f17920ef2009aac170273393a1f667a893de64

Observation 5127e6c6-1d23-4421-ad8e-df41f411260e · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Muse: Multi-modal target speaker extraction with visual cues, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.209791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.720990Z digest=sha256:789dba49d7354d99cf1014a89af44f42c0994dab2651ce2fd27d0dc931f02446

Observation 4de8b763-bce7-4c53-9f63-e2b381b24b73 · outbound

This paper cites Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.193757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.725497Z digest=sha256:b288740ec65e4ec2d8768b835346183e627465a1c5777f169612ec530b7fcf0f

Observation 86797946-3f20-476d-9058-dc0a29b65927 · outbound

This paper cites Attention is all you need in speech separation, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Attention is all you need in speech separation, 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.179881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.729752Z digest=sha256:8ed34899fcbea8cafd3a1ec4812a7be61e556f1a0fe5104442d664ec412b8344

Observation 2306f3dc-d4e5-4afd-b244-0d46e6daa9c1 · outbound

This paper cites Separate in the speech chain: Cross-modal conditional audio-visual target speech extraction, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Separate in the speech chain: Cross-modal conditional audio-visual target speech extraction, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.163493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.734164Z digest=sha256:cb7b557c3f445d69c97a112796717b511e99feb4d07cbf8f1e091db693668405

Observation 42cc8e8e-7598-4613-b881-5b0a41f0a44a · outbound

This paper cites A light weight model for active speaker detection, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization A light weight model for active speaker detection, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.147444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.738490Z digest=sha256:cd54ee986a2a476444e7969f2f163ba6a77060e04146484d0b693cd7c320cac7

Observation efdd07a1-4504-4b6e-9538-2dfe98b7b255 · outbound

This paper cites End-to-end active speaker detection, 2022.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization End-to-end active speaker detection, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.132880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.743048Z digest=sha256:26227eae188c77623c1529704f83fad2965c3b93e0ba747bf24d21ffa1179c8f

Observation f4bdcfad-cf29-4f35-bd6d-5257cdf90f1f · outbound

This paper cites Loconet: Long-short context network for active speaker detection, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Loconet: Long-short context network for active speaker detection, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.118510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.747136Z digest=sha256:28573b0784a3983bfdb72438d76335fbc4311f30681b93f0f077642c2d39b01b

Observation 961a5d6d-51b1-41b1-af05-564ebfbbf10b · outbound

This paper cites Lr-asd: Lightweight and robust network for active speaker detection.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Lr-asd: Lightweight and robust network for active speaker detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.104097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.751398Z digest=sha256:7c08ebd36543521f5d5a2d34312d6403a9cc4505af3f58b4c2cd6eb95975d0c6

Observation ea3bdbd8-7c68-4200-a748-65a4d799d393 · outbound

This paper cites Tdn: Temporal difference networks for efficient action recognition, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Tdn: Temporal difference networks for efficient action recognition, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.089734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.755649Z digest=sha256:cd39f2fee7b186cac9fab24a7abb1dc96745e78851707b93476dfae126d22608

Observation f3be7b5b-adc9-468b-9668-f40f52bc606d · outbound

This paper cites Dauphin, Angela Fan, Michael Auli, and David Grangier.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Dauphin, Angela Fan, Michael Auli, and David Grangier

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.759904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.759904Z digest=sha256:83e3cbc1314acd7ae33d8f7e66427990a2268b8b8de1c0ed08361140337779ea

Observation 23630693-a3ce-4533-b4ae-8efd72aa2bd1 · outbound

This paper cites Musan: A music, speech, and noise corpus, 2015.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Musan: A music, speech, and noise corpus, 2015

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.764101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.764101Z digest=sha256:71e6f234f22ea0cad7a03a039eccda9a7cd47ad3c4851305277719342797aa8b

Observation 04d071a2-fb16-4b7b-bb6d-b6a310bdfb8d · outbound

This paper cites Audio- visual speech recognition with a hybrid ctc/attention architecture, 2018.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio- visual speech recognition with a hybrid ctc/attention architecture, 2018

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.768340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.768340Z digest=sha256:0d235ac8856e428e81ae713d35dc9e734317ab2d11ad4331e79ba000ecb85c47

Observation c07be99f-6f8c-4ca5-b7a2-81137f33c9d7 · outbound

This paper cites Es3: Evolving self-supervised learning of robust audio-visual speech representations.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Es3: Evolving self-supervised learning of robust audio-visual speech representations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.045641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.772581Z digest=sha256:9d22153ca286a348663f36aa563a1e6bf6159b41ce2edff7a557c8a797901db8

Observation 73c5d417-7494-4daf-940f-cf1880addee0 · outbound

This paper cites Syncvsr: Data-efficient visual speech recognition with end-to-end crossmodal audio token synchronization, 2024.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Syncvsr: Data-efficient visual speech recognition with end-to-end crossmodal audio token synchronization, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.031534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.777319Z digest=sha256:8bc687deabe48f66f547247f315fd57f2d1d39ab73dc0f6f38e8dcaa3d3755b1

Observation 6d3df58b-f8ac-4d6b-a411-aca47a226325 · outbound

This paper cites Sub-word level lip reading with visual attention, 2021.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Sub-word level lip reading with visual attention, 2021

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.016589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.781608Z digest=sha256:f985659cc24c8103fecabfb2529ef477e02e1c69784b5c1b5499384b49533fb2

Observation d5b44847-8856-4c5c-98f2-8588f060171a · outbound

This paper cites Deep audio-visual speech recognition.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Deep audio-visual speech recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:15.000505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.787191Z digest=sha256:9432ba1393d8374063f2c4985269f2ac99aff99db5f01b27d146898ba28ffd0e

Observation 88998e02-a846-4a4b-80e2-f2dce41aa850 · outbound

This paper cites Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition, 2022.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.986821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.792140Z digest=sha256:f5499ff0143701e2f9d24b374112288ddf3c6720877f3a797191fbbabfb3d42c

Observation d474d59e-b104-4927-8778-875557e3b091 · outbound

This paper cites Audio-visual efficient conformer for robust speech recognition, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio-visual efficient conformer for robust speech recognition, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.972183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.796585Z digest=sha256:7330d48c0bfe0648050a84cfa98aec07280fbd0fb2f2080d2781273f6bb43c9e

Observation d8750052-a914-4197-980a-68f7e040c139 · outbound

This paper cites Audio-visual speech enhancement and separation by utilizing multi-modal self-supervised embeddings, 2023.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Audio-visual speech enhancement and separation by utilizing multi-modal self-supervised embeddings, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.957484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.801054Z digest=sha256:e340bb3fcc65f48ea9d1850340dcfccb0638e99a3c4d2bfc69bda999c71b9092

Observation 1412804f-e3a4-46f8-be06-6b245d5192ae · outbound

This paper cites Time domain audio visual speech separation, 2019.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Time domain audio visual speech separation, 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:14.942609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:02:14.805352Z digest=sha256:6b9d9bea7c5f660ff656e81a7fcc32c5d984fe1267a19f15f13a197c5c94db4b

Pith citing papers

No inbound Pith citation observations are available.