Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:52.053922Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.02178.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:52.053922Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69629e47-2e64-441f-9e27-0d3ef9f2be0e · outbound
Cocktail-Party Audio-Visual Speech Recognition Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09abe8b3-a228-4a21-9d8a-51e66e20fe00 · outbound
Cocktail-Party Audio-Visual Speech Recognition Task definition Given an input sequence of audio A = {a1, a2,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89a15ed6-0bee-4b18-b1f9-1b79cf888d84 · outbound
Cocktail-Party Audio-Visual Speech Recognition For training, we use LRS2 (train and pretrain sets), V ox2 (train set), and A VYT
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d388c060-6fb0-47e8-a097-b47628032925 · outbound
Cocktail-Party Audio-Visual Speech Recognition The A V-HuBERT CTC/Attention (A V1) model uses the A V-HuBERT large [12] as the encoder, which has 24 transformer blocks, each with 16 attention heads
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 544ba7dc-756e-42e9-b7c2-76383fb4e87e · outbound
Cocktail-Party Audio-Visual Speech Recognition The WERs for models evaluated on the original LRS2 test set are shown in the column where SNR = ∞
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b47fdb0-1712-40ff-9aec-a3e62004d8d3 · outbound
Cocktail-Party Audio-Visual Speech Recognition We highlighted the gap between conventional datasets and real-world cocktail- party scenarios, where target speakers are not always ac- tive
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8aa6569e-f973-45b0-9770-ef4ae4cfc25b · outbound
Cocktail-Party Audio-Visual Speech Recognition How is AI Changing Science? Research in the Era of Learning Algorithms
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 000f8e70-32b7-4932-bca3-bdb3b8fcf14e · outbound
Cocktail-Party Audio-Visual Speech Recognition Knowing who to listen to in speech recognition: Visually guided beamforming,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49f73fe2-8684-49d0-a265-347a8aa056a7 · outbound
Cocktail-Party Audio-Visual Speech Recognition Hearing lips and seeing voices,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0c50f4-0308-40d5-8c9d-1ad72f60d392 · outbound
Cocktail-Party Audio-Visual Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571c6281-057d-4e5e-a4bb-7ac627c59fdc · outbound
Cocktail-Party Audio-Visual Speech Recognition Super-Human Performance in Online Low-latency Recognition of Conversational Speech
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66de8a82-0444-497e-80e7-a82766e5d749 · outbound
Cocktail-Party Audio-Visual Speech Recognition Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4af78f2-617d-404a-ab3a-0487f7473dba · outbound
Cocktail-Party Audio-Visual Speech Recognition Recognition of conversational telephone speech using the janus speech engine,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0750410-d6e9-4e47-84c5-4340803780ca · outbound
Cocktail-Party Audio-Visual Speech Recognition See me, hear me: inte- grating automatic speech recognition and lip-reading,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d39485d-cc8e-4944-80d8-c594e9dc7e1c · outbound
Cocktail-Party Audio-Visual Speech Recognition Multimodal interfaces,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a64442cf-b5ec-473d-8a67-87295b2d83c8 · outbound
Cocktail-Party Audio-Visual Speech Recognition Modeling focus of at- tention for meeting indexing,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 895b5d0d-c927-4c2e-9035-fa9af05703ee · outbound
Cocktail-Party Audio-Visual Speech Recognition Visual track- ing for multimodal human computer interaction,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2fb81885-5478-49af-a611-a35f13ec12ad · outbound
Cocktail-Party Audio-Visual Speech Recognition Estimating focus of attention based on gaze and sound,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 449b42b3-e185-4d2f-8aa5-487dbb195c09 · outbound
Cocktail-Party Audio-Visual Speech Recognition Chil: Computers in the human interaction loop,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d352c00-65e7-45db-9210-f041cb23e2c1 · outbound
Cocktail-Party Audio-Visual Speech Recognition Robust self-supervised audio-visual speech recognition,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 417a4001-a7f1-4c34-bba2-e274b8007a83 · outbound
Cocktail-Party Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c68e33f3-e420-4d17-8582-18862c7170c1 · outbound
Cocktail-Party Audio-Visual Speech Recognition Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce23936d-fff7-4fcc-8e82-78f2229ee0c5 · outbound
Cocktail-Party Audio-Visual Speech Recognition Speaker-targeted audio-visual models for speech recognition in cocktail-party environments,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 616c2b1a-4dc2-409b-b6f2-cbab93fdb9a8 · outbound
Cocktail-Party Audio-Visual Speech Recognition Audio-visual multi-talker speech recognition in a cocktail party,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5412a3ce-84d3-4df4-9229-aed1624a57c4 · outbound
Cocktail-Party Audio-Visual Speech Recognition Robust audio-visual asr with unified cross-modal attention,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4257471-b8c4-48c8-98e7-4c55494e0247 · outbound
Cocktail-Party Audio-Visual Speech Recognition Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c0b254-0e57-4c40-aea5-d2534ecd2d31 · outbound
Cocktail-Party Audio-Visual Speech Recognition Visualvoice: Audio-visual speech sep- aration with cross-modal consistency,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a96f7f83-0fed-4d3a-b55b-31f2ebeec776 · outbound
Cocktail-Party Audio-Visual Speech Recognition Seeing through the conversation: Audio-visual speech separation based on diffu- sion model,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16f1a9e5-2c18-4ed6-9d0c-d6b801f490dc · outbound
Cocktail-Party Audio-Visual Speech Recognition Lip read- ing sentences in the wild,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ad5a0f-a6ca-467e-b941-c5f4ef9d3d18 · outbound
Cocktail-Party Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754f3861-d6af-4fd6-a7d5-268aafc9688d · outbound
Cocktail-Party Audio-Visual Speech Recognition V oxceleb2: Deep speaker recognition,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d149db78-4009-4645-a845-b1eaa41a963e · outbound
Cocktail-Party Audio-Visual Speech Recognition Unified cross-modal at- tention: Robust audio-visual speech recognition and beyond,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5671cdd1-37fa-430c-bbd5-13b737cd4ac9 · outbound
Cocktail-Party Audio-Visual Speech Recognition The first multimodal information based speech processing (misp) challenge: Data, tasks, baselines and results,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85d7a2be-fdbc-4a69-964e-a9c1d392050d · outbound
Cocktail-Party Audio-Visual Speech Recognition Summary on the multimodal information based speech processing (misp) 2022 challenge,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b46bbb2-dd24-4a08-b344-7fac75f5ebd8 · outbound
Cocktail-Party Audio-Visual Speech Recognition Summary on the multimodal information- based speech processing (misp) 2023 challenge,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5122e6c-0b43-417c-9412-5550698e7071 · outbound
Cocktail-Party Audio-Visual Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 044d076e-366b-4787-98de-26b2f36dca3a · outbound
Cocktail-Party Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5a526e-c436-4455-a64f-17f06af5a54a · outbound
Cocktail-Party Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e3ed08e-49da-42ae-bd49-95707080c23d · outbound
Cocktail-Party Audio-Visual Speech Recognition From text segmentation to smart chaptering: A novel benchmark for structuring video transcriptions,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b61790ee-525c-413c-be7c-5c2b623b9c80 · outbound
Cocktail-Party Audio-Visual Speech Recognition A light weight model for active speaker detection,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55c6a2c-03d5-4e7e-86af-cc3423175638 · outbound
Cocktail-Party Audio-Visual Speech Recognition Out of time: Automated lip sync in the wild,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad33d507-50b3-411d-8a83-3b4fc66134b2 · outbound
Cocktail-Party Audio-Visual Speech Recognition Synthetic conversations improve multi-talker asr,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a0f896-1316-491a-a9f5-6c8445a1a8ee · outbound
Cocktail-Party Audio-Visual Speech Recognition Msa-asr: Efficient multilingual speaker attribution with frozen asr models,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.