Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:44:02.577587Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.04635.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:44:02.577587Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:58.304998Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:44:02.849295Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 14534fc1-1de0-45a5-9a32-140aa25824c8 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cda3b0d1-e221-46f6-8a58-e7c8b0932664 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail dataset In this section, we describe the multi-stage pipeline for auto- matically generating a dataset for A VSR model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af47041d-62e8-4410-ac52-cee99ec66afa · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition The first model uses a Conformer-based encoder [31] with a CTC/Attention de- coder [32]
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c55a85f0-9deb-42c5-b28f-c3c9109eb22d · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34ae3160-da82-41d3-b579-faa8a6eb7395 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition natural”, “music
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8e7cc0e-08b4-4f64-b6c7-a5b6f96bd24f · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31069dcb-5d22-4b6d-8860-ed195367982f · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Our data collection and preparation process is fully automated and can be extended to other languages
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b9620e7-3120-47fc-83fe-813600c7f8d6 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Hearing lips and seeing voices,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e025039d-30c9-482c-9372-2276f42acc19 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Integration of acous- tic and visual speech signals using neural networks,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66ef00be-2e43-4000-8ffb-1a7c9f6da7fe · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition See me, hear me: inte- grating automatic speech recognition and lip-reading,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf6772e2-062b-4b11-a5f5-d8f0eaa910dd · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Multimodal interfaces,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ddaa6b4d-bd68-4df0-81f8-e6212c4d5f7f · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lip read- ing sentences in the wild,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1abe6788-433b-426a-b964-dccf5c0adfcd · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition RUSA VIC corpus: Russian audio-visual speech in cars,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97db2532-7544-4c47-9e69-7132363b267b · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Large-scale visual speech recognition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80da9692-e488-4ef2-8012-8ff559f16835 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Avas: Speech database for multimodal recognition applications,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16c16b9d-f679-400a-88c3-c35c431f57ec · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition An arabic visual dataset for visual speech recognition,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a85b86a1-6850-4757-a564-b82e1ddc902d · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Cn-cvs: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation caf27eae-3b0e-4f6a-b147-4f5a72c66176 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lipreading with densenet and resbi-lstm,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5fb88b5-9090-425a-afc2-c04eba3c98d2 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition A cascade sequence- to-sequence model for chinese mandarin lip reading,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6fb9e2-5068-4426-a192-7c4b4f6122bd · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82c188c4-f436-4a9b-8934-fe8c410b2b7b · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Havrus corpus: High-speed recordings of audio-visual russian speech,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bef02caf-cb17-4f85-b162-160266766334 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Towards es- timating the upper bound of visual-speech recognition: The visual lip-reading feasibility database,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42e40033-5729-4718-bbf1-5cdc395bb8ae · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Visual lip reading dataset in turkish,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b858014-4c99-4a7e-8f23-27711c5e815f · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition The ASD model relies on both audio and visual features to deter- mine when a person in the video is actually speaking
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 760ad800-c7ad-4ebc-8cc5-b6a609744b58 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lip reading in the wild,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b42453e-0765-48e7-947b-7acba01e11d4 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Robust self-supervised audio-visual speech recognition,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d456a7ff-1ee3-4d35-b82c-aaed08e9dc57 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 471ae582-469e-466f-8343-a189ba9faf54 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ef293d-e845-4682-afbc-87b2ad79bfc6 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Simple and effective zero-shot cross-lingual phoneme recognition,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2849a824-84b2-42d5-8a38-e82436792501 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition S3fd: Single shot scale-invariant face detector,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abbf2c26-b893-40cc-a4bf-c1e08487b3ea · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition A light weight model for active speaker detection,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fedafee6-38e3-4709-9e24-ab31cf6c6faa · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Out of time: Automated lip sync in the wild,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53c7fd6b-edc6-4728-81b2-38816c36a69c · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Dlib-ml: A machine learning toolkit,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8258663b-4295-4237-949e-95f1acb13e81 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Vietnamese end-to-end speech recognition using wav2vec 2.0,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06e23409-5baf-4e7e-a8e4-c44496ccd595 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Synthetic conversations improve multi-talker asr,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3e1757f-8f27-4af3-8d94-6012dfc76184 · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Msa-asr: Efficient multilingual speaker attribution with frozen asr models,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb1324eb-2107-40b9-8c39-0e37d2b80fcf · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Phowhisper: Automatic speech recognition for vietnamese,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5ceb486-d7ba-4b2a-af87-238661caba7d · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56da5b47-bc35-473c-a8de-6f583bad848d · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d35d24-1f69-4048-856e-89f12b758e9b · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Subword regularization: Improving neural network translation models with multiple subword candidates,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5dc24ec7-717d-4f4b-8817-26656b7bf6cd · outbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 14534fc1-1de0-45a5-9a32-140aa25824c8 · inbound
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.