Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T15:03:07.926870Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2502.01547.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T15:03:07.926870Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:14.608210Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T00:02:14.902604Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 36e66a6d-34ba-4b8e-8440-2d59a16c0f13 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Robust speech recognition via large-scale weak super- vision,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d763d471-8483-4c04-af38-66ca2efec5ff · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Less is more: Accurate speech recognition & translation without web-scale data,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e46d493a-d02d-4e7e-8e52-584060f119c8 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre- Training for Adaptation to Unseen Languages,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8394274d-6913-418b-985b-3c572fdcc44a · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Ml-superb: Multilingual speech universal performance benchmark,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a221308b-4bbd-4d46-a4e3-583e078458d9 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Whisper-AT: Noise- Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df002455-96f9-4915-bb9d-748b31932171 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Large language models are efficient learners of noise-robust speech recognition,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e11d26f-00db-4d7e-b871-6cf861d0359c · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Deep audio-visual speech recognition,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3b4d0265-477c-4caa-b0d7-c12c2e6eb71c · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Audio-visual speech recognition with a hybrid ctc/attention architec- ture,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 44cac0ed-1028-4b4a-ba1a-f9a3878a805e · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition End-to-end audiovisual speech recognition,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02e3376b-7db0-4457-8a1c-c810671286a4 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Discriminative multi-modality speech recognition,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61138ab9-eb69-4a70-8cb4-7faa9da03991 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition End-to-end audio-visual speech recognition with conformers,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e509d71-fbdb-400e-82f7-49eae079021c · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Transformer-Based Video Front- Ends for Audio-Visual Speech Recognition for Single and Muti-Person Video,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8e6dce5-ce6a-450e-b873-b615a666b54a · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Robust Self-Supervised Audio- Visual Speech Recognition,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8a21a7c-fd7e-46df-ad2b-005c0ef57b1f · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 12d95f06-cb2d-4367-813e-ef6a257bb358 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Audio-visual efficient conformer for robust speech recognition,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2244be9d-215d-4123-8dd1-3bdf48305918 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Large Language Models are Strong Audio-Visual Speech Recognition Learners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2bc02ab-d9e4-47a3-b4ed-64447ba14f70 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Whisper-flamingo: Integrating visual features into whisper for audio-visual speech recognition and translation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7309830d-ebb4-4beb-821d-6dddeb9ad7f2 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Learning audio- visual speech representation by masked multimodal cluster prediction,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 546fd007-8919-4b06-997f-967cde156f6e · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Efficient training for multilingual visual speech recognition: Pre-training with discretized visual speech representation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08b9fe21-b622-4a98-b96a-a5814f0f093f · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recogni- tion and robust speech-to-text translation,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8fb9f648-784a-4b41-b679-2c46496d415e · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Flamingo: a visual language model for few-shot learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bb6baf1d-9960-4900-a057-09b9762b7737 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Improving neural networks by preventing co-adaptation of feature detectors
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8380c3a9-2e3c-4e3d-963e-c2e1c6015377 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Moddrop: adaptive multi-modal gesture recognition,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e0bd4aa6-222b-45eb-becd-281d14617ff0 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Recurrent neural network transducer for audio-visual speech recognition,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 29593e31-0d24-4cd0-99a4-bf8e0787951f · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition u-hubert: Unified mixed-modal speech pretrain- ing and zero-shot transfer to unlabeled modality,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d465ae67-070a-484e-a53a-296fb165f05a · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Av-data2vec: Self- supervised learning of audio-visual speech representations with contex- tualized target representations,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c74a130c-812f-4f37-85b5-88e9aaf7124f · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Av-cpl: Continu- ous pseudo-labeling for audio-visual speech recognition,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6b910bf-beed-4376-92a6-424d0d8bf13a · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Attention is all you need,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea984b1-97d7-4c3e-ab50-fc385ea22aac · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Reducing transformer depth on demand with structured dropout,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8d036736-1800-4ead-a958-6487c6c89f7d · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Lrs3-ted: a large-scale dataset for visual speech recognition,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c2b1e04f-f708-4d92-aefe-d2c845d46035 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition The multilingual tedx corpus for speech recognition and translation,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 385b4df5-66e6-4df6-a230-e3185c1c5f51 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Jointly learning visual and auditory speech representations from raw data,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8dbf6df7-8964-413f-b51a-8f046c7e1a1c · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Braven: Improving self-supervised pre-training for visual and auditory speech recognition,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14b41418-4d9d-49d2-b8a6-04cb133fa1d6 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Unified speech recognition: A single model for auditory, visual, and audiovisual inputs,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b97927d3-ede5-4728-9966-51d92880e896 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Visual speech recognition for multiple languages in the wild,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6bccbc4-25ff-42af-8208-12533f65f5e8 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Learning cross-lingual visual speech representations,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7110a677-aff4-4f0e-9b62-3da95a3a408e · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Lip reading for low-resource languages by learning and combining general speech knowledge and language-specific knowledge,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 63b682dd-ec3e-4a89-aa90-7fd72ce28d3b · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Visual speech recognition for languages with limited labeled data using automatic labels from whisper,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a7dce9ac-2401-4d27-8966-c86f8d680b07 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Dlib-ml: A machine learning toolkit,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 48fe8379-40b4-4ab5-92de-393c514e34db · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Lipreading using temporal convolutional networks,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eea94984-04c7-42ce-a740-939e941c3201 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Pytorch: An imperative style, high-performance deep learning library,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b00122c-32b2-4b3e-9bae-659036c7d26b · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition PyTorch Lightning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df6723af-8ffa-4c3d-9ca8-38dec7a295c3 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Decoupled weight decay regularization,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250fce94-1a33-4c98-9acf-8c22242b245f · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Musan: A music, speech, and noise corpus,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 761cd799-449e-4c39-a88f-be137e13379a · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 093781f0-359e-49ae-be0a-06a28840dfe0 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Intuitive multilingual audio-visual speech recognition with a single-trained model,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b7ad998-8934-4a4d-9567-022f9efe3597 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Multilingual audio-visual speech recognition with hybrid ctc/rnn-t fast conformer,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bbc9f116-afe0-4ff3-bec7-3cade1e228ba · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Parameter-efficient cross-language transfer learning for a language- modular audiovisual speech recognition,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4404c09d-99be-479d-be85-4e0fa608f3ef · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Tailored Design of Audio-Visual Speech Recognition Models using Branchformers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee80a905-f169-4e1c-af44-3f8f51b53760 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition Interleaved audio/audiovisual transfer learning for av- asr in low-resourced languages,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec092975-4e34-4db7-9888-f6f1929deaf8 · outbound
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31658b4e-5454-4c42-a0c0-c56795cce0fb · inbound
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.