Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2505.14336.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:28.786395Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:39:34.749620Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eb34c564-0255-47f0-a5af-47e66d1a49da · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 854669e9-b19b-4a3d-afb6-81d7d591a698 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach This is crucial in resource-constrained LLM-based A VSR systems, as we aim to improve performance despite using smaller-scale LLMs and pre-trained encoders
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b7ab5303-54d5-4ca6-9a85-233cd69bf06c · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Transcribe{task prompt}to text
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c0562e3-906b-4377-854b-4debf0c19cc5 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Its key innovation is replacing the lin- ear projector with a Top-K sparse MoE module
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48b325f5-22a0-4758-bcab-bcf4935c0bb3 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af2e464-b6fa-45be-8ebc-79f4ad2c0280 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Deep speech 2: End-to-end speech recognition in english and mandarin,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 77e79aba-3861-4ae5-a3ff-d9b6241e220a · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach End-to-end speech recognition: A sur- vey,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a7ebaeff-76fa-40d6-beee-51a715d9f196 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audio-visual speech modeling for continuous speech recognition,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 188010e7-f95a-4c3e-a6b5-424bef2cbb2e · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Investigation of speech separation as a front- end for noise robust speech recognition,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4af7a8d2-2900-4128-8e72-eb6ca82d3e0d · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audio-visual speech recognition using deep learn- ing,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 35050d51-0909-49ee-bb22-e1d9811c4f07 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Deep audio-visual speech recognition,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4162f35f-1e8f-4e3a-a19e-5cce765eb734 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audio-visual speech recognition with a hybrid ctc/attention architecture,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cb370472-9c29-485b-a705-d25f88a2af39 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach End-to-end audio-visual speech recognition with conformers,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ebdfcbc3-9b9c-4648-bee3-ab9970c64991 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Visual context-driven audio feature enhancement for robust end-to-end audio-visual speech recognition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f0b6873-112f-4b70-9cf0-599666223be1 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Auto-avsr: Audio-visual speech recognition with automatic labels,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 96890755-e6b6-4be8-b254-2881b9b49b1b · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Whisper-flamingo: Integrating visual features into whisper for audio-visual speech recognition and translation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d83c6aa0-e7a5-492c-a3e1-f8c869c0f0b9 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach A survey on self-supervised learning: Algorithms, applications, and future trends,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4242120d-d12d-4b0a-b055-bdd8319be364 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Learning audio-visual speech representation by masked multimodal cluster prediction,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c333ce47-1df5-4dbe-9ef5-2ee24958b68e · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Jointly learning visual and auditory speech representations from raw data,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c8e2a37-9a4f-4736-a29b-c404690f6afa · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2bbe6460-a9ef-466d-a12c-3fded2e3cf02 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Braven: Improving self-supervised pre- training for visual and auditory speech recognition,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a089d9ec-d62e-4726-ada1-d02a18e3969d · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Unified speech recognition: A single model for auditory, visual, and audiovisual inputs,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation feaadf66-b34a-4cb0-8f27-6aa4e0bab2d8 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach GPT-4 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec5a96e-b54c-437f-b06c-851942001bb4 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA: Open and Efficient Foundation Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f279aed-e8f4-4d24-80a7-c207d0f18020 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Improved baselines with visual instruction tuning,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d470f4f4-f794-4237-a9a3-c508819393ac · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach On generative spoken language modeling from raw audio,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72bd5ed8-d435-45e2-8f0d-ac21070670db · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audiogpt: Understanding and generating speech, music, sound, and talking head,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 408ee907-d152-4960-835c-8966203d05aa · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Let’s go real talk: Spoken dialogue model for face- to-face conversation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8170e4d9-a3c6-43a0-8b53-310b3f5797dd · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Developing instruction-following speech language model without speech instruction-tuning data,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b49faad6-117c-4d7a-94ac-ada9c3300905 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach SSR: Alignment-Aware Modality Connector for Speech Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7835cb45-c654-45f2-adb9-91e41bf36593 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach It’s never too late: Fusing acoustic information into large language models for automatic speech recognition,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3043acb5-fa44-4907-9105-98c275627aef · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Large language models are efficient learners of noise-robust speech recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fde77f27-3377-4510-a71b-9d197e0e108b · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f32b6a4-acba-4693-adaf-d1b547c9d30f · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Connecting speech encoder and large language model for asr,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b27c3acf-289d-4f57-b72d-08758ab950f9 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Prompting large language models with speech recognition abilities,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8b5be17-de52-4402-9183-5d8945b7a122 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech process- ing,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4bd81aca-8e3e-4140-8c0c-b766d7a78aea · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Large language models are strong audio- visual speech recognition learners,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 74c03ad7-df57-4179-97a5-233f70fed895 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0fded2c8-7c41-4daf-8b4e-16ea2954e66a · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f6b4946b-c23a-41ea-82be-6456a319d174 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Gshard: Scaling giant models with condi- tional computation and automatic sharding,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ee17e29d-dfee-4a80-9d27-a387a5a7a269 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Efficient fine-tuning of audio spectrogram transformers via soft mixture of adapters,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba451649-cbec-430b-9c6f-ea5250ed1c15 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6287b0f-f201-490d-b5b4-5e8dc12f5861 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Mixture of A Million Experts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e2b32c4-466f-4a0d-a84f-f9d077ad07d9 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Olmoe: Open mixture-of-experts lan- guage models,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 160b6d9e-fe8d-43f6-85ae-7b74809bd176 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e513e606-4937-4e6c-a932-ad186f76fced · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Chartmoe: Mixture of expert connector for advanced chart understanding,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 34b72aa8-89c5-4db9-b9c7-da0fbd3ffdd5 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Dense connector for mllms,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1884dd08-84ba-4140-bcec-d54b4ef10f95 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Visual instruction tuning,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b7540b4-6bb1-47d6-903e-38d789347c5d · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Vila: On pre-training for visual language models,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c8812366-15d3-45be-8a37-11d69eec5b6e · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbefa5de-ac70-46fe-8edc-5698f6e910c5 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Meteor: Mamba-based traversal of rationale for large language and vision models,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a04070d8-7a4a-48f0-aaea-7ca62d1dd6ca · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Lora: Low-rank adaptation of large language mod- els,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c76be01e-d01f-4fa6-9cc7-8aa6b74ce86f · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a45b99a-6465-4660-92f8-72b2677fe431 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae50937-c189-4c57-afcb-9d96beb8aa9e · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LRS3-TED: a large-scale dataset for visual speech recognition
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a5e078-3f87-49dc-bd75-d48bb2d935d7 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Robust speech recognition via large-scale weak supervision,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cf986061-082d-400a-9f3a-ca53fa4dd693 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach The Llama 3 Herd of Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6916c243-6717-40fb-86ef-3cb2febef3aa · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Towards a unified view of parameter-efficient transfer learning,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 938640d8-9e79-4b54-b176-dde2deed5f17 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Parameter-efficient transfer learning of au- dio spectrogram transformers,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c07ba69b-84a4-4031-a186-ef95e4cb43fe · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088c4cf1-edc1-41ea-9160-20f5dfcdbfe9 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Mixtures of experts for audio-visual learning,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 10189419-0fe1-4ad1-9f36-c209cea88af3 · outbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb34c564-0255-47f0-a5af-47e66d1a49da · inbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.