Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:08:19.017587Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19037.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:08:19.017587Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T17:45:51.528645Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T06:11:01.750278Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c3e573cf-65f3-49a9-b1e5-55b9c13ca43f · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f90e351b-e603-4ff9-9c94-b1c93912ee2a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ea53f2-f7e7-40b5-b809-13cb0a4817f0 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Next-gpt: Any-to-any multimodal llm,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b42e5b2e-62da-443f-88ee-1b53fa616297 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8cdab4-09ad-4202-a4bb-e82f7aad9105 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1e7b3b71-36b7-492f-8ccf-44d5c78c4b62 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a0b88f1b-85f6-4be0-b476-0c7d23eddc74 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speech recognition meets large language model: Benchmarking, models, and exploration,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 678efdf9-8030-46c2-8235-4433d439abe0 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Hubert: Self- supervised speech representation learning by masked prediction of hidden units,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c7dd4b52-4a83-4b33-bfb7-190f8bcc8de3 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? A comprehensive review of multimodal large language models: Performance and challenges across different tasks,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 913df5e4-5cd2-486e-8865-3de22a8503af · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d24fb121-cf20-4d97-9c81-0863c290eca7 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b3c3f9-e185-4528-8319-63a29e8699c4 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are strong audio-visual speech recognition learners,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0d0b914e-cf80-4843-919d-bcef5a4926cf · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1e39b58-d6a1-443a-9657-d0a941bdac5f · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are efficient learners of noise-robust speech recognition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07648979-b6d1-4176-bf51-4eba9e608b8a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avatar: Un- constrained audiovisual speech recognition,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a779a849-e73d-4472-b0b9-d18e1f4b451c · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6cbf5381-b990-4157-ab70-8e85debd40f0 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba: Linear-time sequence mod- eling with selective state spaces,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9454dc90-f285-4674-9527-e0c0beed982c · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Improved baselines with visual instruction tuning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d51339a1-297b-4b06-8b8c-81d0a4500bdf · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81f9b08-7c84-4b94-a780-d72efe7c4c06 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aaf9ea8d-38c8-43ae-bf8b-257ce7e9322b · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6587fa31-2a7e-4651-b08e-900b35ef08b0 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cobra: Extending mamba to multi-modal large language model for efficient inference,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 10d2330a-fdf2-48ab-a2a9-e37e4f767d12 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7637d7e-45e8-4168-a67f-e961fdc2d588 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Vl-mamba: Exploring state space models for multimodal learning,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c1523190-95e2-410a-8d53-10aacc321d34 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3636cc29-cd3c-4e8c-94e0-57b2dabd6e74 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Diffusion- lm improves controllable text generation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5be19410-7fcc-4241-916c-2b76a9b8247d · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large language diffusion models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 272cceb3-fdc4-461b-bff8-8b12b6551c15 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2d55106d-d58e-4a28-934b-ad3fc2bedbe8 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba in speech: Towards an alternative to self-attention,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fb5446a7-e0b9-4e22-90fd-ab3eb3f0b433 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fb09a6e2-85e7-488a-aa59-e4d6fb97ec9f · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50ac2db-981b-41c4-8db5-b16c887c2a86 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avformer: Injecting vision into frozen speech models for zero-shot av-asr,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b793b600-57fd-4929-b18c-e0f972b4440a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gesture-aware zero-shot speech recognition for patients with language disorders,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2792da03-5baf-4951-a1fc-32dca1273452 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9ea407-1853-49bf-9f5d-8d534cac714e · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00be9c3e-7574-41a0-bc77-9ec09d650be3 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Robust speech recognition via large-scale weak supervision,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 47ed3191-fbb7-4a33-8c57-74b2f4f57603 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Perceptual score: What data modalities does your model perceive?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 68229392-882f-44df-8fb1-1c45af54fd8f · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? OPT: Open Pre-trained Transformer Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4445125b-d4d2-4759-b990-34cf2ae851af · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Librispeech: an asr corpus based on public domain audio books,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 94bb2653-48d2-44cb-9284-6e7eebf93853 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cvss corpus and massively multilingual speech-to- speech translation,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6654fe36-8e69-4982-97bd-87b4c9347b63 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408f3226-4033-4374-8f2e-e8539d49e997 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Microsoft coco: Common objects in context,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c239201b-f9af-417f-8de9-6f4f4bb0558e · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61914bb-bbae-4dcc-8c67-944074878096 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Learning audio-visual speech representation by masked multimodal cluster prediction,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d0de09d2-5e37-47ea-a87c-4aafa2b5146a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-shot text- to-image generation,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e84adce2-67ee-4889-b71d-f08a113f675e · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 34763197-be4c-459c-9712-e0ff39790d65 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Decoupled weight decay regularization,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9ad27258-a6d2-41f0-960b-2870186be976 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Deep audio- visual speech recognition,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4774422d-0fbb-4c12-87da-50305f889e5a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f1d409a-8add-4c12-93ca-11ffb967f691 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Spoken moments: Learning joint audio-visual representations from video descriptions,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8f5bff19-e950-4cb5-a898-7b14a40d218b · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? pyttsx3: Offline text to speech (tts) converter for python,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 502a7683-4b4f-4e05-b863-413ee6eca90a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MUSAN: A Music, Speech, and Noise Corpus
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89498a63-f6cb-43e0-b679-fece7763bde2 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Easyocr: Ready-to-use ocr with 80+ supported lan- guages,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 93f3b218-634d-478c-9940-26d54c6e24bd · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- moments in time: Learning and interpreting models for multi-action video understanding,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc02efc9-5ce5-420c-9be2-a3f62988bd0a · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Auto- avsr: Audio-visual speech recognition with automatic labels,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8c26f861-47c6-4525-9ff7-878f0c730255 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Flamingo: a visual language model for few-shot learning,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 89a83efc-6b20-4894-a2fa-2bca2d5eb5c8 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ead41f1-829e-4bb4-bb49-02985ed27086 · outbound
MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19493417-f083-4e90-ac6b-3d3122752e56 · inbound
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.