Pith. sign in

Paper Citation Record · LEDGER

MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

As of 20 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19037 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:08:19.017587Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:45:51.528645Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T06:11:01.750278Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3e573cf-65f3-49a9-b1e5-55b9c13ca43f · outbound

This paper cites Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- modal speech transformer decoders: When do multiple modalities improve accuracy?

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-15T18:08:19.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.766507Z digest=sha256:538f45beab20fe4b72afde1193241f5a7f556631057588498fb91e3bd52063a3

Observation f90e351b-e603-4ff9-9c94-b1c93912ee2a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.771390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.771390Z digest=sha256:9513f6d5ac2a45c9e53c29403746db65a7489f494326fee5800359fc7129f617

Observation 21ea53f2-f7e7-40b5-b809-13cb0a4817f0 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Next-gpt: Any-to-any multimodal llm,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.826497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.775950Z digest=sha256:bb4161165ded80e99cc9aae3f5c57eb4b5ae026d1a39c66afac7b9608c73def3

Observation b42e5b2e-62da-443f-88ee-1b53fa616297 · outbound

This paper cites Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.780134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.780134Z digest=sha256:721e5f294e974f474e2d03c767fb8ecca2f16d8875311eb5fdead23b5059396e

Observation 3f8cdab4-09ad-4202-a4bb-e82f7aad9105 · outbound

This paper cites Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:08:19.194368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.784896Z digest=sha256:f69460d74c533a3a464085825eda7c62567d233f2f438942a1f4732cba3a83f9

Observation 1e7b3b71-36b7-492f-8ccf-44d5c78c4b62 · outbound

This paper cites Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Exploring speech recognition, translation, and understanding with discret e speech units: A comparative study,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.814059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.789458Z digest=sha256:4b3768f764521284180bd80e3e74bd3cf588b32cd1b9520ccd6cb3bf0ccaffd6

Observation a0b88f1b-85f6-4be0-b476-0c7d23eddc74 · outbound

This paper cites Speech recognition meets large language model: Benchmarking, models, and exploration,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speech recognition meets large language model: Benchmarking, models, and exploration,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.801207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.794719Z digest=sha256:90ff920f77dc9eb219007b6c0c3b989d1685646fca5c6e4c0cb4db01abe6162d

Observation 678efdf9-8030-46c2-8235-4433d439abe0 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.786852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.799622Z digest=sha256:1c9a283739189400b4e0bc2ac91ffe41bed271947166830da37447794e69e917

Observation c7dd4b52-4a83-4b33-bfb7-190f8bcc8de3 · outbound

This paper cites A comprehensive review of multimodal large language models: Performance and challenges across different tasks,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? A comprehensive review of multimodal large language models: Performance and challenges across different tasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.774604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.803629Z digest=sha256:8f56c8c48b41282c2ac0dc876d3da973f3ca455f857aa85b51ce2923c9cd9468

Observation 913df5e4-5cd2-486e-8865-3de22a8503af · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.807653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.807653Z digest=sha256:d6f71fd4c00ac1fb99fbcc7ad1094f46ce642df5ab7d43331fba314b9763bfc5

Observation d24fb121-cf20-4d97-9c81-0863c290eca7 · outbound

This paper cites MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.812315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.812315Z digest=sha256:429e0c5f8782768f29e16dcf4001a4e47aca008b13621ce1668f088cb28a406d

Observation 07b3c3f9-e185-4528-8319-63a29e8699c4 · outbound

This paper cites Large lan- guage models are strong audio-visual speech recognition learners,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are strong audio-visual speech recognition learners,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.760517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.816674Z digest=sha256:fddcc2fb1a1b99a5945bf96791cc8ce9cc5cc8b9cd9df24966bf30e8bd4ce6cc

Observation 0d0b914e-cf80-4843-919d-bcef5a4926cf · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Watch or listen: Robust audio-visual speech recognition with visua l corruption modeling and reliability scoring,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.747066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.820439Z digest=sha256:e3672045ef43ff94a80c0ae340797c8f57125bfc6e48c922f14e088c0fa62ea1

Observation d1e39b58-d6a1-443a-9657-d0a941bdac5f · outbound

This paper cites Large lan- guage models are efficient learners of noise-robust speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large lan- guage models are efficient learners of noise-robust speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.733800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.824792Z digest=sha256:a51e2e5679a73faec928258b1f439c2fb88abb08da12be69ae206ec66a7ede7e

Observation 07648979-b6d1-4176-bf51-4eba9e608b8a · outbound

This paper cites Avatar: Un- constrained audiovisual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avatar: Un- constrained audiovisual speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.720447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.829101Z digest=sha256:11ebf8e4b0bc466d942574bc005f61a401e15a42b169f3c0d0c6cc7c5a7cddd6

Observation a779a849-e73d-4472-b0b9-d18e1f4b451c · outbound

This paper cites Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mmger: Multi-modal and multi-granularity generative error correction with ll m for joint accent and speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.707244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.833367Z digest=sha256:13a5d423837bbb8c0c6032b16030bd538c42fac305806c767d07fb6d1ac07ad2

Observation 6cbf5381-b990-4157-ab70-8e85debd40f0 · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba: Linear-time sequence mod- eling with selective state spaces,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.694982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.837512Z digest=sha256:f62bf0a520c6af8f076c587c7b71341cd6c3251054f171453c0b44dbc752a02e

Observation 9454dc90-f285-4674-9527-e0c0beed982c · outbound

This paper cites Improved baselines with visual instruction tuning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Improved baselines with visual instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.841709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.841709Z digest=sha256:cbf71efdf2922d0dc4498c205ba744df5e9c684418011c4da9988a6c38af4055

Observation d51339a1-297b-4b06-8b8c-81d0a4500bdf · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.845694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.845694Z digest=sha256:ad8b0e528e94320f40a066880c77c7debdf44d86b59625b24e8c65e871aaee6b

Observation d81f9b08-7c84-4b94-a780-d72efe7c4c06 · outbound

This paper cites Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mm-interleaved: Interleaved image-text generative modeling via multi- modal feature synchronizer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.674964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.850564Z digest=sha256:a02301bdb71204ad7a61815363094e44a5673465bac4f6ea3983c63b9d52676a

Observation aaf9ea8d-38c8-43ae-bf8b-257ce7e9322b · outbound

This paper cites Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Transformers are ssms: Generalized IEEE TRANSACTIONS ON MULTIMEDIA, VOL. XXX, AUGUST 2021 10 models and efficient algorithms through structured state space duality,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.660556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.856094Z digest=sha256:61442c9099c1c762c4e27816479b4608da4687ef9753d1e8d4121f8cc0271796

Observation 6587fa31-2a7e-4651-b08e-900b35ef08b0 · outbound

This paper cites Cobra: Extending mamba to multi-modal large language model for efficient inference,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cobra: Extending mamba to multi-modal large language model for efficient inference,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.648770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.860457Z digest=sha256:2aed7d80c6ca844f75b7d71c734c1aca6c223c7fd7cac59382425c3e62944cef

Observation 10d2330a-fdf2-48ab-a2a9-e37e4f767d12 · outbound

This paper cites U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.864679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.864679Z digest=sha256:45613da4cbe9fbd983be18f80640a61138218b24659f8203659bdeec13f4ba48

Observation b7637d7e-45e8-4168-a67f-e961fdc2d588 · outbound

This paper cites Vl-mamba: Exploring state space models for multimodal learning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Vl-mamba: Exploring state space models for multimodal learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.637526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.868502Z digest=sha256:ba9e4351b4700a79e1ed2c7c78fc82251f11cccb06566c97c4d75d1d07d32c0c

Observation c1523190-95e2-410a-8d53-10aacc321d34 · outbound

This paper cites Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Foundations & trends in multimodal machine learning: Principles, chal- lenges, and open questions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.625715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.872168Z digest=sha256:9aa37d9affb2d963dbd00d9a0856f3347ebcc8caa8b03c6b7b29c62d23cf0d34

Observation 3636cc29-cd3c-4e8c-94e0-57b2dabd6e74 · outbound

This paper cites Diffusion- lm improves controllable text generation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Diffusion- lm improves controllable text generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.612881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.885465Z digest=sha256:e0a6ce79e6d8d9f1b032cd0d33917fd6f92ce420f6ab217c8700e12446df5309

Observation 5be19410-7fcc-4241-916c-2b76a9b8247d · outbound

This paper cites Large language diffusion models,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Large language diffusion models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.600420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.890194Z digest=sha256:51a7239f8c666ecb3ef77402f42132b6c06a3e98706f3d28ae05a2b164f1f6de

Observation 272cceb3-fdc4-461b-bff8-8b12b6551c15 · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Speechgpt: Empow- ering large language models with intrinsic cross-modal conversational abilities,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.589094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.895150Z digest=sha256:d96e0a5ca77ca9ab72acb0771bfe6409d447137a701cca198de07d92c2ad2a07

Observation 2d55106d-d58e-4a28-934b-ad3fc2bedbe8 · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Mamba in speech: Towards an alternative to self-attention,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.575908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.899530Z digest=sha256:4f533313214f7c7572cfd85bd30fee052451f0ea0f5f0654441c3887aaa11542

Observation fb5446a7-e0b9-4e22-90fd-ab3eb3f0b433 · outbound

This paper cites Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.564962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.903286Z digest=sha256:21fe13ae3ed94b997bf8d90a44e706155889f0bd672a16b0ab17040bfbbf985e

Observation fb09a6e2-85e7-488a-aa59-e4d6fb97ec9f · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.907558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.907558Z digest=sha256:5b62ba7f12b9b5204e7ed8800b7625e7d3a70e283f2318796434f60c2c91411d

Observation b50ac2db-981b-41c4-8db5-b16c887c2a86 · outbound

This paper cites Avformer: Injecting vision into frozen speech models for zero-shot av-asr,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Avformer: Injecting vision into frozen speech models for zero-shot av-asr,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.553143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.911702Z digest=sha256:e22dd2f6b65200422b7a62fa2970fa11327b74cdac6f80f0d9242b8c507e60f7

Observation b793b600-57fd-4929-b18c-e0f972b4440a · outbound

This paper cites Gesture-aware zero-shot speech recognition for patients with language disorders,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Gesture-aware zero-shot speech recognition for patients with language disorders,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.541550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.916276Z digest=sha256:1012e70a726e737614828feb5bd58352bf85aa807bbe6949b63204e027c97526

Observation 2792da03-5baf-4951-a1fc-32dca1273452 · outbound

This paper cites Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.920776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.920776Z digest=sha256:7d4750c865502a6ad89903d6b457c2bd852b5111bfefb6cca92a39d4082a7347

Observation fc9ea407-1853-49bf-9f5d-8d534cac714e · outbound

This paper cites Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.924921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.924921Z digest=sha256:38b6fa8e702579bbaef0ed0791357fb8c5554c418375e48ce5abca2e69d9190d

Observation 00be9c3e-7574-41a0-bc77-9ec09d650be3 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Robust speech recognition via large-scale weak supervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.529812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.929951Z digest=sha256:1b626b8e0e7de28ec52548173ddd7e69dad7b2bd7b7f087ffe52632d11fce439

Observation 47ed3191-fbb7-4a33-8c57-74b2f4f57603 · outbound

This paper cites Perceptual score: What data modalities does your model perceive?.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Perceptual score: What data modalities does your model perceive?

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.518476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.933403Z digest=sha256:fbcca2b03e95634f9ffc5a4d8b319a6ea3da8b69e409a4b50cb839321392a5c3

Observation 68229392-882f-44df-8fb1-1c45af54fd8f · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? OPT: Open Pre-trained Transformer Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.936763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.936763Z digest=sha256:1e2c5221cc6218e16b77505eacd20b70e1b2020d7264caa32e8e68c5b2882a3d

Observation 4445125b-d4d2-4759-b990-34cf2ae851af · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Librispeech: an asr corpus based on public domain audio books,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.507533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.940735Z digest=sha256:b528233c04b18c839432df03a4a89cfdde0f694d6808adeb98bba4505cab675c

Observation 94bb2653-48d2-44cb-9284-6e7eebf93853 · outbound

This paper cites Cvss corpus and massively multilingual speech-to- speech translation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Cvss corpus and massively multilingual speech-to- speech translation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.496644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.945358Z digest=sha256:6c6a37b0c65336ce3bf37afe0e4ac7bb5ffc8f579328c4df3ca7434ca01ed542

Observation 6654fe36-8e69-4982-97bd-87b4c9347b63 · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.949250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.949250Z digest=sha256:bfc00585c76e00134a69038d3387f1a56f4aa9ec86a9ce39c5014747325b25a0

Observation 408f3226-4033-4374-8f2e-e8539d49e997 · outbound

This paper cites Microsoft coco: Common objects in context,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Microsoft coco: Common objects in context,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.485293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.953240Z digest=sha256:29759afff5ce3d8ef7ff129506b5fc6d3ad3bd234148ea590da64037e726eb4d

Observation c239201b-f9af-417f-8de9-6f4f4bb0558e · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.957514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.957514Z digest=sha256:76cebfd28c9bfe6891837a3ea09745846b6321cac406c3c7bf6988565ea82720

Observation f61914bb-bbae-4dcc-8c67-944074878096 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.472960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.961349Z digest=sha256:68daa7308734999b7642cd2727df53dbc7c9ce20dca3c27c91af83030f44c51f

Observation d0de09d2-5e37-47ea-a87c-4aafa2b5146a · outbound

This paper cites Zero-shot text- to-image generation,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Zero-shot text- to-image generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.461940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.965387Z digest=sha256:3b4fc61756ddab5b8b2fe8098a14e835e0e209d68aa1fc556173f22251f4f269

Observation e84adce2-67ee-4889-b71d-f08a113f675e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.450637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.969045Z digest=sha256:4e363c31f87b1116bb532fa723d884b8802a36a85163a8167608c1394bf7bd1f

Observation 34763197-be4c-459c-9712-e0ff39790d65 · outbound

This paper cites Decoupled weight decay regularization,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Decoupled weight decay regularization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.437858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.972445Z digest=sha256:6ac1d0333bf5057ba6a2f71ca3f7acee966e7ddd3f69dfe2bb5a7595a636bb1a

Observation 9ad27258-a6d2-41f0-960b-2870186be976 · outbound

This paper cites Deep audio- visual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Deep audio- visual speech recognition,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.426624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.976171Z digest=sha256:a6fd32e0185e7a4226d394c00e68609024588a9d8fadc3c21e7cc3944bc3ffe5

Observation 4774422d-0fbb-4c12-87da-50305f889e5a · outbound

This paper cites Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Slideavsr: A dataset of paper explanation videos for audio-visual speech recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.414781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.980618Z digest=sha256:45560f39523b63ab20d90484a9b2a9b0bc3f45a9458bf42ee5f9da24939da233

Observation 4f1d409a-8add-4c12-93ca-11ffb967f691 · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Spoken moments: Learning joint audio-visual representations from video descriptions,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.402749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.984832Z digest=sha256:cc8d65b9f20e07f7e42d61552f071e924d3c944b04df62e75130aa8f816c0976

Observation 8f5bff19-e950-4cb5-a898-7b14a40d218b · outbound

This paper cites pyttsx3: Offline text to speech (tts) converter for python,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? pyttsx3: Offline text to speech (tts) converter for python,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.390891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.988723Z digest=sha256:417a083aee5203dbc9927b8f4e713573bc7a78fa18f3ef66071a796ad9486182

Observation 502a7683-4b4f-4e05-b863-413ee6eca90a · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? MUSAN: A Music, Speech, and Noise Corpus

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:18.992534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:18.992534Z digest=sha256:2c4a2200554d99b98a3acc28eb4917cc17102a6ea340f68970f45feeda528877

Observation 89498a63-f6cb-43e0-b679-fece7763bde2 · outbound

This paper cites Easyocr: Ready-to-use ocr with 80+ supported lan- guages,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Easyocr: Ready-to-use ocr with 80+ supported lan- guages,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.377923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:18.996449Z digest=sha256:a7712828d2544c6ad97b01fa122267befa013b0a2515c1efccc7a1442fdfa52d

Observation 93f3b218-634d-478c-9940-26d54c6e24bd · outbound

This paper cites Multi- moments in time: Learning and interpreting models for multi-action video understanding,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Multi- moments in time: Learning and interpreting models for multi-action video understanding,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.366134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:19.000266Z digest=sha256:8e5134cccc683c17e479dd3b1281936dd398edb00294cbb48dab3527f91863cf

Observation cc02efc9-5ce5-420c-9be2-a3f62988bd0a · outbound

This paper cites Auto- avsr: Audio-visual speech recognition with automatic labels,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Auto- avsr: Audio-visual speech recognition with automatic labels,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.353708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:19.005817Z digest=sha256:a6e83639c71c3bc08634c463374f97c1d09671dda00681eef26294ad715a20e9

Observation 8c26f861-47c6-4525-9ff7-878f0c730255 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Flamingo: a visual language model for few-shot learning,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.341067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:19.009408Z digest=sha256:0b646c096bda0af0f417ddd1863eff2ed888c94db9e36d302cab5f8e87852585

Observation 89a83efc-6b20-4894-a2fa-2bca2d5eb5c8 · outbound

This paper cites Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:19.013669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:19.013669Z digest=sha256:7657fd19a706358b890ecf04923396ef24f08cca9f7908cf9535c7a1c22b33ff

Observation 6ead41f1-829e-4bb4-bb49-02985ed27086 · outbound

This paper cites Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,.

MLLM-based Speech Recognition: When and How is Multimodality Beneficial? Image first or text first? optimising the sequencing of modalities in large language model prompting and reasoning tasks,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:08:19.329115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:08:19.017587Z digest=sha256:8c3bb0e03168d0469a1096078edc2a8fc6a9e4ac6a831f213064aa60137908f6

Pith citing papers

Observation 19493417-f083-4e90-ac6b-3d3122752e56 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-10T01:09:09.297559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:f1476ea3b7b04acef5ef926b30669db488802a6a50d357000e5b97bcae762f4b