Pith. sign in

Paper Citation Record · LEDGER

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2505.14336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14336 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:34.111978Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:28.786395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:39:34.749620Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb34c564-0255-47f0-a5af-47e66d1a49da · outbound

This paper cites Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:39:34.821428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:28.786395Z digest=sha256:828f2488a76e9b065512846c1e274d2bed18905c9c8daf82662e53ccced183c9

Observation 854669e9-b19b-4a3d-afb6-81d7d591a698 · outbound

This paper cites This is crucial in resource-constrained LLM-based A VSR systems, as we aim to improve performance despite using smaller-scale LLMs and pre-trained encoders.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach This is crucial in resource-constrained LLM-based A VSR systems, as we aim to improve performance despite using smaller-scale LLMs and pre-trained encoders

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:43.754098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:28.842189Z digest=sha256:b50cbec1ed9117596cd42107937f2e26c1c5f338f1dcc10b1a890d6417355878

Observation b7ab5303-54d5-4ca6-9a85-233cd69bf06c · outbound

This paper cites Transcribe{task prompt}to text.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Transcribe{task prompt}to text

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:39:43.622848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:28.959989Z digest=sha256:7d6a03e25f8eb90fe5d450079ab0d2e093d422df61d89e850166decd5a169ee1

Observation 4c0562e3-906b-4377-854b-4debf0c19cc5 · outbound

This paper cites Its key innovation is replacing the lin- ear projector with a Top-K sparse MoE module.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Its key innovation is replacing the lin- ear projector with a Top-K sparse MoE module

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:43.513781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.067190Z digest=sha256:3ec9634b86896436c6f2b754eb8db5a6bb9f8a37747b7a80cab8e1c998fb0d6b

Observation 48b325f5-22a0-4758-bcab-bcf4935c0bb3 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:29.177686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:29.177686Z digest=sha256:a39fb50eddeb24ef49497e93e98dfbc85534dcc3a7b9f260798435c13435e233

Observation 6af2e464-b6fa-45be-8ebc-79f4ad2c0280 · outbound

This paper cites Deep speech 2: End-to-end speech recognition in english and mandarin,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Deep speech 2: End-to-end speech recognition in english and mandarin,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:43.274160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.296827Z digest=sha256:8f7c83b70854be0ecee516460ee8f325796d0fb493f21720443bfab491bf2f25

Observation 77e79aba-3861-4ae5-a3ff-d9b6241e220a · outbound

This paper cites End-to-end speech recognition: A sur- vey,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach End-to-end speech recognition: A sur- vey,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:43.064274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.392445Z digest=sha256:be9926e86898b7cf0af19f0814dc53fd628543f9bc3a7bf4910f601800979689

Observation a7ebaeff-76fa-40d6-beee-51a715d9f196 · outbound

This paper cites Audio-visual speech modeling for continuous speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audio-visual speech modeling for continuous speech recognition,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:42.921045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.512583Z digest=sha256:fb8d2b8c62636b8a49ddb45fd0ed11b01088203084c6827abc3f2b9230fcce16

Observation 188010e7-f95a-4c3e-a6b5-424bef2cbb2e · outbound

This paper cites Investigation of speech separation as a front- end for noise robust speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Investigation of speech separation as a front- end for noise robust speech recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:42.815885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.617090Z digest=sha256:f714d38d9f2e73d6bcc70b77389a0144477a38def340c30e301d3a52ec3014bc

Observation 4af7a8d2-2900-4128-8e72-eb6ca82d3e0d · outbound

This paper cites Audio-visual speech recognition using deep learn- ing,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audio-visual speech recognition using deep learn- ing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:42.636962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.718054Z digest=sha256:195c96354abbef1cce10febf83e34b7dd3d81c8b4b8e60e83590e7f4cf012f47

Observation 35050d51-0909-49ee-bb22-e1d9811c4f07 · outbound

This paper cites Deep audio-visual speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Deep audio-visual speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:42.460290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.793596Z digest=sha256:e0e8147202679c9a5e4a7f4f996454c28c806d9c72d6bbf9736272e4dca91bdb

Observation 4162f35f-1e8f-4e3a-a19e-5cce765eb734 · outbound

This paper cites Audio-visual speech recognition with a hybrid ctc/attention architecture,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audio-visual speech recognition with a hybrid ctc/attention architecture,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:42.312711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.884046Z digest=sha256:1b211b72a13c14dd95a1939444f128bd37996b77636b9549837751cff8a35cf7

Observation cb370472-9c29-485b-a705-d25f88a2af39 · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach End-to-end audio-visual speech recognition with conformers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:42.134659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:29.960022Z digest=sha256:d95abbed02940fdff90621657f1d694f822bd9ece36749734dd533920e1476ff

Observation ebdfcbc3-9b9c-4648-bee3-ab9970c64991 · outbound

This paper cites Visual context-driven audio feature enhancement for robust end-to-end audio-visual speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Visual context-driven audio feature enhancement for robust end-to-end audio-visual speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:41.826894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.022183Z digest=sha256:96eb9de11633b48ad0b0fdf90e8b33ad6f0750bcb0d01e7866eaf4c01c385bd4

Observation 9f0b6873-112f-4b70-9cf0-599666223be1 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:41.559836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.099358Z digest=sha256:3b5f7040355be055ccd9984c6ad03b05bd394b8d3daee38d648c683a5f3da6a8

Observation 96890755-e6b6-4be8-b254-2881b9b49b1b · outbound

This paper cites Whisper-flamingo: Integrating visual features into whisper for audio-visual speech recognition and translation,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Whisper-flamingo: Integrating visual features into whisper for audio-visual speech recognition and translation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:41.292636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.179611Z digest=sha256:3532d56db563913425b3da897b03fadb9615849d40749b0611abbd361a823e0e

Observation d83c6aa0-e7a5-492c-a3e1-f8c869c0f0b9 · outbound

This paper cites A survey on self-supervised learning: Algorithms, applications, and future trends,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach A survey on self-supervised learning: Algorithms, applications, and future trends,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:40.975092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.265926Z digest=sha256:e5165aa79e9669407a7af046518b0552c23938bd72627fab39e5496ec74457b6

Observation 4242120d-d12d-4b0a-b055-bdd8319be364 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:40.725609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.329908Z digest=sha256:d1e1d6d38dfd740b50f6b0aba297b0206ae542631af4985296da1df90771edf6

Observation c333ce47-1df5-4dbe-9ef5-2ee24958b68e · outbound

This paper cites Jointly learning visual and auditory speech representations from raw data,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Jointly learning visual and auditory speech representations from raw data,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:40.495493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.413872Z digest=sha256:75bf69b9cbe347efccb902f41c7ef9f61f8f9f3031c5384019263302a64d507d

Observation 5c8e2a37-9a4f-4736-a29b-c404690f6afa · outbound

This paper cites u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:40.289862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.498623Z digest=sha256:f0a65e186c890916315801186c6872c5d60e2603d6dc126e9420acce234567d7

Observation 2bbe6460-a9ef-466d-a12c-3fded2e3cf02 · outbound

This paper cites Braven: Improving self-supervised pre- training for visual and auditory speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Braven: Improving self-supervised pre- training for visual and auditory speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:40.122153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.583391Z digest=sha256:028ef7f406aaab17511522d3ac2e0fdf804acc18ec7a284569835ee89dcf7d47

Observation a089d9ec-d62e-4726-ada1-d02a18e3969d · outbound

This paper cites Unified speech recognition: A single model for auditory, visual, and audiovisual inputs,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Unified speech recognition: A single model for auditory, visual, and audiovisual inputs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:39.932682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.647934Z digest=sha256:7eec13ca7b54face6feee655e019bc8c1e2c01f51f5317a16d2d0650a7d616d2

Observation feaadf66-b34a-4cb0-8f27-6aa4e0bab2d8 · outbound

This paper cites GPT-4 Technical Report.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:30.744252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:30.744252Z digest=sha256:82bc4c67ee119812b68ac2007fcc218114affdd7e33d032eefb94c44198304d5

Observation fec5a96e-b54c-437f-b06c-851942001bb4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:30.816401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:30.816401Z digest=sha256:70f6fe0f378f076c13b60c13bcf4169663eb97fb18003d9a07883ba1ce0a8778

Observation 9f279aed-e8f4-4d24-80a7-c207d0f18020 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Improved baselines with visual instruction tuning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:39.770707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:30.899234Z digest=sha256:06a4af3f8e3d4a5514d55e4c327b026f5b192ebd8c58bb53994ee76bf1fd6600

Observation d470f4f4-f794-4237-a9a3-c508819393ac · outbound

This paper cites On generative spoken language modeling from raw audio,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach On generative spoken language modeling from raw audio,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:30.987118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:30.987118Z digest=sha256:dbd797a9f224b2bda138304d61575049ac685ab5e5e55a56cb79a114fe7df2cc

Observation 72bd5ed8-d435-45e2-8f0d-ac21070670db · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:39.596788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.081916Z digest=sha256:265671dec70e01dc080569ff69fb160d707ee037dad780a44f27da9dc62a82eb

Observation 408ee907-d152-4960-835c-8966203d05aa · outbound

This paper cites Let’s go real talk: Spoken dialogue model for face- to-face conversation,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Let’s go real talk: Spoken dialogue model for face- to-face conversation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:39.428741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.143555Z digest=sha256:085906752576515582524ea81dbcf79956627a380c3d586d8dc862f7a403e013

Observation 8170e4d9-a3c6-43a0-8b53-310b3f5797dd · outbound

This paper cites Developing instruction-following speech language model without speech instruction-tuning data,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Developing instruction-following speech language model without speech instruction-tuning data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:39.229595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.222800Z digest=sha256:67dd1ff327ff0282e9f00029b8564e3e39054bf6f3478ee4e90175c3d9a55215

Observation b49faad6-117c-4d7a-94ac-ada9c3300905 · outbound

This paper cites SSR: Alignment-Aware Modality Connector for Speech Language Models.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach SSR: Alignment-Aware Modality Connector for Speech Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:31.292721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:31.292721Z digest=sha256:2f693f192a0970c2d70c49ef3c114eb74ce6388a0900f8257bb57fc4c9264bc6

Observation 7835cb45-c654-45f2-adb9-91e41bf36593 · outbound

This paper cites It’s never too late: Fusing acoustic information into large language models for automatic speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach It’s never too late: Fusing acoustic information into large language models for automatic speech recognition,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:39.011320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.368874Z digest=sha256:59a0cf0f6e093800e1d4444eccd1ea4aa3794f22dfe012eda12943574c92b977

Observation 3043acb5-fa44-4907-9105-98c275627aef · outbound

This paper cites Large language models are efficient learners of noise-robust speech recognition,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Large language models are efficient learners of noise-robust speech recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:38.784593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.451865Z digest=sha256:525a45f2c2569a4804355379592d9ba10d1000454a041603f0b37c88b8db715b

Observation fde77f27-3377-4510-a71b-9d197e0e108b · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:31.525701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:31.525701Z digest=sha256:d762b5162a2fbb518edf6ff8d0ce9a533f3eeaae52cc8275e35c5953c3da2e84

Observation 9f32b6a4-acba-4693-adaf-d1b547c9d30f · outbound

This paper cites Connecting speech encoder and large language model for asr,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Connecting speech encoder and large language model for asr,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:38.586326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.586155Z digest=sha256:06271e27951caf44631568c9d2c5db7c3bb24944b28a7a9d72486e63e73aa835

Observation b27c3acf-289d-4f57-b72d-08758ab950f9 · outbound

This paper cites Prompting large language models with speech recognition abilities,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Prompting large language models with speech recognition abilities,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:38.393062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.667864Z digest=sha256:359f605b2cfc19ff62f9b807d628e93d4f605c9c5bc15810cd7ad3d161e13459

Observation e8b5be17-de52-4402-9183-5d8945b7a122 · outbound

This paper cites Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech process- ing,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Where visual speech meets language: Vsp-llm framework for efficient and context-aware visual speech process- ing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:38.209014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.772230Z digest=sha256:9038f2b136a7dab320e0a3d88e92d7b4d640cdeb3b2463a428774961e21c6dd6

Observation 4bd81aca-8e3e-4140-8c0c-b766d7a78aea · outbound

This paper cites Large language models are strong audio- visual speech recognition learners,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Large language models are strong audio- visual speech recognition learners,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:37.984074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.860619Z digest=sha256:9edec4257e47d8998e831f109233a1bc3eccad89e04c2e1e2f9117d1bd232067

Observation 74c03ad7-df57-4179-97a5-233f70fed895 · outbound

This paper cites Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:34.600776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:31.946915Z digest=sha256:2ac11fef2f3c20e8cbb7e2759243732bfc3519cb5b950cb5a0efd13519eed9d9

Observation 0fded2c8-7c41-4daf-8b4e-16ea2954e66a · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:37.759019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.029232Z digest=sha256:d2f1bd208331d16be92fc560b8703c40f2c4cdd23cf00a4c1b1dbb67dedd7dbb

Observation f6b4946b-c23a-41ea-82be-6456a319d174 · outbound

This paper cites Gshard: Scaling giant models with condi- tional computation and automatic sharding,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Gshard: Scaling giant models with condi- tional computation and automatic sharding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:37.572919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.102664Z digest=sha256:b1306e5593cf755f4990f5376c66fe2f9dbae5263e4016371352446c036ab547

Observation ee17e29d-dfee-4a80-9d27-a387a5a7a269 · outbound

This paper cites Efficient fine-tuning of audio spectrogram transformers via soft mixture of adapters,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Efficient fine-tuning of audio spectrogram transformers via soft mixture of adapters,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:37.370592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.197118Z digest=sha256:8715a442d58d58bcac512b60b580ea42ac5593641f8202542077bc16cd72f5fc

Observation ba451649-cbec-430b-9c6f-ea5250ed1c15 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:32.285143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:32.285143Z digest=sha256:b54c0e73806f3265e54fcce8b95747148104d03bc42714aab12164c1d2a0f349

Observation c6287b0f-f201-490d-b5b4-5e8dc12f5861 · outbound

This paper cites Mixture of A Million Experts.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Mixture of A Million Experts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:32.373581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:32.373581Z digest=sha256:da8d6f321866118c9a08a04fbf8b36eb784fb138d545b2320c9837ac4cb0a4d9

Observation 6e2b32c4-466f-4a0d-a84f-f9d077ad07d9 · outbound

This paper cites Olmoe: Open mixture-of-experts lan- guage models,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Olmoe: Open mixture-of-experts lan- guage models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:37.161671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.488903Z digest=sha256:131a7f3a6263d804d17c9432b23025d0e50a28d00eb68ccd6617b4cc8950eb07

Observation 160b6d9e-fe8d-43f6-85ae-7b74809bd176 · outbound

This paper cites Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:36.979109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.629949Z digest=sha256:a5c68d429cff959b83fe64809a1c90e1987106e806cf8276728644fc9198ac47

Observation e513e606-4937-4e6c-a932-ad186f76fced · outbound

This paper cites Chartmoe: Mixture of expert connector for advanced chart understanding,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Chartmoe: Mixture of expert connector for advanced chart understanding,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:36.730999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.719550Z digest=sha256:bba6b0614cccafc01fc5b28ff971531eb4a1252d9ca786b82a4feff7fc4fc117

Observation 34b72aa8-89c5-4db9-b9c7-da0fbd3ffdd5 · outbound

This paper cites Dense connector for mllms,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Dense connector for mllms,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:36.560704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.804638Z digest=sha256:468303620a67300027e069707bb59451d0b7dcb4cb4e6de8549c11f14f94a347

Observation 1884dd08-84ba-4140-bcec-d54b4ef10f95 · outbound

This paper cites Visual instruction tuning,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Visual instruction tuning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:36.381234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:32.910843Z digest=sha256:e5a111c6df66c130d43f2df0dfbc508b10ee786d100338b007507fd850a1cd0c

Observation 5b7540b4-6bb1-47d6-903e-38d789347c5d · outbound

This paper cites Vila: On pre-training for visual language models,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Vila: On pre-training for visual language models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:36.213955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:33.015404Z digest=sha256:aff5ae649f7242e44a7162f88ffda1eda0a379f199559b9012310e1fce3ce059

Observation c8812366-15d3-45be-8a37-11d69eec5b6e · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.091881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.091881Z digest=sha256:45a33d2c10abe5024542387dd97733d042ddb289a24d67d597b5cc06a338d772

Observation dbefa5de-ac70-46fe-8edc-5698f6e910c5 · outbound

This paper cites Meteor: Mamba-based traversal of rationale for large language and vision models,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Meteor: Mamba-based traversal of rationale for large language and vision models,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:35.958993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:33.156800Z digest=sha256:37022b2d56616a7369b113d3ff925240d6427a332fb43f00439db8eba00b8997

Observation a04070d8-7a4a-48f0-aaea-7ca62d1dd6ca · outbound

This paper cites Lora: Low-rank adaptation of large language mod- els,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Lora: Low-rank adaptation of large language mod- els,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:35.751900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:33.253776Z digest=sha256:1e8803aba451a1b33713ad42259e3c6bdcfe8e033b766420b7bbfab712498554

Observation c76be01e-d01f-4fa6-9cc7-8aa6b74ce86f · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.358907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.358907Z digest=sha256:ad905aedf62891c9d08f7c98226d323f3089cfe6a51a83d5ec66f1d11d0e8550

Observation 4a45b99a-6465-4660-92f8-72b2677fe431 · outbound

This paper cites MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.418185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.418185Z digest=sha256:253e5f76b218580cecc81f9b78a856b069c323c8f246d58afe16498f41608d80

Observation aae50937-c189-4c57-afcb-9d96beb8aa9e · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LRS3-TED: a large-scale dataset for visual speech recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.522351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.522351Z digest=sha256:65f80db615bc77630aee93d83126d741daaaebfd791d3c4192085da6dd74082e

Observation 26a5e078-3f87-49dc-bd75-d48bb2d935d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Robust speech recognition via large-scale weak supervision,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:35.548830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:33.634650Z digest=sha256:b1bb77056004d68a1c8ed232cedc894991338ff1a71c88605cc4b3e5ae6923b6

Observation cf986061-082d-400a-9f3a-ca53fa4dd693 · outbound

This paper cites The Llama 3 Herd of Models.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach The Llama 3 Herd of Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.700092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.700092Z digest=sha256:06895f2cb9f944112dd02681c95cd4b93b305f04b1d400223865cb52d7aabbd5

Observation 6916c243-6717-40fb-86ef-3cb2febef3aa · outbound

This paper cites Towards a unified view of parameter-efficient transfer learning,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Towards a unified view of parameter-efficient transfer learning,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:35.376169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:33.787501Z digest=sha256:b1f81e95281a5d85a67e430ca0d6ae9f2f24e731c1c229399cd2b1cf4c75d3cb

Observation 938640d8-9e79-4b54-b176-dde2deed5f17 · outbound

This paper cites Parameter-efficient transfer learning of au- dio spectrogram transformers,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Parameter-efficient transfer learning of au- dio spectrogram transformers,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:35.200745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:33.881568Z digest=sha256:828e8b48317cd25fa6579dbabdd1d0132b886c2f8569fcad50f1c9de0c344130

Observation c07ba69b-84a4-4031-a186-ef95e4cb43fe · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.966621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.966621Z digest=sha256:8c90e83a413377c9621eabeae11171f7191f43eae3323e97ceb91f07a3c57fe7

Observation 088c4cf1-edc1-41ea-9160-20f5dfcdbfe9 · outbound

This paper cites Mixtures of experts for audio-visual learning,.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Mixtures of experts for audio-visual learning,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:39:34.993972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:34.034276Z digest=sha256:ca388c2df631144c393596c276a659c6a176fd572eea3f2f6f3437fb4eb29839

Observation 10189419-0fe1-4ad1-9f36-c209cea88af3 · outbound

This paper cites MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:34.345367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:34.111978Z digest=sha256:61254361f5f90343ed1f0613c5df3a29ed1133f2745cc6697c7b5360fb703339

Pith citing papers

Observation eb34c564-0255-47f0-a5af-47e66d1a49da · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:39:34.821428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:39:28.786395Z digest=sha256:828f2488a76e9b065512846c1e274d2bed18905c9c8daf82662e53ccced183c9