Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:50.951353Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2506.02012.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:50.951353Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:48.987998Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T13:26:51.425005Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b37aed34-d145-4be2-98d0-b26ab68836dc · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Typically, a VSR system takes a silent video containing the speaker’s lip movements as input and outputs the corresponding text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f16fa30a-8bec-4a7e-b1e9-821092d3c397 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b6b68f59-8a76-4fc6-b1c2-b6866386d3ee · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Data To evaluate the effectiveness of the proposed methods, we need to select an appropriate dataset that enables LLMs to demon- strate their potential
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 762b0e2f-eb5c-4727-8bdf-ccb35005817d · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Results on Scaling Test We first conduct a Scaling Test to investigate how the LLM scale affects VSR performance
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c3fdda2-e049-4577-b9b5-63b50531825a · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing By a comprehensive experimental study, several interesting conclu- sions can be drawn
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a90a3d3c-9de4-430d-acb4-60865026ff96 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing LipNet: End-to-End Sentence-level Lipreading
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0bc5c44-0d9d-4e87-8190-505f49c24f5a · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Large-scale visual speech recognition,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 70d4497e-80f7-4555-9aac-065e795b019b · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Auto-A VSR: Audio-visual speech recognition with automatic labels,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9e9b772-6ece-44e2-8493-5a65aa9f11f3 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Zero- shot fake video detection by audio-visual consistency,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8eb19c12-609b-4f9d-a03f-b4d2349dd87f · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing End-to-end audio-visual speech recognition with conformers,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da87d841-f1f5-4358-889f-125ba353a560 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Conformer: Convolution-augmented transformer for speech recognition,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8b2fb7-6e6b-4dcd-8f94-a7f8f59eeecc · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Audio-visual speech recognition with a hybrid CTC/attention architecture,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d56787f-5e08-4d60-b1db-07a5e87ece83 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Attention is all you need,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49acefc7-7e40-4a56-827e-894eac7a893f · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing SyncVSR: Data-efficient visual speech recognition with end-to-end cross- modal audio token synchronization,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 150bd5e0-5047-4913-b84f-9f45720ad619 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b40599-a2f9-453a-b2b0-10e4c6ee0d0a · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Video understanding with large language models: A survey,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0117793b-5de8-4a38-8bf8-8c9c07a7b2ee · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Audio-Visual LLM for Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db23810-0159-4806-8878-5b2ca4b1e425 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e6cc251-cb88-4cba-aca1-49a25c431ee7 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing In- teractive video search with multi-modal LLM video captioning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5341b810-45ce-41fc-b7fd-ce13ed146c14 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Improv- ing image captioning descriptiveness by ranking and LLM-based fusion,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed53de28-daf9-497b-8404-adcd6a9563a2 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing From Alt-text to real context: Revolutionizing image captioning using the potential of LLM,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3005850-b162-48f3-ae2f-f8932267cbe1 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Image captioning using multimodal LLMs,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0327afcf-c734-44be-94d0-24058c03960e · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Large Language Models are Strong Audio-Visual Speech Recognition Learners
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ab2c1c-d7f1-48a4-92cc-31733056dfad · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Where visual speech meets language: VSP-LLM framework for efficient and context- aware visual speech processing,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 00a6e8bb-a13c-4354-baf0-8e26871288e2 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Learning audio-visual speech representation by masked multimodal cluster prediction,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c891675-f42a-477c-8555-633857dc29f1 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing QLoRA: efficient finetuning of quantized LLMs,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 163a93d6-7ced-4505-9690-2c48bdb8f7ad · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing The Llama 3 Herd of Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc8b4bba-044c-49b5-84b1-eada9dd0c73f · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ce7bd1-b68d-4f4e-aaa7-1dca9bbc100b · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88951f03-dd06-4f25-ab6f-b33e1550af26 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing SALM: Speech- augmented language model with in-context learning for speech recognition and translation,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0a989c0-1396-4f1e-9cba-a0d4c16c8906 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing End-to-end speech recognition contextualization with large language models,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e71e487-a1dc-4eca-8346-8eb7ef13852a · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Eliciting the Priors of Large Language Models using Iterated In-Context Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9773ea2a-8ec3-432e-9950-a73c074c1dbe · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Deep audio-visual speech recognition,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82560c5f-b2c8-4036-90d3-a6af7d469a01 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing CNVSRC 2023: The first Chinese continuous visual speech recognition challenge,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e126a1f0-ed55-4bd6-8622-f923f61a0da1 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing CN-CVS: A Mandarin audio-visual dataset for large vocabulary continuous visual to speech synthesis,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2459821c-f360-47ef-bbe0-f41d26cc0bb9 · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Reti- naface: Single-shot multi-level face localisation in the wild,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0233da9d-c52d-48dc-a9b5-20952f3ad4de · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing How far are we from solving the 2D & 3D face alignment problem?(and a dataset of 230,000 3D facial landmarks),
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfc314a5-4cd6-4b34-8c8d-07d91929297d · outbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16fa30a-8bec-4a7e-b1e9-821092d3c397 · inbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.