Pith. sign in

Paper Citation Record · LEDGER

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2506.02012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02012 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:50.951353Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:48.987998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:26:51.425005Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b37aed34-d145-4be2-98d0-b26ab68836dc · outbound

This paper cites Typically, a VSR system takes a silent video containing the speaker’s lip movements as input and outputs the corresponding text.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Typically, a VSR system takes a silent video containing the speaker’s lip movements as input and outputs the corresponding text

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:54.549002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:48.891052Z digest=sha256:7cf9791ddbdccc3d1252d25e6be1536c2fbfdc173d523755ea4d3549c50fe844

Observation f16fa30a-8bec-4a7e-b1e9-821092d3c397 · outbound

This paper cites Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:26:51.504112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:48.987998Z digest=sha256:738cfb42d103d8aafbfb5980aaf2389a4d3dcf4bd72003a8c210ea9d6bc958c9

Observation b6b68f59-8a76-4fc6-b1c2-b6866386d3ee · outbound

This paper cites Data To evaluate the effectiveness of the proposed methods, we need to select an appropriate dataset that enables LLMs to demon- strate their potential.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Data To evaluate the effectiveness of the proposed methods, we need to select an appropriate dataset that enables LLMs to demon- strate their potential

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:54.446779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.122743Z digest=sha256:5b72be1bbc4a139e7abae9cad520e0171ecfe82711203ddf786a9b6b50bd5aa4

Observation 762b0e2f-eb5c-4727-8bdf-ccb35005817d · outbound

This paper cites Results on Scaling Test We first conduct a Scaling Test to investigate how the LLM scale affects VSR performance.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Results on Scaling Test We first conduct a Scaling Test to investigate how the LLM scale affects VSR performance

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:54.347995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.219210Z digest=sha256:a79109375a0ea37abaaec4b40898fa7ab6d38e71d30787611a00b271be12b229

Observation 5c3fdda2-e049-4577-b9b5-63b50531825a · outbound

This paper cites By a comprehensive experimental study, several interesting conclu- sions can be drawn.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing By a comprehensive experimental study, several interesting conclu- sions can be drawn

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:54.227348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.339547Z digest=sha256:2c78e341b8e2c22c4c21c29821b806d8792778b14d6c47c42995371ef766ed00

Observation a90a3d3c-9de4-430d-acb4-60865026ff96 · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing LipNet: End-to-End Sentence-level Lipreading

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:49.462849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:49.462849Z digest=sha256:8278e5166d9fb83933512478a12b6a2523d3341014b4cfc12a350731da6d4940

Observation c0bc5c44-0d9d-4e87-8190-505f49c24f5a · outbound

This paper cites Large-scale visual speech recognition,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Large-scale visual speech recognition,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:54.105230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.577502Z digest=sha256:3de9750cc1c415c48cf78bdf18bc112f783abe903746b0e64b369698c301dd3f

Observation 70d4497e-80f7-4555-9aac-065e795b019b · outbound

This paper cites Auto-A VSR: Audio-visual speech recognition with automatic labels,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Auto-A VSR: Audio-visual speech recognition with automatic labels,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:53.906245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.710060Z digest=sha256:33997f90a289cc9fa31b2cbb78c5c69b2ac8edea9738287d3582cc8c06fc055d

Observation a9e9b772-6ece-44e2-8493-5a65aa9f11f3 · outbound

This paper cites Zero- shot fake video detection by audio-visual consistency,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Zero- shot fake video detection by audio-visual consistency,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:53.747874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.801766Z digest=sha256:ea5f08aeac37906f07a3c15d1494148ba27018673b3eceb2e5b9b67d785cb696

Observation 8eb19c12-609b-4f9d-a03f-b4d2349dd87f · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing End-to-end audio-visual speech recognition with conformers,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:53.530159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.830794Z digest=sha256:bbfd14571acd83adf8924b6d43ca3869da5defa9af2f0155ce155c345d843d5d

Observation da87d841-f1f5-4358-889f-125ba353a560 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Conformer: Convolution-augmented transformer for speech recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:49.868685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:49.868685Z digest=sha256:5e2a9fa3038d052465c61cd6a54639b141a3cf45fd78f216ce6aa56d93ad6bfa

Observation 6b8b2fb7-6e6b-4dcd-8f94-a7f8f59eeecc · outbound

This paper cites Audio-visual speech recognition with a hybrid CTC/attention architecture,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Audio-visual speech recognition with a hybrid CTC/attention architecture,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:53.368223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:49.912997Z digest=sha256:0d5a805b581b1db58e9e5aef24f44793bcbdeda68a0a1f0bea3fad28c1d6e0fb

Observation 3d56787f-5e08-4d60-b1db-07a5e87ece83 · outbound

This paper cites Attention is all you need,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Attention is all you need,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:49.952570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:49.952570Z digest=sha256:cfd485a3d4e503042154d83df349a1de5d59fd653e9bf36a21d4756af6c33584

Observation 49acefc7-7e40-4a56-827e-894eac7a893f · outbound

This paper cites SyncVSR: Data-efficient visual speech recognition with end-to-end cross- modal audio token synchronization,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing SyncVSR: Data-efficient visual speech recognition with end-to-end cross- modal audio token synchronization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:53.191811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.008690Z digest=sha256:43e4e3601263f8d827221547227af75964d31b6cdea78990ae6ae1d6b40ed080

Observation 150bd5e0-5047-4913-b84f-9f45720ad619 · outbound

This paper cites AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.042373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.042373Z digest=sha256:a9bdf446cccb78fd1e07cb771dbaf92a02e61d6ae0e362d3b6ce5e1637bd2e45

Observation f9b40599-a2f9-453a-b2b0-10e4c6ee0d0a · outbound

This paper cites Video understanding with large language models: A survey,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Video understanding with large language models: A survey,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.076082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.076082Z digest=sha256:17a674d31ed3ee36d6b700114a972b834390ef64c5127be9e95ba1daaa459dac

Observation 0117793b-5de8-4a38-8bf8-8c9c07a7b2ee · outbound

This paper cites Audio-Visual LLM for Video Understanding.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Audio-Visual LLM for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.104456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.104456Z digest=sha256:951c64ef5b5805ac0433e343fa5ff3b1b0f736fff9509f810f21b2beb8a1ac81

Observation 5db23810-0159-4806-8878-5b2ca4b1e425 · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.136607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.136607Z digest=sha256:fce0d24f1fe84992ba42e596625e2c7999f93e7e82c90f21e780a77fd182cd4b

Observation 5e6cc251-cb88-4cba-aca1-49a25c431ee7 · outbound

This paper cites In- teractive video search with multi-modal LLM video captioning,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing In- teractive video search with multi-modal LLM video captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:53.084681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.164958Z digest=sha256:b07a5b94a03113f4a181e09b94492737c98721a56e97a5ba166a0d8b8b65ce5d

Observation 5341b810-45ce-41fc-b7fd-ce13ed146c14 · outbound

This paper cites Improv- ing image captioning descriptiveness by ranking and LLM-based fusion,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Improv- ing image captioning descriptiveness by ranking and LLM-based fusion,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.215701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.215701Z digest=sha256:2b2876465eadfb52659442023e9c6907cf0966a71c4a15f51cdbdf644c20c928

Observation ed53de28-daf9-497b-8404-adcd6a9563a2 · outbound

This paper cites From Alt-text to real context: Revolutionizing image captioning using the potential of LLM,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing From Alt-text to real context: Revolutionizing image captioning using the potential of LLM,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.962299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.249129Z digest=sha256:6b4fa974f7d7d474b839cb0e40ff62c6d2fe659c97d99b4f4c301567a0fc932a

Observation b3005850-b162-48f3-ae2f-f8932267cbe1 · outbound

This paper cites Image captioning using multimodal LLMs,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Image captioning using multimodal LLMs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.867854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.285952Z digest=sha256:569dce9cbab0df119af5b2703bf1cadc78a06246aebceb237e92641aa337e7e0

Observation 0327afcf-c734-44be-94d0-24058c03960e · outbound

This paper cites Large Language Models are Strong Audio-Visual Speech Recognition Learners.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Large Language Models are Strong Audio-Visual Speech Recognition Learners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.325971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.325971Z digest=sha256:de4e25581934d16141780c8f11ff1177981f23684f965ccea87e561a09112df0

Observation 82ab2c1c-d7f1-48a4-92cc-31733056dfad · outbound

This paper cites Where visual speech meets language: VSP-LLM framework for efficient and context- aware visual speech processing,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Where visual speech meets language: VSP-LLM framework for efficient and context- aware visual speech processing,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.779608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.365046Z digest=sha256:ebdff8d27af4fe36471f4d1bef71b44fffcaf8208eb312c871714f27c51fe942

Observation 00a6e8bb-a13c-4354-baf0-8e26871288e2 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.690197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.399320Z digest=sha256:d018b828bfcddd5d9647d6646f2121ec7dcc0468007701efbf2f6d37fd13b7c1

Observation 6c891675-f42a-477c-8555-633857dc29f1 · outbound

This paper cites QLoRA: efficient finetuning of quantized LLMs,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing QLoRA: efficient finetuning of quantized LLMs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.588466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.444681Z digest=sha256:236c05a32adafaf1895ee6be79e9988fdf3b2a1607c6771ab519d6fdc971d883

Observation 163a93d6-7ced-4505-9690-2c48bdb8f7ad · outbound

This paper cites The Llama 3 Herd of Models.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.497376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.497376Z digest=sha256:a3ad0c892bbb0ad445610cc66719a0084e327fde26d776689a7a37c8d0c8fe2b

Observation bc8b4bba-044c-49b5-84b1-eada9dd0c73f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.531636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.531636Z digest=sha256:4638d065389ea40d205aef0f35392c6cb6bc9c224e2562dc04a702118ea2d99d

Observation e5ce7bd1-b68d-4f4e-aaa7-1dca9bbc100b · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.575343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.575343Z digest=sha256:f04a25b11542ae8a76fbce9fed197a1b7ab5a755ed725bbf20b8d64d5ea73067

Observation 88951f03-dd06-4f25-ab6f-b33e1550af26 · outbound

This paper cites SALM: Speech- augmented language model with in-context learning for speech recognition and translation,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing SALM: Speech- augmented language model with in-context learning for speech recognition and translation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.486867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.621678Z digest=sha256:3b48c236e609c5678c06a7cc274a53d4802ed508c7d9422e81182538a0ad2dd6

Observation e0a989c0-1396-4f1e-9cba-a0d4c16c8906 · outbound

This paper cites End-to-end speech recognition contextualization with large language models,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing End-to-end speech recognition contextualization with large language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.373358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.638136Z digest=sha256:fbfae5939101ab34d334036452ed25fadfa3bb3fc04f6898052095f89ce62437

Observation 1e71e487-a1dc-4eca-8346-8eb7ef13852a · outbound

This paper cites Eliciting the Priors of Large Language Models using Iterated In-Context Learning.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Eliciting the Priors of Large Language Models using Iterated In-Context Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.653716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.653716Z digest=sha256:d04c21bf2baacebbf860e38d42569412818f9bbb3a5a25388572274a289d57dd

Observation 9773ea2a-8ec3-432e-9950-a73c074c1dbe · outbound

This paper cites Deep audio-visual speech recognition,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Deep audio-visual speech recognition,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.274212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.682959Z digest=sha256:5596c227852efc228e66a4488d913864f9c9bee27e17344c94096658ae634260

Observation 82560c5f-b2c8-4036-90d3-a6af7d469a01 · outbound

This paper cites CNVSRC 2023: The first Chinese continuous visual speech recognition challenge,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing CNVSRC 2023: The first Chinese continuous visual speech recognition challenge,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.170063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.713141Z digest=sha256:2f8ebcdb4610de18f6d27e9e7ef0ca44fe6418cd9a6429149f620229360b195c

Observation e126a1f0-ed55-4bd6-8622-f923f61a0da1 · outbound

This paper cites CN-CVS: A Mandarin audio-visual dataset for large vocabulary continuous visual to speech synthesis,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing CN-CVS: A Mandarin audio-visual dataset for large vocabulary continuous visual to speech synthesis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:52.062046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.767581Z digest=sha256:542170e11bc991fbd0c22f920e39c125caecf7dde2a01baac48a7e228a12212e

Observation 2459821c-f360-47ef-bbe0-f41d26cc0bb9 · outbound

This paper cites Reti- naface: Single-shot multi-level face localisation in the wild,.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Reti- naface: Single-shot multi-level face localisation in the wild,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:51.921380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.865248Z digest=sha256:f2ac863932608b1b5e86f3da65f87a15f273133dcdebd69eb374653959d5de7a

Observation 0233da9d-c52d-48dc-a9b5-20952f3ad4de · outbound

This paper cites How far are we from solving the 2D & 3D face alignment problem?(and a dataset of 230,000 3D facial landmarks),.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing How far are we from solving the 2D & 3D face alignment problem?(and a dataset of 230,000 3D facial landmarks),

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:51.677890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:50.916906Z digest=sha256:5a77fbd9dcf45b8eaa9d497687e0471e8174ae55ec5b9ff43b7242d29abb4803

Observation cfc314a5-4cd6-4b34-8c8d-07d91929297d · outbound

This paper cites Qwen2.5 Technical Report.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.951353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.951353Z digest=sha256:8066db39b1f4eecc23b0817ad1474afbf517954818915dfa1410ec152ab7b63b

Pith citing papers

Observation f16fa30a-8bec-4a7e-b1e9-821092d3c397 · inbound

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing cites this paper.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:26:51.504112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:26:48.987998Z digest=sha256:738cfb42d103d8aafbfb5980aaf2389a4d3dcf4bd72003a8c210ea9d6bc958c9