Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.105122Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.13062.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.105122Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:30.954743Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T20:25:31.298438Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4ea2c479-b15d-413c-acd7-527324f5090a · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5731167-1c17-4eed-965a-c55997dc66e0 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model SFT for SV AD Given a silent video V = I T t=1 withT frames, the SV AD task aims to generate a corresponding audio descriptionCaudio
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eaed0417-0dd9-4561-86bf-b27b85b69929 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b1f99f45-1435-4309-ad95-58a9ab79dc68 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 284182dc-e2cc-4900-a2d2-3ec29bfd01a3 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce585bf-1cbd-46e4-8dea-10197763c2d2 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Multisensory- guided associative learning enhances multisensory representation in primary auditory cortex,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 77e07d20-e03f-49b3-88ca-f220eee99961 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362b949f-dad6-4244-a3b8-c006b8b180b2 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model VideoChat: Chat-Centric Video Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2196f4a7-e8cf-44d3-8da0-5a3c6c9a569c · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c19d327-59f8-4445-8d81-db6b18681e49 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae3ddbc-b652-4a76-839e-9901d10ec1b6 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Video-to-Audio Generation with Hidden Alignment
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15cd0ab-403f-4041-baad-11197ebdf900 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44878e4b-545a-41f0-a16e-fceac3f57f41 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model V2a- mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation feace7b4-450d-4946-b42d-21e6045664e0 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Foleygen: Visually-guided audio generation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 18472efb-57f1-4f3f-b574-bca98d354bf9 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a35509-3055-4ad8-a498-f9160489027d · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Text-to-Audio Generation Synchronized with Videos
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb3c112-9498-4a3f-b0fe-a785e763ebd2 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03c0787-f0c9-4648-ac43-3779aea4f037 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Read, Watch and Scream! Sound Generation from Text and Video
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f584c6a1-2a11-49c5-8f75-1f9adfd55165 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e1dd92-2121-47ad-9bc6-fa069eb0e06c · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7707387e-fe86-4ddf-82f1-7b08dc136b3c · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678a54ef-0ac2-4661-a5e2-11c78754244a · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Automatic video captioning using tree hierarchical deep convolutional neu- ral network and asrnn-bi-directional lstm,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dbbae53a-3ae4-421f-9fee-05f82fe597b7 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Streaming dense video captioning,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d77a8053-30e2-4f0f-a42a-4f421eb3052c · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82eef265-2b49-4bf9-bda8-ddcdf46e4150 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LoRA: Low-Rank Adaptation of Large Language Models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dd4e23ca-97e8-42fd-9630-4ff4ae785269 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 74fa0396-0ea8-4937-835c-55415ff7c5d2 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9b4b0f06-7b50-45f8-89ba-2a35155a16aa · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d49c75-0db9-4a8b-a2c5-6836a03b46f8 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243a8b76-21f5-40d4-8d5f-b7a273d96fda · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 37f408ac-3dd9-45af-b12c-26f66d1d4934 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Chain-of-thought prompting elicits reasoning in large language models,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd2efa4d-726f-4c0b-8bd7-2e9e2e40e59f · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Audiocaps: Generat- ing captions for audios in the wild,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8375c5c-9baf-456c-9441-576008fcf55d · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model GPT-4o System Card
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e3f8bc-e07e-4d65-96e2-5a71e5504bab · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4209607-048a-4018-8810-7e9937d9f066 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Bleu: a method for automatic evaluation of machine translation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ca474cb5-5e15-4d39-8400-74b900f0eab9 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c7563e6a-e57e-4d09-988a-1600ab5f93cf · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Rouge: A package for automatic evaluation of sum- maries,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c4be74c4-4337-4175-8cf9-e9c46ee7b2af · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Cider: Consensus-based image description evaluation,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0a9bf1a8-5169-4aa7-8c10-824cf5de2467 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Spice: Se- mantic propositional image caption evaluation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b3fb4b55-5c87-4e36-b638-651ffbeb94f6 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Diverse and aligned audio-to-video generation via text-to-video model adaptation,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aef9d9d5-9aa1-4124-a0f2-5894db01ff40 · outbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model The Llama 3 Herd of Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea2c479-b15d-413c-acd7-527324f5090a · inbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.