Pith. sign in

Paper Citation Record · LEDGER

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

As of 16 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.13062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13062 v3

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.105122Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:30.954743Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:25:31.298438Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ea2c479-b15d-413c-acd7-527324f5090a · outbound

This paper cites Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:25:31.304236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.954743Z digest=sha256:8ae0318599c6caa75d3f249bf928e3b986310276f93cb0848348baf0598970eb

Observation f5731167-1c17-4eed-965a-c55997dc66e0 · outbound

This paper cites SFT for SV AD Given a silent video V = I T t=1 withT frames, the SV AD task aims to generate a corresponding audio descriptionCaudio.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model SFT for SV AD Given a silent video V = I T t=1 withT frames, the SV AD task aims to generate a corresponding audio descriptionCaudio

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.521636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.963218Z digest=sha256:4e52286f3a96988f3ab5cd915b0bfdc1d804936840e0939b823b37dd86b0dfec

Observation eaed0417-0dd9-4561-86bf-b27b85b69929 · outbound

This paper cites an unresolved cited work.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Unresolved cited work

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:25:31.510788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.967354Z digest=sha256:08a31aabf213a72793d6b41052a573e0e0aa460f727efb31411d71c2b6e81232

Observation b1f99f45-1435-4309-ad95-58a9ab79dc68 · outbound

This paper cites an unresolved cited work.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:25:31.500577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.971335Z digest=sha256:a19bb6bd8b6f73fa153bc7ef2782c513c71678b55bf165354a22a6f2fce3b8d4

Observation 284182dc-e2cc-4900-a2d2-3ec29bfd01a3 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:30.990915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:30.990915Z digest=sha256:a1e151c14bb6764a049036b1764c188a15a1d8ca70b04c3614f5efe25f790148

Observation 1ce585bf-1cbd-46e4-8dea-10197763c2d2 · outbound

This paper cites Multisensory- guided associative learning enhances multisensory representation in primary auditory cortex,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Multisensory- guided associative learning enhances multisensory representation in primary auditory cortex,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.489915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.974975Z digest=sha256:f88c6ec2881b53dfc11cf34ae505cd4dc26e515ab5b4df8a63f6c4bb36a6867f

Observation 77e07d20-e03f-49b3-88ca-f220eee99961 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:30.978656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:30.978656Z digest=sha256:9cf8efc8b6f44545cc1e899f398a5b926b235e56591192d0bc706e82c16c3c3a

Observation 362b949f-dad6-4244-a3b8-c006b8b180b2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model VideoChat: Chat-Centric Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:30.982484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:30.982484Z digest=sha256:16e399407bcc9ba5fe32cefc692b0a3338ca55ddba9510cf7c2f8326f0765c11

Observation 2196f4a7-e8cf-44d3-8da0-5a3c6c9a569c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:30.986800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:30.986800Z digest=sha256:aa9189207df7db9cd0f960feaf8c729c3f26a066803c4ab5a5c3b23d3e4c025b

Observation 4c19d327-59f8-4445-8d81-db6b18681e49 · outbound

This paper cites STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.011778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.011778Z digest=sha256:f2a40c268fbfbe021274cf73338e5858d03f4b2d37084a48c631b980e9948246

Observation 9ae3ddbc-b652-4a76-839e-9901d10ec1b6 · outbound

This paper cites Video-to-Audio Generation with Hidden Alignment.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Video-to-Audio Generation with Hidden Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:30.995667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:30.995667Z digest=sha256:f11547b39f135352b1a7c134447cfb82915a8166b5b837f967ed2a6f2d551e79

Observation f15cd0ab-403f-4041-baad-11197ebdf900 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:30.999910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:30.999910Z digest=sha256:1bbcca1968a808a7d7ab123a3fe33834a3e922a1eac05e87e7c4154a4586e0bb

Observation 44878e4b-545a-41f0-a16e-fceac3f57f41 · outbound

This paper cites V2a- mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model V2a- mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.478939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.004061Z digest=sha256:c335a7c93c5974eb65fd1d42ad14015369e86d0188ad815d307090a0ea4868ef

Observation feace7b4-450d-4946-b42d-21e6045664e0 · outbound

This paper cites Foleygen: Visually-guided audio generation,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Foleygen: Visually-guided audio generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.468363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.007650Z digest=sha256:332c83244ffe62abed3eda44d68e4233864ee1ed37fba487053dfb790b848839

Observation 18472efb-57f1-4f3f-b574-bca98d354bf9 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.030495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.030495Z digest=sha256:e4771b683f785d3ea57ffa51e2598abaf269a6b2f9f0b2d70288fd8d4a650ddd

Observation 25a35509-3055-4ad8-a498-f9160489027d · outbound

This paper cites Text-to-Audio Generation Synchronized with Videos.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Text-to-Audio Generation Synchronized with Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.015693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.015693Z digest=sha256:a2451c41d067416a1c8a1a14d2cf8f8a28d9e76b1bfa9847569fead7fdd28bdb

Observation 8fb3c112-9498-4a3f-b0fe-a785e763ebd2 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.019247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.019247Z digest=sha256:d45003af1e89121a11a497fd0f690ac82a173e90d95ec714e78bf1b565ab32cb

Observation a03c0787-f0c9-4648-ac43-3779aea4f037 · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Read, Watch and Scream! Sound Generation from Text and Video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.022815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.022815Z digest=sha256:12b1a86741b36333864db9b9a86637be1ee1e0f82d458199ff1a8890d45dd9c4

Observation f584c6a1-2a11-49c5-8f75-1f9adfd55165 · outbound

This paper cites FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.026317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.026317Z digest=sha256:1431d20a531471ccfce8e349ab3252ec6d0c537c2bd77b4902f419c85aa7da2d

Observation 92e1dd92-2121-47ad-9bc6-fa069eb0e06c · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.048189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.048189Z digest=sha256:5128408d99af2c2818239e50e1f44b94b206766c647a3774db51b0ab6c28af9b

Observation 7707387e-fe86-4ddf-82f1-7b08dc136b3c · outbound

This paper cites Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.033984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.033984Z digest=sha256:2bffd488b1f6a19ae208827e0ed80a20a4f6e2f790ba47130b7145ed92d5c910

Observation 678a54ef-0ac2-4661-a5e2-11c78754244a · outbound

This paper cites Automatic video captioning using tree hierarchical deep convolutional neu- ral network and asrnn-bi-directional lstm,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Automatic video captioning using tree hierarchical deep convolutional neu- ral network and asrnn-bi-directional lstm,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.445167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.037314Z digest=sha256:e15f2d5e56a639cff06daa8957a6fdb0fe6fed5f3e4a200d0a2022329dc71076

Observation dbbae53a-3ae4-421f-9fee-05f82fe597b7 · outbound

This paper cites Streaming dense video captioning,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Streaming dense video captioning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.434478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.040976Z digest=sha256:6603584c8d76caf9e60fbdf79d4594d66644715cd02042c816108baebf01b6d1

Observation d77a8053-30e2-4f0f-a42a-4f421eb3052c · outbound

This paper cites AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.044527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.044527Z digest=sha256:56bb6278054679918e5d2e79cfc08d4e7c3bba0e753a187563145f57284a07b4

Observation 82eef265-2b49-4bf9-bda8-ddcdf46e4150 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model LoRA: Low-Rank Adaptation of Large Language Models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.403181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.066625Z digest=sha256:2f189e60677cb01c1b12e2d7610f438ecdf9c519421aced85d1a31b000dd210e

Observation dd4e23ca-97e8-42fd-9630-4ff4ae785269 · outbound

This paper cites an unresolved cited work.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:25:31.531676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.959376Z digest=sha256:26a185f0841da0fc687bdf98e88e4804b6f46f262378e83c29f0ef5cc95b93a8

Observation 74fa0396-0ea8-4937-835c-55415ff7c5d2 · outbound

This paper cites An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.424081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.051720Z digest=sha256:27925b7eda7bd50e8d1cb88a3c9389a0db0ef5485b1ddfb7f88dd76cf985f108

Observation 9b4b0f06-7b50-45f8-89ba-2a35155a16aa · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.055296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.055296Z digest=sha256:bf0cc5a60a9c1ab1a21e4b9ce30db964676ead416d3686628b03f6d7418b6fa9

Observation 85d49c75-0db9-4a8b-a2c5-6836a03b46f8 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.059117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.059117Z digest=sha256:9d1304a89835cbbeaedf68b11e63ecd39d1c175a01eb0c288d066bb1620913a0

Observation 243a8b76-21f5-40d4-8d5f-b7a273d96fda · outbound

This paper cites Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.413892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.063165Z digest=sha256:056e2151ab8b6a1d5b3d6055845f5fd0c5d4dce889bee9f466cb3a22f5c82ba1

Observation 37f408ac-3dd9-45af-b12c-26f66d1d4934 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Chain-of-thought prompting elicits reasoning in large language models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.070010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.070010Z digest=sha256:f1e7b3cc17ef50dd76fb7bce74333fe3681ef3385378f3af589d04324462ffc6

Observation bd2efa4d-726f-4c0b-8bd7-2e9e2e40e59f · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Audiocaps: Generat- ing captions for audios in the wild,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.073254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.073254Z digest=sha256:a2c127d023f7f65453777c7d0c25df71ddc081a4d24899c572c4bc348811e9a2

Observation a8375c5c-9baf-456c-9441-576008fcf55d · outbound

This paper cites GPT-4o System Card.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model GPT-4o System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.076500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.076500Z digest=sha256:bc75d8f984c35ce80ce5a9e41bb56856dbd4aefa5e963afd9a43286e93f4368b

Observation 98e3f8bc-e07e-4d65-96e2-5a71e5504bab · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.380074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.079772Z digest=sha256:82483b0974d216dcaba10aa99a82b29bfbda9ccabab712ceeddfa69887791c9c

Observation e4209607-048a-4018-8810-7e9937d9f066 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Bleu: a method for automatic evaluation of machine translation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.369004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.083009Z digest=sha256:b5daba77925a4c9bd9da3c36091320ba1d3ed404083f703ba69f72643f0baaae

Observation ca474cb5-5e15-4d39-8400-74b900f0eab9 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.357837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.086629Z digest=sha256:9d980476879df229cf866a4bb29bdd5b6421cdbf09fd2a93992609adb6d4573e

Observation c7563e6a-e57e-4d09-988a-1600ab5f93cf · outbound

This paper cites Rouge: A package for automatic evaluation of sum- maries,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Rouge: A package for automatic evaluation of sum- maries,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.346844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.089834Z digest=sha256:b33427ab32a48a99c7fa1d7aaa2e6d9eec0c04d0326d5acb29c80d6eb12341ce

Observation c4be74c4-4337-4175-8cf9-e9c46ee7b2af · outbound

This paper cites Cider: Consensus-based image description evaluation,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Cider: Consensus-based image description evaluation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.336845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.093343Z digest=sha256:ccef99d3296ec447a8d165bcdea035a464c932df53cab5da49dd3cca367ec2e5

Observation 0a9bf1a8-5169-4aa7-8c10-824cf5de2467 · outbound

This paper cites Spice: Se- mantic propositional image caption evaluation,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Spice: Se- mantic propositional image caption evaluation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.326264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.097455Z digest=sha256:44ca7a05bf7ca6c08cfc00e1a88739ee175b8779b121c6499055da0788251ded

Observation b3fb4b55-5c87-4e36-b638-651ffbeb94f6 · outbound

This paper cites Diverse and aligned audio-to-video generation via text-to-video model adaptation,.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Diverse and aligned audio-to-video generation via text-to-video model adaptation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:25:31.315646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:31.100975Z digest=sha256:9a0859303b1318fecca56fef13702f36d67942742d51c5bf4e11a6c9d11ee214

Observation aef9d9d5-9aa1-4124-a0f2-5894db01ff40 · outbound

This paper cites The Llama 3 Herd of Models.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model The Llama 3 Herd of Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.105122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.105122Z digest=sha256:50f7a367d767dd91dea3f510750b678b75320cceaf2d6a76d50003529c8b9bdf

Pith citing papers

Observation 4ea2c479-b15d-413c-acd7-527324f5090a · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:25:31.304236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:25:30.954743Z digest=sha256:8ae0318599c6caa75d3f249bf928e3b986310276f93cb0848348baf0598970eb