Pith. sign in

Paper Citation Record · LEDGER

Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2404.16305.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.16305 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:27.373999Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:03:15.203891Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e21430f2-83eb-4bc1-9cbf-c25bf59de308 · inbound

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment cites this paper.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.934025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.934025Z digest=sha256:20f64145657aac13774821eaea01cbda6c137be368496f578d5aa74d0b83bf8a

Observation faad95cb-ed2f-4b8d-8d8c-e712d49709e3 · inbound

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text cites this paper.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.373999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.373999Z digest=sha256:e049e0e4332ca80c363342beecfbe82d0a1bd75b5215033c67b8d8e6f9bf3119

Observation f8009c33-1708-4b14-8d49-2048beb884b1 · inbound

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment cites this paper.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.108555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.108555Z digest=sha256:eb168890136ab9cc69af82a7697f6fe7b18afd9d14acd52c81699b9e06a709a6

Observation 7a208eb1-b5ab-4f60-b6cb-f59eda251448 · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:52.764864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:52.764864Z digest=sha256:d2a310e95dc3b77965d91ba52a4cbc85c41145659975f58faef09cb0cca79377

Observation c340d833-540e-49a6-ac7c-09537bd9008c · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.206803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:78596393171e69023aa39573ba3d7fc1b145210c7f729fec95bfeb7c4bfdcd9f

Observation f646db97-c6ea-43c5-877a-b25cd2f2f765 · inbound

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance cites this paper.

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:20.091923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T06:24:21.611234Z digest=sha256:56746ffed48679d778e68827fc8df1febef3de3f0c4ea45b76b478ad6995e5f1