Pith. sign in

Paper Citation Record · LEDGER

Text-to-Audio Generation Synchronized with Videos

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2403.07938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.07938 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:27.412504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T16:14:15.247445Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b6b1c667-3311-4ac6-be88-5c8557f19490 · inbound

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation cites this paper.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation Synchronized with Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.994918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.994918Z digest=sha256:fedfebc3fc474ee4e6a6abd3655d685c81ce52dcfe9f2f4f15d58e0aa85af6cd

Observation 0d3da6c7-1d27-4941-8bd5-a2b01f48b431 · inbound

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text cites this paper.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Text-to-Audio Generation Synchronized with Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.412504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.412504Z digest=sha256:c525ff6e0d1063b7144d72af9dd2431390491eab29a19d68589a3b74b0b409d3

Observation e5c31a37-908b-4686-8d02-3fd9ea7ef719 · inbound

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis cites this paper.

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis Text-to-Audio Generation Synchronized with Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:35:57.784363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:35:57.784363Z digest=sha256:139997021de7b19ca0bc9aafa68bba28409a66c6e0a1a69076c99a9eefc4589a

Observation 3e7a7405-a2a5-460c-864f-92d9f622b98a · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation Text-to-Audio Generation Synchronized with Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.876985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:874d8439a0da403a07fe3c86350c42c12aec4582172fdc834bf8699dcd10568d

Observation 25aad33e-09f3-4920-8c94-ed11eaeaab10 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation Text-to-Audio Generation Synchronized with Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.249489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:1760a48f2646b9c800c09cc89af01e1fdfdede6917cf640e9085192958358e93

Observation cddac56f-c799-495b-bd5e-c2d6e51f4a91 · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models Text-to-Audio Generation Synchronized with Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.474589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:35d2e68734ea15a6372625d7f0990be241387a1583c7f3968b4f52728c4ea8c5