Pith. sign in

Paper Citation Record · LEDGER

A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2409.17550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.17550 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:47:55.955024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:05:48.394109Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7325ee3-c854-42b1-b248-3037bbaf9549 · inbound

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment cites this paper.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.955024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.955024Z digest=sha256:836ad16af5bc8112dae333f9fbb64423b39c684c348c53c96ef167ba71777b1c

Observation 3f26d0e5-5a5e-488e-8b79-6cf965eb5261 · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.302374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.302374Z digest=sha256:9270f0bf60deeecd9839312a8dea5b71aea365ff6ff684b0d6518e3b8910e8b2

Observation c00952e7-6c12-4923-bec4-ce6fee53ae84 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:05:48.399732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T00:05:47.686964Z digest=sha256:fbbb51c21f911f4b02a7bccf48a423233798b165341259b0b249c457a907808b

Observation 295887da-3bab-4ee3-bbf4-dd8932dc092b · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.293402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.293402Z digest=sha256:19690af18573e469fed3ad86cdb25cc17175f4ebb856d51870680cb818b731db

Observation cee6463b-6430-4500-a5af-9c211722155e · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.290355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.290355Z digest=sha256:9d6f5e44203281b2ac746f7879eedafc334364fee19c80cacc2e3c27cd2be9b4