Pith. sign in

Paper Citation Record · LEDGER

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2409.19132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.19132 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.575530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:28:18.779237Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5b295371-2952-460f-a006-c9e040a762ad · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.575530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.575530Z digest=sha256:4829fbe1b33b3a6807db15cd6955bfeb61adcb12302a78007c014fd51bdab942

Observation 76428a5b-8200-43c9-a5fe-e62957cc6005 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.763813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:90cc83bed3e4d903fb50cf063349225164b1a0c10ce21280c0369c1836c920ea

Observation c7ae6a2e-3729-4f90-9057-e3bc1fc0f94c · inbound

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals cites this paper.

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:50.844186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:30:30.431777Z digest=sha256:859b611fdc694f9fd41db9a4bd8a79019112e880d01384f2251d927ecf509467

Observation 4c6031e5-0730-446d-a82b-6215795da607 · inbound

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning cites this paper.

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:10:59.444346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:46:43.584943Z digest=sha256:94906dd12c9534eb136d82aebe98ff3931a820e87988e1ba6973f9c304f2db18

Observation 4c469386-d564-4fbb-88f1-64ad2f9fc04e · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.234269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:b75ca5a89ff00e81bc5248bb9f3eb5992272020e60ce0a33cd617ad2f107e18d

Observation 9041fab1-11ca-4481-aad0-9e7d6e411fcc · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.780712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:d5a3cfe90988fe9693117f814935e6f16882bd8a658abf0698db14f2d12ebc34