Pith. sign in

Paper Citation Record · LEDGER

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2508.00733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00733 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:25:58.872482Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.612458Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5db16c02-bee4-4b43-8764-32f82764cc54 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.872482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.872482Z digest=sha256:569eac527e799c63a4616a3d14121849b7836818a1517d0db1be8985af593cde

Observation 92cd37fc-20d2-409e-97b6-218e161ecbb2 · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.031207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:39c914a4b28fdac7cdd468d1d7003f7d5f701b008465f207516a64f1bf09ac2c

Observation 3cdf2ce2-884a-4cdd-8752-e50636cef16e · inbound

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories cites this paper.

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.838588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:02:52.748220Z digest=sha256:97e33e59cac991372c315063390897253b2730b8848d5e5e09599811d4030b38

Observation d4725871-7853-4a8c-b72c-0e3b8675d5eb · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.060801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:9286778f08609b8496cdbe457964a7dc435c59eee623fc86dba5ee9f264bcc6a

Observation 3baa02ea-0a50-43c0-be32-347d30bc946a · inbound

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer cites this paper.

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.355982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T21:17:48.886421Z digest=sha256:5e98f683115b5836ac5dc6108d2e4ffe7de76674de335889bc292554303d4824

Observation 1ae5565b-83bd-4351-8c3a-fa0453af18b8 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.615037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:33e91d96dc813346c4a15a1d08c825b384ba096eaf8e40de339e9b3c335faa4c