Pith. sign in

Paper Citation Record · LEDGER

DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2503.22265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.22265 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:03.363252Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.609674Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9b7cf25e-f4f1-4e12-95bd-7241ae4fd2e3 · inbound

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing cites this paper.

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:03.363252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:03.363252Z digest=sha256:2b74ef87aa985eae265324d44c4a5465f506dfabf5e463c67cfd007d1bf58642

Observation cf95400b-5535-4a27-8787-6e8f5fe34cb5 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.967160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.967160Z digest=sha256:645c810d3820fa66c801dce51da42ff260eb8956691f9bf55ac8a7cbdcafd0a5

Observation ea2fe1d3-4c7d-4479-8ac7-4d133687f277 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.201964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:620a08719d4c1ebd3633a6bdb9463455a3d2daffc3486c3df4fba0a9466f2569

Observation eea46098-d7d4-4b19-8d6c-04de3a768fae · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.541154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:a33a76ab40f1b1b07932b6fbbe3ba341e2c70a7344bb5f5f1341a7127fb4d521

Observation 22c41ad1-d414-458e-8d08-7e50d83c24fb · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.611351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:a93553a02055082241f861b4e3c84f473e046a7491c5273e5808c92e991b6222