Pith. sign in

Paper Citation Record · LEDGER

AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2308.05734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.05734 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.495779Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:37:36.063263Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fd68f75f-f03a-41d5-9719-75abebc37671 · inbound

Training-Free Multi-User Generative Semantic Communications via Null-Space Diffusion Sampling cites this paper.

Training-Free Multi-User Generative Semantic Communications via Null-Space Diffusion Sampling AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:43:43.424894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:38:42.441433Z digest=sha256:fe778129cb1311857ec4df256ba2372776c0b4e72962e26f7f4638f48d1f0903

Observation 641e40c4-6c03-4697-812a-0c7803250b6b · inbound

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing cites this paper.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.495779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.495779Z digest=sha256:8996649c643c83730a316745ba28d3d7f0e874581c9bef36817ffabadb590a66

Observation 8bb57f59-89c8-42be-b281-b63b528f5c1d · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.545619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.545619Z digest=sha256:b118769de851ab84adfd6efb772869dab38853a92b6c52b2f639d3843fdeffb2

Observation 7580dc87-cd7f-4eac-a138-9a22eb42f6cf · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.810718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:10fd139bf58f42849211066f4fd47de7ae46df7a3383cb63bab93230ca99bb5c

Observation f772624d-cf2f-4a4a-be1a-20483ddd75d3 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.909257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:fcbff20ac6d5a8cff4d688090dad8a58f819197221af29dfb34bed6634ff7ab6

Observation efa4768f-3971-4717-b90a-b10cfdcfce84 · inbound

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models cites this paper.

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:36.064677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T14:58:27.176375Z digest=sha256:07298ebf8f17ab8c15c889d7a0f10082d83d0b3d40beba2e19cf08a2be3e4992

Observation d9051232-1bdf-4919-a565-e9ca816cee75 · inbound

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training cites this paper.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.545673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T04:21:38.825926Z digest=sha256:bd13cb53e87c544c54913d9b9a85d5fb6fc30f0692410ae08fd726e780b626c7

Observation 6655ad29-f814-4a34-aaf4-c09ed6593027 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:8e5abea283a9977af26858a52e280dbd2fe5de04d1d91e73cc291165e4d7b7e5

Observation 9a6d927d-40b9-4c36-ab0a-c965da0af895 · inbound

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration cites this paper.

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.967115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.967115Z digest=sha256:60f5a90d18536f797a23248c4a2621a16320436f5450488d1e0dc739b64edb8c