Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2312.15821.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:21:56.979989Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation c1ec0f3c-064d-4904-b008-082d18969987 · inbound
Movie Gen: A Cast of Media Foundation Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82917358-2cb2-4ecd-a980-b3ad05c3c46f · inbound
Flow Matching Guide and Code Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e0fe9cf-f618-4751-95ca-173c00fee50e · inbound
Overview of the Amphion Toolkit (v0.2) Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6b83bd-d021-4081-b844-7266387cb578 · inbound
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · inbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2415d845-f9a1-4ae3-bb0b-742007a90c25 · inbound
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163393c2-c867-4746-b125-fe627ac4467d · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de5ba83-5a11-4c40-9eda-c9d19a19f0fa · inbound
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · inbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab885eb-2fbd-44b9-a22a-87a76227beed · inbound
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053015d6-0d7b-49f1-b8b8-ec50576d9198 · inbound
LoRP-TTS: Low-Rank Personalized Text-To-Speech Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5236f2fc-4bde-4a3f-a90b-7afba7cc9d0f · inbound
RenderBox: Expressive Performance Rendering with Text Control Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae32be0-7d26-4b9d-b5f7-32d78822003b · inbound
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec86a7c-90b7-4ae1-af69-b106424ed52e · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cad407e-8458-4f77-ac9a-8ffe4e44edda · inbound
In-the-wild Audio Spatialization with Flexible Text-guided Localization Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a35ce5e-51c7-41c7-b247-0f98169daa3c · inbound
InfiniteAudio: Infinite-Length Audio Generation with Consistency Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c4f62d-2778-41f1-879b-66735e0cfc0f · inbound
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · inbound
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77c8120-21a5-4283-bef7-4f137d735d91 · inbound
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca6c30e-5db6-455d-9495-cd915c13d794 · inbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3fd9f59-84fe-4f00-8876-9f085173922a · inbound
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1be8c0-385a-42bf-8d7f-0378490c6428 · inbound
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c609199-248d-450a-8ba7-1316a92b2f4f · inbound
Testing chatbots on the creation of encoders for audio conditioned image generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c5340e-9405-4a0c-a097-db088a4fb847 · inbound
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe87faf-64e1-4cfe-9174-d11e04825029 · inbound
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54156058-05fa-4072-93ba-39b32cde10fd · inbound
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d2d6a9-4041-44da-9dfb-4a344fb57ac6 · inbound
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ed3b4b-f6c3-4c91-ba67-bc68f8dad714 · inbound
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b384b49b-ef6e-4168-93ae-0a5b55b3b17d · inbound
Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ffc0fb2c-350d-44f9-86f3-c1c906de81fc · inbound
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0746874f-c294-4504-8dd5-241db0182cff · inbound
A unified perspective on fine-tuning and sampling with diffusion and flow models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d819b93-7627-48bd-b92b-619c12b3a0ee · inbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66c3dcb2-33d5-4afa-a404-39d9a448298f · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fae6d979-9042-4207-ac15-07057f9144df · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation adaf4418-7ee9-45c4-a77b-e9c57146a96f · inbound
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c1d55db-e97a-4f06-b3ca-42f193b15c9e · inbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 444d884a-4575-45da-9dac-ac58abdf689f · inbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 688eb36f-523f-4e77-a7fc-6e0b0fb11064 · inbound
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 087fe4f2-13f7-432f-bd41-19685fa2e3c3 · inbound
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 435d69f2-c82f-4b7e-b121-cff147ec5199 · inbound
EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 590a8127-654d-4387-90f6-df1426c6ce91 · inbound
VoxCPM2 Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c590f6f-ae5c-47a6-bbc3-0c9cbd2afdb7 · inbound
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2278720b-7402-4a96-ade9-baf10dabe26a · inbound
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 980e0768-d242-4eb0-a098-4155afd207f7 · inbound
Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3d23cec-77d5-403f-9493-0bf03a9ec394 · inbound
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4231bb2-8c63-4c46-82fd-17b2e69f233a · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a7202ec-e2cb-4f8f-b476-2fe65d39ecad · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d288361-f52f-4898-a045-1092ca4af2de · inbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e525aefd-e97b-4c32-93ce-dc1f3fa08ff8 · inbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.