Pith. sign in

Paper Citation Record · LEDGER

Audiobox: Unified Audio Generation with Natural Language Prompts

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2312.15821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.15821 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:21:56.979989Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1ec0f3c-064d-4904-b008-082d18969987 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.065148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:380643a6250b14e00a9198d44f2508efc7667eeced2917bf8ddb9cbbf7bfa903

Observation 82917358-2cb2-4ecd-a980-b3ad05c3c46f · inbound

Flow Matching Guide and Code cites this paper.

Flow Matching Guide and Code Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:28:14.138973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T10:28:14.014706Z digest=sha256:8635b60d004c18ae17ae5fb2d775ea48bea4dbd6cd09d9b4519780437419bc91

Observation 3e0fe9cf-f618-4751-95ca-173c00fee50e · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.979989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.979989Z digest=sha256:f1e5b7fb2dfffaecaf915603acbfaad0689b1e7289b86505aadfc6ac4d792b62

Observation 1f6b83bd-d021-4081-b844-7266387cb578 · inbound

CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions cites this paper.

CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T10:58:29.045502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:58:29.045502Z digest=sha256:2b0987960f5d96ae429b7da06f53108977c4ff1b60678eab6dbf75dc90ce1357

Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · inbound

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video cites this paper.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.331087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.331087Z digest=sha256:5e99b97607ba5961d8cbe417bea41c3bfe29f071a17ae694e09be1af0fcf2155

Observation 2415d845-f9a1-4ae3-bb0b-742007a90c25 · inbound

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation cites this paper.

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T18:51:12.656812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:51:12.656812Z digest=sha256:46a09632bc667ed9e961e5fc2dab499f9cf07c71d9cc09a9dd8a1834df1848ed

Observation 163393c2-c867-4746-b125-fe627ac4467d · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.761884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.761884Z digest=sha256:971533c7f63cb447c72c95e8843c0cade35ebceae993d680b201b153b89cb324

Observation 1de5ba83-5a11-4c40-9eda-c9d19a19f0fa · inbound

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound cites this paper.

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:30:51.367209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T00:30:51.265062Z digest=sha256:f20b539fbfb09c5b832853605046fdf7f808353f574c52639aea9e4071aff8bd

Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.205829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.205829Z digest=sha256:07e69173bfabe1b7825373eb03d46529f9880199fa47a5b4f0c7425abe6a9b3b

Observation 2ab885eb-2fbd-44b9-a22a-87a76227beed · inbound

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction cites this paper.

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:06:32.937055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:06:32.937055Z digest=sha256:970f402c89bfd0f72079bc4f76463d0990a01c522409a3fcbec38f968d1d2bff

Observation 053015d6-0d7b-49f1-b8b8-ec50576d9198 · inbound

LoRP-TTS: Low-Rank Personalized Text-To-Speech cites this paper.

LoRP-TTS: Low-Rank Personalized Text-To-Speech Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:21:44.261324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:21:44.261324Z digest=sha256:9fdc3f99b73cc4fe03d2c3bc0e6a599834bd7bfe67fdee649d3c4c83dd7281f9

Observation 5236f2fc-4bde-4a3f-a90b-7afba7cc9d0f · inbound

RenderBox: Expressive Performance Rendering with Text Control cites this paper.

RenderBox: Expressive Performance Rendering with Text Control Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T11:51:21.819764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:51:21.819764Z digest=sha256:8aa3e555008a0fda5db929d3a435f487bbf546971908f414639ae1ec45a9bdc8

Observation 9ae32be0-7d26-4b9d-b5f7-32d78822003b · inbound

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations cites this paper.

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.653144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:42.653144Z digest=sha256:e8b47510e721da8210cee9d7637a34cfb32f943384ed6365af7e3f60c22cb189

Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.842801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.842801Z digest=sha256:a0feec5d0bcd8f17d2c769bbd38f2e90791dadaf6a6809d857f520472cbddb25

Observation 7ec86a7c-90b7-4ae1-af69-b106424ed52e · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:04.049255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:04.049255Z digest=sha256:7991dd82458f8f5f8ec9bf6eb2941a78f7f3750bb0d7630c8bd63f7acd166b90

Observation 0cad407e-8458-4f77-ac9a-8ffe4e44edda · inbound

In-the-wild Audio Spatialization with Flexible Text-guided Localization cites this paper.

In-the-wild Audio Spatialization with Flexible Text-guided Localization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:49.503496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:49.503496Z digest=sha256:d6937b68194c38f3b144b1f05b4a5c4fbeca96fda96d258f39b2b9e5ac307fa0

Observation 8a35ce5e-51c7-41c7-b247-0f98169daa3c · inbound

InfiniteAudio: Infinite-Length Audio Generation with Consistency cites this paper.

InfiniteAudio: Infinite-Length Audio Generation with Consistency Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:29.298666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:29.298666Z digest=sha256:b6da271817901fa8324b42d0132ca8277ab84d1a3b9efa792f4c459eec08e55c

Observation 25c4f62d-2778-41f1-879b-66735e0cfc0f · inbound

Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation cites this paper.

Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:24.856067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:24.856067Z digest=sha256:3f5f4c65030e96d60d83b80a2767f16a3fe70fb682357ffad59addd28d1c9217

Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.234441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.234441Z digest=sha256:7adcbdcf0a7e004abb97c238f8679d8682811e2af2d3d25278fef3e83040a488

Observation d77c8120-21a5-4283-bef7-4f137d735d91 · inbound

Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation cites this paper.

Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:17:14.641009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:17:14.641009Z digest=sha256:fa19a9b7a295466114a967a76fa9527c3623ef882992003bcb9f94703f827dee

Observation cca6c30e-5db6-455d-9495-cd915c13d794 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.451574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.451574Z digest=sha256:c428fa62854d1a695bf45ff1037261290558c697734affc6437b1cb189ffbc97

Observation f3fd9f59-84fe-4f00-8876-9f085173922a · inbound

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization cites this paper.

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:49.645421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:40:49.645421Z digest=sha256:dd531710b5704e61ebb22d743321f6ce8bbe326a854abda0ded639e7dca67090

Observation 5e1be8c0-385a-42bf-8d7f-0378490c6428 · inbound

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models cites this paper.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.672162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:7919ea0f4771136106919dc8e867d39c8aedf36e2094879e8684f1a9d1763c13

Observation 0c609199-248d-450a-8ba7-1316a92b2f4f · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.303876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.303876Z digest=sha256:893b616274db4bd90ccb739e13f9681ae02bd5c6083038886422ca5b3c538b16

Observation 85c5340e-9405-4a0c-a097-db088a4fb847 · inbound

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement cites this paper.

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:29:17.809858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:29:17.809858Z digest=sha256:d079e925ca5d6645d85c700ed71e19af5df680d5e192f18652be7eb27033ea8c

Observation abe87faf-64e1-4cfe-9174-d11e04825029 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.922024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.922024Z digest=sha256:fdf612d8cc238bdfd76d83913cde7d77d7ab4ff4730eccf6be02237375991b13

Observation 54156058-05fa-4072-93ba-39b32cde10fd · inbound

FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation cites this paper.

FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:14.257145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:14.257145Z digest=sha256:1f9a30a84d5ff7ff4de58b3461cc3fc128b05ab55c9ec5d03beef6cbe737ac85

Observation 78d2d6a9-4041-44da-9dfb-4a344fb57ac6 · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.115934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.115934Z digest=sha256:404cb06e2b01fdb53ede23e6d3fc07ae91f56960479bbf304000aa87890c8c24

Observation 42ed3b4b-f6c3-4c91-ba67-bc68f8dad714 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:51.864433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:f001288bea73433f0b95748c9497fd46b7679ab57cd4da86fb7d34c5f0c5cc0c

Observation b384b49b-ef6e-4168-93ae-0a5b55b3b17d · inbound

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control cites this paper.

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:28:04.692081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:25:53.037164Z digest=sha256:d3af94e10fa783149615443aa35d9a68dddd44181db72db687e110abf34e21e0

Observation ffc0fb2c-350d-44f9-86f3-c1c906de81fc · inbound

PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing cites this paper.

PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:56:04.361972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:21:43.379525Z digest=sha256:efcea5e5c860c87ec39b7be81a6273845d4ecf7a6b38b4a24d755a0c45b27392

Observation 0746874f-c294-4504-8dd5-241db0182cff · inbound

A unified perspective on fine-tuning and sampling with diffusion and flow models cites this paper.

A unified perspective on fine-tuning and sampling with diffusion and flow models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:21.370267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:43:42.331642Z digest=sha256:f4f308f27d8a38204d7085069c1fd77956ded48a983aa7ea597ef2fbdf9c1141

Observation 1d819b93-7627-48bd-b92b-619c12b3a0ee · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:41.933376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:200467603b833f2ec41e5b4ec42bee92d587df9ffc0916dbd9c922a51050404a

Observation 66c3dcb2-33d5-4afa-a404-39d9a448298f · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:45.292578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:cb555c86108a1f10b7e4600313d88dc978593455ffa05e320c80659602400259

Observation fae6d979-9042-4207-ac15-07057f9144df · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.877386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:566b02767293012832bac41ef7b876414ca052381d7b27b97e5fea6c81d9a2af

Observation adaf4418-7ee9-45c4-a77b-e9c57146a96f · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.342726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:923c785903838e640a9393705459038f8caba2d5431a6e764013e528e3c3d49f

Observation 6c1d55db-e97a-4f06-b3ca-42f193b15c9e · inbound

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models cites this paper.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:41.261751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:dba80ebe69df5b61344d2e23549b7b4022562aeb80125c6a2290cf0269e91a5b

Observation 444d884a-4575-45da-9dac-ac58abdf689f · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.930200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:c41cb364615def39176e0323e3aac0b432f932e1d163f04dcadfa62128cb2264

Observation 688eb36f-523f-4e77-a7fc-6e0b0fb11064 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.626271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:2610f2533debf01c43ccc5b71b806d38668df4411e2534326bf5c54a6a38ec2e

Observation 087fe4f2-13f7-432f-bd41-19685fa2e3c3 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.917213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:dbc899997d326fbb0b05d0bc7ef144cfb18b6eb6499617db5d19cf6e90d5057e

Observation 435d69f2-c82f-4b7e-b121-cff147ec5199 · inbound

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement cites this paper.

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T12:42:08.983227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T12:34:06.024192Z digest=sha256:2528897cb71243a889db83772c595bfc04ffd210311d98314d00cd58a87c6bcd

Observation 590a8127-654d-4387-90f6-df1426c6ce91 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.711855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:9b3894a57627bee552df8285af7acf8b4e3b451ee448548989dc92712c9edde0

Observation 2c590f6f-ae5c-47a6-bbc3-0c9cbd2afdb7 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.662075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:c622a2951cc7611feb390a8bb9c152c8ea6a37e5f9a83f25ead0602315773556

Observation 2278720b-7402-4a96-ade9-baf10dabe26a · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.917841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:2450b321b19e31859a08a9ab39fcf42f861927e68a14fb7efd215d94dd4c82d6

Observation 980e0768-d242-4eb0-a098-4155afd207f7 · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.956747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:52dd673405ab79bb6fb697d49f8195196829fe0a61d49d989239006e75e5dac6

Observation a3d23cec-77d5-403f-9493-0bf03a9ec394 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:da5a4cc7ded5ac0dd1f1cbdb48cbe6f19eb787d38931d6f29e92fd3c7db85e5a

Observation a4231bb2-8c63-4c46-82fd-17b2e69f233a · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.480374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:7a25df0d806f988f9a1258888f22ea118a5940cc8d3ab9a06a4348ea6ec4e4ae

Observation 7a7202ec-e2cb-4f8f-b476-2fe65d39ecad · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:fa79725cc15609332d8aeae43062a3df0324d9d30cf0529e7b1124f92d9ad546

Observation 4d288361-f52f-4898-a045-1092ca4af2de · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-30T14:08:14.459751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:08:14.459751Z digest=sha256:2065f052729b1758f52a2388f5cfe48b4fd9a5a8ff500e867ce22a4e0b07f922

Observation e525aefd-e97b-4c32-93ce-dc1f3fa08ff8 · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T10:17:26.573299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:17:26.573299Z digest=sha256:35b05b91bd3cdc0aff638839d94c4d4dcd3ed4b405e61e5e4d635bdb381ecbaa