Pith. sign in

Paper Citation Record · LEDGER

UniAudio: An Audio Foundation Model Toward Universal Audio Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2310.00704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.00704 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:16.151902Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e4360b2b-7a9d-4c88-ab5f-5dff3bc5fc37 · inbound

Moshi: a speech-text foundation model for real-time dialogue cites this paper.

Moshi: a speech-text foundation model for real-time dialogue UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:13:22.565669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T08:13:21.962488Z digest=sha256:f7db0ab554aa21cee12e3afb0cac36678e9d7fe397b961d302664a884c654483

Observation 0c1be9db-808d-43b2-8d74-453dc411fda3 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.112262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:294e5e8e46ffb3cd55358833c7135ae9544719f8c77818f9081bed1478196b21

Observation 6c5afbac-e9d2-4862-8f5f-cab96a22e33d · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:16.151902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:16.151902Z digest=sha256:fcba45993c0e4e7030f67c0eedfa9c182a76e895eb7a7579f36b35c98d4c4921

Observation 48e7124c-05e4-43c9-bb0d-10ed4f8b5f88 · inbound

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion cites this paper.

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:08.829562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:08.829562Z digest=sha256:e3be8ac2f9a478b00f9ae265387bbd2b509c8411e2af5b355121471923f02efc

Observation c87ec37f-a114-40b6-acff-57d356b935c6 · inbound

In-the-wild Audio Spatialization with Flexible Text-guided Localization cites this paper.

In-the-wild Audio Spatialization with Flexible Text-guided Localization UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:49.805442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:49.805442Z digest=sha256:8507ecde220da12237ed486ee5afdf4cd6476536833af00e1d83c2e3aeb4d680

Observation 4eb02c9c-0049-4412-b20a-1d0f0b385f28 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:58.032031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:58.032031Z digest=sha256:10db4b5411abf312b00239ae53e593d3ccb2c0636a5a00eaff848adae00dd852

Observation 93b38794-9dc5-4c3f-b4e9-219efef4dac9 · inbound

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion cites this paper.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:24.284181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:24.284181Z digest=sha256:ab65f3a49e69043cd5c0a4edcaa5339ebebc3daf6b9a4ff593bd6877d65fbdc4

Observation e9596670-4fda-4713-bc55-68aeb0e506e2 · inbound

Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration cites this paper.

Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:33:55.754374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:33:55.754374Z digest=sha256:ab257df1fb7f665f2ec719a810e1cb1d7a29384fe543ca8400016119338fc7b8

Observation 4572b879-d5a4-4980-9e0d-41350217fb08 · inbound

Scaling Laws of Motion Forecasting and Planning -- Technical Report cites this paper.

Scaling Laws of Motion Forecasting and Planning -- Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:23:32.781421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:23:32.781421Z digest=sha256:ddbe3a4e73563fb6ce4392c9816e335d452171ab0f7b4dd337dc25ef015f45d8

Observation 981006d2-d1a6-4647-8332-b80cd2f36eda · inbound

Exploring State-Space-Model based Language Model in Music Generation cites this paper.

Exploring State-Space-Model based Language Model in Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.387244Z digest=sha256:8d5180d81d8ee318c2af04371db93ee528297a454fc27067229717915812a4fc

Observation 72949c76-f247-4187-949e-bf301ca95cc7 · inbound

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment cites this paper.

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:07.244165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:14:07.244165Z digest=sha256:44923b482e1671fd0c99fc9229d8709557bceb6b43bcd9cf4fd848109e8a17e9

Observation d682d2a4-1638-47e1-b3b7-c04993354b10 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.524187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.524187Z digest=sha256:072940764f07059b754b5c211d0c435166b218244ba5cf134da00c1d1561b4d8

Observation c6e79082-fc53-44c9-8f1e-6ec750618d11 · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.500277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.500277Z digest=sha256:f2ef4029cccc30e252f6911a3a1d6537291724d71bab707c00c85b8b0fcff263

Observation 118c0019-ded0-417d-97e5-48b2b48b530c · inbound

Small Data Explainer -- The impact of small data methods in everyday life cites this paper.

Small Data Explainer -- The impact of small data methods in everyday life UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:28.943831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:05:28.943831Z digest=sha256:662fddb50c776a846235b53da14a1e6fc5b588a8e605f63f5bb107ae99b4da33

Observation 506e6954-1a24-43c5-aaa2-a2c3c992b00e · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.479314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.479314Z digest=sha256:e7b8593e07ee43f4a579bebe7ad5f1c6df4c131dce61375777336ee83b2559e5

Observation 8ff6b37b-4359-4374-8562-1d9f47b94cca · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.199707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.199707Z digest=sha256:dcbcb077a4cc9c179bd37e7a65efaf4bd7aa217e7d58a5852c8bd83e63d6e68a

Observation ecd473e6-b5a8-42f0-9157-10e01aae665a · inbound

Your Spending Needs Attention: Modeling Financial Habits with Transformers cites this paper.

Your Spending Needs Attention: Modeling Financial Habits with Transformers UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:59:05.807002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:59:05.807002Z digest=sha256:e712b360af851a306fd811a3928b8fe1dd83c89420f4bc7b8ed45c9a54cc9f44

Observation 404acb60-0c1c-40f7-93ad-3c6b2d3cc334 · inbound

Ego-centric Predictive Model Conditioned on Hand Trajectories cites this paper.

Ego-centric Predictive Model Conditioned on Hand Trajectories UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:25.573605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:29:25.573605Z digest=sha256:1ede15c71e51786bade8d982f24544b0020b3bdf12b0f4721e66c9155664a215

Observation e6197b20-9dbe-4590-945c-9a1418149def · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.543438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:1f21c940b9e4183e7c23970c90e63c86789e8370e8303dc56fa9bfd9608ff1ab

Observation a6ccd059-323c-4831-b26c-401f61f95b9c · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.870189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.870189Z digest=sha256:c32d998440267e927843956555e1e1a61f75b787b03efdee3e7b72848b2f7f69

Observation 7c4c84d1-c061-4e10-a641-3c466918b1cd · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.644203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:3952fbb5e9257a32b87011425f709d087673c983c4d512f60fa26c58f6ae8a84

Observation cefc0374-dafb-439b-9a9f-a548839f2bdf · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.106539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:21504acb2113630ed353b608f5380309b1ff487ea10fec36ac5fd47c8fcee1a2

Observation b07c37f9-d40f-46f2-91cb-41d3251b0bbc · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:42.653482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:42.653482Z digest=sha256:2aa47a0e0b5644814382d773b02f93cae180ea263854958ebf5d30e7225a22f5

Observation 3fb4263e-3399-4e4f-b0ee-b933c3288000 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.473180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.473180Z digest=sha256:9f241e492e1204dd9c953f5e672cb8bb91f59178ac9a448a2236e0b4e3e6d63c

Observation c6f8bd97-d23e-4c84-ae04-9a6fff638289 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.178157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:339153ac92a297277124fcccc6608fa3eea75a5bf25c59add8160e8f4fcfb5c1

Observation bbcbef27-93d7-41b0-986d-1c53182bed16 · inbound

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards cites this paper.

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:03:33.272230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:03:33.272230Z digest=sha256:50b6f5306d394015d4f2aeeb4390b6f6cfa3db80bf1710e489d72db1227762be

Observation 71031306-d505-4e03-98b8-542f9d4370e0 · inbound

Simultaneous Speech-to-Speech Translation Without Aligned Data cites this paper.

Simultaneous Speech-to-Speech Translation Without Aligned Data UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:57:01.237575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:57:01.237575Z digest=sha256:7384f895a79cef12438b375b44e70ed05492b9b9cafad07d58e5fc3fd55f1aa0

Observation 0f88f706-4546-4784-8303-9f521e2f6234 · inbound

Fine-grained Soundscape Control for Augmented Hearing cites this paper.

Fine-grained Soundscape Control for Augmented Hearing UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T20:00:06.136667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:00:06.136667Z digest=sha256:2f02a9b734097daa63d76973afa6a3ee506ad97a86a866481f5e1aee31273ef2

Observation 1f131168-718f-4581-ae3f-3eab37b2d5eb · inbound

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cites this paper.

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.928071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:58:26.634355Z digest=sha256:c599943b969f5785ead6c06e06aff4eb848558f232481bfac7ed7c33ddb73477

Observation c0fb6750-12f8-47aa-b72c-659f602c332e · inbound

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation cites this paper.

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:47:43.034952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:47:32.003545Z digest=sha256:7f24ea6a33412f3dc17688ecb7e83f107e6d0328f648758e346c6b7629235d45

Observation e2101993-a06f-4051-818d-136ce8d283bc · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.934369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:70bdf7476bba6139a55b38e1150107b45541b491a3bc6806ad681d2b2c967a3b

Observation 323b2215-b214-4cd6-a90a-ff2a2dc5a7c7 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.586510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:6f5aa80fec373680de5c3cfb66916381b24d506880aefbe5a38303f9705bece1

Observation 3f22033f-a2cb-4491-8224-66d7c5fdcebc · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.522604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:7f3b3025eab84aaa16f8c9e8a4f0ea99db0a2b1c271ba0d1920e770073c3cbb7

Observation 7746aa85-7f8c-453c-b45b-f0bc2f686713 · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.873935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:ab3da3113e12882bb644dc1ed5f770e2246d2035838010f67de5247f273b04b1

Observation e05d13f6-79a7-4bf1-a267-e5cfad63baa1 · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.989373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:43ae7527dd164e7c8c5e2944c52b9271d52515ea4db27ee816cab4b7a77242bf

Observation 22c60912-0aec-4dea-9cd9-2e2be7da006e · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.397390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:b8853e430efa0cee79ffd9274b5857c9bb9c7b56e1f5a38510d33b2b24c4fc0a

Observation 59d0817a-b1a7-4eff-8df7-4707f059cc02 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:6da0c83e5674e20e725e41e4201a55f35eaaf2ba52bdaec3c99e49d313526fd8

Observation 2aa98018-0654-4444-959f-a554c30ce503 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:7d33cd544802f2ff4e57216b8788d61a390b76becd25072f63acac7effc52750

Observation ba1e27fc-e26b-4c2d-8126-fa61f865c25f · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-30T14:08:14.574038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:08:14.574038Z digest=sha256:f0d35e419dc3becb6263ce5977fc1b99658675563a4125a01b061d1eba3e1cad

Observation e8563ad7-2fc0-47f6-abc6-efa67a33c4ee · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T10:17:27.671246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:17:27.671246Z digest=sha256:f32be0ca8f1316b029c162004751ca9c9b67430efca654478b8d5809e04207ff

Observation f6a686b1-0ac0-4803-9929-bcd8db3a2396 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:30.088106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:30.088106Z digest=sha256:d331bb22079a5327096f1a02fb1a0bce660a330830693bf0fd23628fd28d6b95

Observation bcd2f08e-ee7f-4ed1-b3c2-e482f20237cf · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:49.016529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:49.016529Z digest=sha256:6a43f810809bae6256d9d93c02c6c67ebba6c2ba039b5fdb3d7531eb3a715419