Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T10:28:18.202974Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2605.28063.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T10:28:18.202974Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ece5bcda-246a-4892-8c05-8ffd6a467d9c · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d413b2f-096d-407f-9f01-caefb6e09dae · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 371d6347-b1b1-4839-9dc8-6ea913c27278 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f534a123-42ed-408f-8bca-eecc927ea264 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Qwen3-TTS Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bfea9cc9-0dc8-4d42-8dc2-c4457ece23b4 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiogen: Textually guided audio generation,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04477e7-b9e9-4068-a968-b89a18add4f8 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Plumbley
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 46e73937-64d8-44ac-9a81-f22cf4d1f8ab · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a9413d29-44bb-4b6b-9150-39d4fdc32e0b · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e4932df2-b7c3-4b7e-adab-71f687d3de51 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b0d47770-c9c5-4768-85e5-a4b28b290007 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 444d884a-4575-45da-9dac-ac58abdf689f · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b30c9e72-229f-4392-99ef-d8ffa8b9726a · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0741283b-44c7-485a-b48f-3f377bb160fa · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts A udio C aps: Generating captions for audios in the wild
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dfbf643f-4677-4750-9cc7-a5636ba2e02e · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ac3978-4a6f-4105-9708-88906345c40e · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio- Language Multimodal Research
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0fec8409-fc71-4ffe-8518-67e7d803d2bf · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c1bc5f90-f798-4570-9dff-5051fbf68911 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Freeaudio: Training-free timing planning for controllable long-form text-to-audio generation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a491638-ff52-4895-90a5-e3ffaea17fac · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8d226f0-b228-4cda-8687-c8de081ba039 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts model-predicted CoT
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 15ef98f5-b922-4d06-a182-15a3368a65c9 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Flexivoice: Enabling flexible style control in zero- shot tts with natural language instructions.arXiv preprint arXiv:2601.04656,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c51c5e4-1523-44d3-a286-74cf8ac7a6d0 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2101993-a06f-4051-818d-136ce8d283bc · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d90cdc08-f892-4acc-9532-5626d9dd7b9e · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Fugatto 1: Foundational generative audio transformer opus 1,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312a9d9a-08e2-4680-8c09-50e62f97e871 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Available: https://openreview.net/forum?id=B2Fqu7Y2cd
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032ef048-05c2-49d6-bc39-abfb96283acf · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Chain-of-thought prompting elicits reasoning in large language models,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d00847-865e-4617-89f8-b018de417282 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Cot-vtm: Visual-to-music generation with chain-of-thought reasoning,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a775bd8c-8635-4455-abaf-a63e67bd9e69 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts In: Pro- ceedings of the 33rd ACM International Conference on Multimedia (ACM MM)
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4f459bcf-c373-403e-8527-2132986c57b2 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Ov-instructtts: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4f9adc71-74b0-4299-bd3e-49e20a41e968 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69930641-b79b-4333-8aa6-b23ca2d6a1b4 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Training Large Language Models to Reason in a Continuous Latent Space
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 53010181-7d25-4ccd-b174-76a5bbb84785 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025a
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a369a4ab-51fb-4b34-9c01-f58b2a0c16d6 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6bb8d847-8a32-48fa-b849-63909e5dacac · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Gemmeke, Daniel P
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f1828948-1ee8-4a80-b90a-2366ec50d8d0 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Robust speech recognition via large-scale weak supervision,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf62ae6b-102a-437f-8fa0-7184b8bbf937 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Simple and controllable music generation,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08905fd2-4446-4e2b-ac4a-c5d44b33ca3a · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ec4ee320-d471-4e3b-aafb-48d439831d16 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77870d47-ec8f-40a9-8618-a66372ee6102 · outbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
No inbound Pith citation observations are available.