Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T10:17:28.260156Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2607.27011.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T10:17:28.260156Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 506e5225-c9d6-4243-81cc-8c15eefdde8f · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46f2807-32b6-422f-900d-14e96376af32 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Seed-TTS: A family of high-quality versatile speech generation models, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571238a2-8fe4-426c-a2e6-561df36dc1c1 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Neural codec language models are zero-shot text to speech synthesizers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399657bb-fc47-4d67-b213-0facfc60978c · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa75e47-0324-4170-ad01-cf3379c6abf8 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report MMAudio: Taming multimodal joint training for high-quality video-to-audio synthesis, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f789e2c-bf5c-48c4-927e-fdd2b64b9cad · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Simple and controllable music generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9124b13d-8210-4db4-90f6-5dcd7675ff24 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Pushing the frontier of full-song generation: Hierarchical autoregressive planning meets flow-matching rendering, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559b2844-9070-4cf3-82fe-d2fb17e524fa · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fc7bfc-07bd-404b-898d-d6bcd9e6194d · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1ea40e-2747-4209-8c63-a91e2855608e · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Stable Audio Open
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908822d6-a2d9-471f-a6be-2704ccb034da · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Text-to-audio generation using instruction guided latent diffusion model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 649f3337-45fc-4fbd-ab70-b4c5d393151a · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Qwen3-TTS technical report, 2026
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a877c28-b456-4422-b7ef-20c3774796b9 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918f5643-c159-4f51-b0f6-a4920efb92eb · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Algorithms to measure audio programme loudness and true-peak audio level
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb5a1c3-91e5-40f3-8571-8b9c2973ab33 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report ControlAudio: Tackling text-guided, timing-indicated and intelligible audio generation via progressive diffusion modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b7adc2-dbc5-4ef7-a62a-0117b0267918 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report AudioCaps: Generating captions for audios in the wild
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98c2186-ab36-4733-9c07-05177109401a · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report High-fidelity audio compression with improved RVQGAN.Advances in Neural Information Processing Systems, 36:27980–27993, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8d524d-a15b-4cfa-91e0-554c26c1d62f · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report V oicebox: Text-guided multilingual universal speech generation at scale.Advances in Neural Information Processing Systems, 36:14005–14034, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4772290-af44-48b9-9494-a3e1ee86784b · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c33759-9f18-43eb-a923-fa8c07c07fac · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 488ddc19-8962-4c2c-b4a3-f2cd8294aef1 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report AudioLDM: text-to-audio generation with latent diffusion models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ea7b7a-1e63-4a91-98e6-913236c357c4 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report AudioLDM 2: Learning holistic audio generation with self- supervised pretraining.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32:2871– 2883, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0f1623-307a-4354-aa3f-18bfb283f993 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report WavJourney: Compositional audio creation with large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f108f7c-28a8-4cab-90af-39bd440f6e69 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report UniMoE-Audio: Unified speech and music generation with dynamic-capacity MoE, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a7be74d-50e3-4437-a25c-bb32e0a5353e · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8279863-c476-4a39-839c-9ca3c40aa482 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24688499-59ac-41fc-a9b7-904c60937d8a · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Semantic-V AE: Semantic-alignment latent representation for better speech synthesis.arXiv preprint arXiv:2509.22167, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc9cd39-4fcf-425f-97ed-979f9ad00833 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report LibriSpeech: An ASR corpus based on public domain audio books
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4620f63-e242-4900-acfc-50dc690bca6b · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report SAME: A Semantically-Aligned Music Autoencoder
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d43403-7e90-42b5-ab88-eab86becbd06 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report VibeV oice: Expressive podcast generation with next-token diffusion
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c13025-0f5b-4f24-83d3-2905c2adf977 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report UniSonate: A unified model for speech, music, and sound effect generation with text instructions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caca9060-b16f-4511-b0b1-699fe3144e56 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Roman, Christopher Ick, Sivan Ding, Adrian S
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90680f8f-0407-4dd7-864c-9e771caff121 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Scaper: A library for soundscape synthesis and augmentation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afcc044-010a-4536-965e-faab9a9064fa · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146f7151-ba20-49ed-beed-69abb25a8011 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Borderless Long Speech Synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1ffba2-04b1-453a-bea8-375646529e51 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report F5R-TTS: Improving flow-matching based text-to-speech with group relative policy optimization, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00eae277-ad11-46e0-b217-7f0d97936aa7 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b3aead-c6c7-4e08-a7f4-f8435bcd69bd · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report MiMo-Audio: Audio language models are few-shot learners, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14c4f44c-7394-4a0b-a5a2-08ea3b89ad89 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Qwen3.5-Omni technical report, 2026
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb4c030-5a72-463c-9709-fd96937ecdca · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db043c97-9178-46f6-a245-c9813ea56194 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report AudioX: A Unified Framework for Anything-to-Audio Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab84040-62d1-44e9-8333-e1d56672a6fc · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audio-Omni: Extending multi-modal understanding to versatile audio generation and editing
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83543aec-9eb5-4cb2-b2da-c1a0b0cffac5 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9acd1e7-5f6d-46d4-8603-af9b1d27b2c7 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Sound event detection in domestic environments with weakly labeled data and soundscape synthesis
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e525aefd-e97b-4c32-93ce-dc1f3fa08ff8 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57127be6-4a80-4ffd-ae61-3d18154f0bf6 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Back to ear: Perceptually driven high fidelity music reconstruction.arXiv preprint arXiv:2509.14912, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4aa9d1d-9861-4cef-8e4b-0548686ec53b · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report MaskGCT: Zero-shot text-to-speech with masked generative codec transformer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2d7cd0-36d1-444d-8359-b5c767d84b25 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report T2A-Feedback: Improving basic capabilities of text-to-audio generation via fine-grained AI feedback
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 156f046e-0a83-4824-98dc-6cd6e8161118 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88406e0-3ec9-4202-89df-41240fc1d1d4 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d56944-f84f-4556-9aa3-f5b66fb74e28 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Step-Audio 2 technical report, 2025
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f557683b-bebe-4083-b6fb-c666f8b4ff7a · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19dcdb3e-a643-47f2-8b55-8b436ea9c748 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Qwen-Audio-3.0-TTS: Freely controllable and highly robust speech synthesis with multi-stage training paradigm, 2026
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40af82c-65ae-410f-98f7-8c43a16c6f99 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report FireRedTTS-2: Towards long conversational speech generation for podcast and chatbot, 2025
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91888c30-2bf3-4f61-a36e-f577327a5173 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report AudioTime: A temporally-aligned audio-text benchmark dataset
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90be1984-ca7f-4df1-b88e-721d526812dd · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report PicoAudio: Enabling precise temporal control- lability in text-to-audio generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb54e2d-7cf5-499e-80e2-7baa3ac3bbaf · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report LongCat-AudioDiT: High-fidelity diffusion text-to-speech in the waveform latent space, 2026
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeaf2ec1-ef1f-4ba1-acc8-c26b5e89467c · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Qwen3-Omni technical report, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63553ad5-8f05-41ed-8c4a-1a99eb56d355 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report UniFlow-Audio: Unified flow matching for audio generation from omni-modalities
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4808a534-3806-41fc-9a0f-1c071d4f21f2 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Ming-UniAudio: Speech LLM for joint understanding, generation and editing with unified representation, 2025
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8563ad7-2fc0-47f6-abc6-efa67a33c4ee · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f655ec0b-3cab-4f0b-bab2-4d273d4b5237 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report MiniMax-Speech: Intrinsic zero-shot text-to-speech with a learnable speaker encoder, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f6d30a-dbdd-4e62-a223-a7b9b07521af · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456441f0-d541-4f23-9631-989e123b8ba4 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report PicoAudio2: Temporal controllable text-to-audio generation with natural language description
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffaa0108-ed1c-425b-b5f6-2c178d8c08d4 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report V oxCPM2 technical report, 2026
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b4c814f-5bfb-4d8d-85b7-90e2353d9b59 · outbound
Qwen-Audio-3.0-Gen-Preview Technical Report Unresolved cited work
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.