Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:49.220696Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2608.02023.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:49.220696Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 121 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e5b0fa62-cc15-4b0d-8bdb-819fafbdea60 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e07fb11-6f18-4911-a4a6-c6b6b5f1b191 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Ultimate vocal remover.GitHub repository, 2020
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e059a6e2-3ede-44b1-84b5-45540c4c9427 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f70143-b4a3-409d-83a3-d9f70d7afe28 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98db6794-4613-41d4-b1f3-a4e8e93a67df · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Curriculum learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b577c6-f76c-49ab-8ec8-705439e43e4b · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Seed2.0, 2026
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825e6a81-4696-4ec4-9c5d-072f943e5227 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Seed Speech ASR 2.0 Documentation.BytePlus documentation, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bc06fd7-2e8a-4e25-90cd-136a05fd081e · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FlexiVoice: Enabling flexible style control in zero-shot TTS with natural language instructions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8863186-fa0b-4565-b00b-518392991fb4 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebceb8d1-6133-48d2-9af2-c7551f683a26 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a543f5-a0a0-44b8-b3fb-d3e4ff8eec92 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verifi- cation and diarization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2115d0ec-3fc9-4b3c-9ab8-63d41e39a447 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SeniorTalk: A chinese conversation dataset with rich annotations for super-aged seniors
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 26617ea3-3047-4c0c-89fa-e2d4ecbb251a · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab600fc-9fff-4482-87cb-ad71cb6b8404 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a395f1-300c-42b5-9e9a-970f52391e57 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a894298-99d8-4ad0-9db8-60155983ecf2 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks An Unsupervised Autoregressive Model for Speech Representation Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5538af60-d184-442e-9b29-e5546f3af888 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4d01e6-0e87-4803-bcab-35f737f7ba19 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406b16fb-4c00-4ca5-b61a-2ad9403cd726 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks High Fidelity Neural Audio Compression
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629d49a3-6ef6-4718-bcc5-644e9df7c498 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e7f5f51-4028-4126-bf09-e77a18dda934 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2c3f76-4745-4fa3-b25e-6a030932b29c · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275a5ee0-2cdb-44f1-beeb-8887c4937a8c · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Stable audio open
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ea08738-fe24-4a9c-b9d0-96302d56204a · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Stable Audio 3
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ce8da0-1d8f-4a2a-8137-1dab0c692588 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Falk, Chenxi Zheng, and Wai-Yip Chan
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468a3151-13de-494f-bb8d-c1a995014d54 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3559cc-902c-4a5e-8aa7-a43b9f23cc53 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc33832-c003-4685-a1b1-7ee3ac42e8c2 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fsd50k: an open dataset of human-labeled sound events.IEEE/ACM Transactionson Audio, Speech, and Language Processing, 30:829–852, 2021
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0518d6f5-8d75-48c3-b78a-d633529e7370 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ACE-Step: A Step Towards Music Generation Foundation Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c21aa88-2213-4606-98eb-e49dc3c1c900 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gemini 2.5 Pro model card, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ba9792-4aab-40b9-9f65-7f4614df886a · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gemini 3 Pro model card, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f428483-1094-474a-94e4-b9c1f8069bcc · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gemini 3.5 Flash model card, 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398c03d8-690b-4022-b369-76dbf280892e · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gray and John D
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 13cb153f-3be4-453e-b761-803c1b53ce09 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MRSAudio: A large-scale multimodal recorded spatial audio dataset with refined annotations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aaee3ad-b073-4f0a-8d1b-5cdf76fa6c1b · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0675b870-9e75-46ce-a1c9-36c8511e58e6 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks STARS: A unified framework for singing transcription, alignment, and refined style annotation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b288ef5-2c01-4a1f-824a-60db308d8ab5 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Classifier-Free Diffusion Guidance
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e949b5c7-b8f6-4f56-b981-51364bb9d405 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks VoiceSculptor: Your voice, designed by you.arXiv preprint arXiv:2601.10629, 2026
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccfcd34-cce5-4033-b454-67182e2d72ec · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Word Level Timestamp Generation for Automatic Speech Recognition and Translation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9abdc930-d034-4b06-b470-9efcf5667b65 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks python-pinyin: pypinyin, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad53429-5f47-4f33-9b3a-5a55b1bb9063 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfba38d7-261f-471b-8c81-8e48249828ab · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MOSS-VoiceGenerator: Create realistic voices with natural language descriptions.arXiv preprint arXiv:2603.28086, 2026
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76e05c2-caee-4bb0-833f-acfcd4c9bf27 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f142beb-f802-45d1-952f-0c0e75de527d · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd961c4-c379-4c63-bc1a-f899dad96973 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Categorical Reparameterization with Gumbel-Softmax
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7de596-658d-464a-afbe-8d5945ac41c2 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65bf9d44-5877-453f-a5c2-dbfd4a626359 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 145311e9-5980-45c6-83c6-68dd594e8f83 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Latent-domain predictive neural speech coding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d65616-ac24-4a57-9f4f-cfd2485a6be1 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3d068e-ceb0-4675-abbf-063188f005fa · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 913701e3-7d01-4e72-8f4d-9caa3d9c1a73 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MoonCast: High-Quality Zero-Shot Podcast Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aee4add-7bbf-4338-8bcb-e00b8f05c8a9 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Libriheavy: A 50,000 hours asr corpus with punctuation casing and context
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2372d0f5-d560-4c5a-a9f2-25412638214b · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AudioCaps: Generating captions for audios in the wild
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36bcbb4a-06ce-49ff-8ac1-e50046674a45 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31010d63-83be-4fe6-a15e-d470da4b3634 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a0e8681-68a2-4b02-8ca6-9a95ba03e568 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Kubichek
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 89fb67de-fd30-4c14-b4fc-fa250198c48f · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59066311-494b-4bd7-b822-fd5f4f5b1025 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks High-Fidelity Audio Compression with Improved RVQGAN
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 274b58e5-e371-4c0c-bf41-58c2f4508c73 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Parler-TTS.GitHub repository, 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9dbd49-c119-4d3b-b0cf-76080032ac4d · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Towards streaming synchronized spatial audio generation via autoregressive diffusion transformer
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f66b5f-a00e-4e28-b50c-be9cec58a48d · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Robust Singing Voice Transcription Serves Synthesis
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35ae698-ac55-4c9a-a820-2a26e6f09cd4 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a4190413-a002-4726-96a9-ee972d6a1490 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Flow matching for generative modeling
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24b2542-123b-447a-a743-6345e1ce0e7a · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniMoE- Audio: Unified speech and music generation with dynamic-capacity MoE.arXiv preprint arXiv:2510.13344, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48989418-7daa-453e-99e1-7a3eebd4211b · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Decoupled weight decay regularization
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33c18b7-0caa-4bd3-9a9e-2e5011ebca0f · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5bb9fd7-b8dd-44bc-9ec0-195f0b295945 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MOSS-TTSD
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 528e3f56-ed6d-458b-8d18-2d5f0e93ab3e · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Representation Learning with Contrastive Predictive Coding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0397d39e-6d29-4f14-96eb-005d53cb285a · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks A multimodal evaluation framework for spatial audio playback systems: From localization to listener preference
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131edcb5-940e-44b9-96ce-f9d35819cbee · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Audio Editing in the Era of Foundation Models: A Survey
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 26b3406d-42d4-44ca-bf3e-58a62695a12e · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ac8b1340-c8f1-4778-87e3-5c8b05add0ea · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Librispeech: an asr corpus based on public domain audio books
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cecd235-87b3-40c3-8f52-32a819df9aab · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SAME: A Semantically-Aligned Music Autoencoder
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f2379a-1267-444b-ae27-72f2a45d69cd · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Scalable diffusion models with transformers
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b72284d2-4250-44f5-b7c9-cbabb5aba932 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks VibeVoice Technical Report
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe2f0d6-c32e-4975-87b4-98fba54c03ec · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Qwen3-TTS Technical Report
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c51ebe-666a-45fd-ab12-5885f1515156 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Nemo forced aligner and its application to word alignment for subtitle generation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87082ea-d441-4166-be99-8a6cc1be3a89 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0857a230-ecd5-4c80-aacf-0b37d099e946 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks OV-InstructTTS: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459, 2026
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb35624b-16cb-42de-b677-ebf4912623bf · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e388abd-9eed-4c5f-bef9-13498a556294 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 265648e2-a0c1-4985-8af3-e4ba1b4ce885 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95b1244-7755-434f-a7bb-b75e33b5c60f · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd704b5-fe2f-401b-9b97-bb112eec6049 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Improving the Diffusability of Autoencoders
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c175b092-2a48-4944-b38b-757f4ae97ea6 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks The 2018 signal separation evaluation campaign
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db877a3-803c-4a61-bfc3-812b220892e1 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks An algorithm for intelligibility prediction of time–frequency weighted noisy speech.IEEE Transactions on audio, speech, and language processing, 19(7): 2125–2136, 2011
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5326b022-af52-4660-bbc9-502176ebfb76 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Seedance 2.0: Advancing Video Generation for World Complexity
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9facae7-a25b-4d97-bbf6-c3f310d05aae · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde9d30c-28e3-48db-bfa2-a2f7981db857 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45fb951-7073-44ac-a50a-849f3e67cf70 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cabb25-a778-4980-9293-5f01de63f73c · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59ab591-ccf8-488e-80f9-9d7ac157491b · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Soulx-podcast: Towards realistic long-form podcasts with dialectal and paralinguistic diversity
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc61158c-106d-4638-89d6-178270575d88 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc6b733-63c5-4f8b-85b2-a6837ba82714 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Secap: Speech emotion captioning with large language model
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e934298-5e24-4464-893f-e74432a7a2e8 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2bae1a7d-e835-4a69-95b2-ea307491fa90 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Qwen3 Technical Report
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1619c124-c65d-48d3-ac09-b57a83db686b · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bcd2f08e-ee7f-4ed1-b3c2-e482f20237cf · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10357896-112d-4d39-bf8e-1e6cb590b7d1 · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6cb030b0-22d5-417f-8608-fd026805c0db · outbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Zezario, Szu-Wei Fu, Chiou-Shann Fuh, Yu Tsao, and Hsin-Min Wang
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.