Pith. sign in

Paper Citation Record · LEDGER

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

As of 22 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2608.02023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02023 v2

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:49.220696Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 121 outbound references displayed

  • verified exact4
  • verified fuzzy3
  • unresolved91
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5b0fa62-cc15-4b0d-8bdb-819fafbdea60 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.102675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.102675Z digest=sha256:dd059ed588542c53c592ba133f49ae2e0912fa74ef2f8fe6569d0d687a3b1e45

Observation 3e07fb11-6f18-4911-a4a6-c6b6b5f1b191 · outbound

This paper cites Ultimate vocal remover.GitHub repository, 2020.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Ultimate vocal remover.GitHub repository, 2020

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.196456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.196456Z digest=sha256:8209cd54b4555357f4b067a031534339c293411117b1c0be972d7b88f1ed57ee

Observation e059a6e2-3ede-44b1-84b5-45540c4c9427 · outbound

This paper cites Qwen Technical Report.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.280587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.280587Z digest=sha256:47c86cc9b6c875904a54cc0ced43444b6c6c7f263fdc360430a58fc2ae180675

Observation 50f70143-b4a3-409d-83a3-d9f70d7afe28 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.363603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.363603Z digest=sha256:deded2aa8e572caa1168d90c940f39cc462451d9e1f68df8eb484028b127b202

Observation 98db6794-4613-41d4-b1f3-a4e8e93a67df · outbound

This paper cites Curriculum learning.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Curriculum learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.453324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.453324Z digest=sha256:cdcf02f7e73d060b1b617f011a42af27559153cea5a847c7b87cf9ca805a0666

Observation b0b577c6-f76c-49ab-8ec8-705439e43e4b · outbound

This paper cites Seed2.0, 2026.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Seed2.0, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.534869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.534869Z digest=sha256:524edf90e5e160b996b8709b89b41044231321c2e713c18a19f0ca040e34f7dd

Observation 825e6a81-4696-4ec4-9c5d-072f943e5227 · outbound

This paper cites Seed Speech ASR 2.0 Documentation.BytePlus documentation, 2026.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Seed Speech ASR 2.0 Documentation.BytePlus documentation, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.593752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.593752Z digest=sha256:cc7f4f422dc5b70823fb9562bcb6073fddd4b457baef9eb82301f18720ef6634

Observation 7bc06fd7-2e8a-4e25-90cd-136a05fd081e · outbound

This paper cites FlexiVoice: Enabling flexible style control in zero-shot TTS with natural language instructions.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FlexiVoice: Enabling flexible style control in zero-shot TTS with natural language instructions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.675718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.675718Z digest=sha256:e577df24925d2422df2d4a94d9fdd697eb50828c039366cb87df84ce080bfe68

Observation b8863186-fa0b-4565-b00b-518392991fb4 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.762978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.762978Z digest=sha256:2723e0756e090569268d9562a4bd882b7ea0589d9e111d638bfbe6ec05c53a5c

Observation ebceb8d1-6133-48d2-9af2-c7551f683a26 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.850191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.850191Z digest=sha256:7d81d84e7d04d5c95c89aba6c49df92e7d95300a90aae52474632902d7d0d6d9

Observation 80a543f5-a0a0-44b8-b3fb-d3e4ff8eec92 · outbound

This paper cites 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verifi- cation and diarization.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verifi- cation and diarization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.931436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.931436Z digest=sha256:30bce2eb56fb7d8720dd7014870f7b21e51f133e7def2d940ae311da63d29f9c

Observation 2115d0ec-3fc9-4b3c-9ab8-63d41e39a447 · outbound

This paper cites SeniorTalk: A chinese conversation dataset with rich annotations for super-aged seniors.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SeniorTalk: A chinese conversation dataset with rich annotations for super-aged seniors

Reference 12

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:14:53.843147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:41.989969Z digest=sha256:645c49cf81542bbe7eda7cadc47ada039cbe1f4737879b16abf3aae5d52c2378

Observation 26617ea3-3047-4c0c-89fa-e2d4ecbb251a · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.071504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.071504Z digest=sha256:91f4fd862d4226f9cdd8a048145948348547dd7b420250bbec5905bf3a9fe31a

Observation bab600fc-9fff-4482-87cb-ad71cb6b8404 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.146192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.146192Z digest=sha256:a6b30b87819a24b3b64d44be2c3c40af3c8ae5a4b7ed7b192332cc985f298535

Observation 56a395f1-300c-42b5-9e9a-970f52391e57 · outbound

This paper cites an unresolved cited work.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.210992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.210992Z digest=sha256:49ffb3c953813a13bba5c8fa162f5e295509c733f6f8038d2c5136b99dedf44a

Observation 8a894298-99d8-4ad0-9db8-60155983ecf2 · outbound

This paper cites An Unsupervised Autoregressive Model for Speech Representation Learning.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks An Unsupervised Autoregressive Model for Speech Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.280938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.280938Z digest=sha256:35d60516d206b4ece910ba377e6004c417cd9f85d780c5ba38ab15f1459cfa79

Observation 5538af60-d184-442e-9b29-e5546f3af888 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.345506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.345506Z digest=sha256:bcf59b974f3b068d70ee9e8b221ca8ac36ad7eeb2db3e5b37de81c6809599fd7

Observation fa4d01e6-0e87-4803-bcab-35f737f7ba19 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.422224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.422224Z digest=sha256:8ad4268675973ec72e6122b27bbfc84167c5d20487af942996a00103252cdd3c

Observation 406b16fb-4c00-4ca5-b61a-2ad9403cd726 · outbound

This paper cites High Fidelity Neural Audio Compression.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks High Fidelity Neural Audio Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.483675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.483675Z digest=sha256:88ae23197958eab07045e2b704b55929a3225a4eb1ed302e82dae3e11af45918

Observation 629d49a3-6ef6-4718-bcc5-644e9df7c498 · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.547078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.547078Z digest=sha256:ca1763449c2b00bd3ffd878c5e33f98437620f099da7036bccc104959584f0fc

Observation 5e7f5f51-4028-4126-bf09-e77a18dda934 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.592339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.592339Z digest=sha256:ba4c1bab92e1187f292cdcf3f5aaabf3a5f2cc29517e9faa3ca8ab4c1cf83b2c

Observation 9a2c3f76-4745-4fa3-b25e-6a030932b29c · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.646582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.646582Z digest=sha256:d19b3574813e7c4876aa71f818e4a4476141632cce881623f02ec1a6bb5b4d63

Observation 275a5ee0-2cdb-44f1-beeb-8887c4937a8c · outbound

This paper cites Stable audio open.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Stable audio open

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.701090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.701090Z digest=sha256:54f8e002a65435a783ffe0fa5ff95b524ef4b364e4eae73f7e5e8b351214b81a

Observation 1ea08738-fe24-4a9c-b9d0-96302d56204a · outbound

This paper cites Stable Audio 3.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Stable Audio 3

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.763577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.763577Z digest=sha256:8784bc7f6056e74d637b761cd554996b107254332373dbe2418a90a7378089c3

Observation e8ce8da0-1d8f-4a2a-8137-1dab0c692588 · outbound

This paper cites Falk, Chenxi Zheng, and Wai-Yip Chan.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Falk, Chenxi Zheng, and Wai-Yip Chan

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.814904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.814904Z digest=sha256:6bf875059e93d5e19dde56fff6b08f7f76a7ca92b9058bcbc8a3b375ab6b7c27

Observation 468a3151-13de-494f-bb8d-c1a995014d54 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:42.943789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:42.943789Z digest=sha256:abb5d17e6b7f8d3dbd7fdbab6166e016f0e9ac188045438bcd0ac23bd17a1ca8

Observation ee3559cc-902c-4a5e-8aa7-a43b9f23cc53 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.018619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.018619Z digest=sha256:2c0f50d2480c07a3ea986c561369aadfc724029fbbbc4c064cfd75d8828ff2e9

Observation 9fc33832-c003-4685-a1b1-7ee3ac42e8c2 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events.IEEE/ACM Transactionson Audio, Speech, and Language Processing, 30:829–852, 2021.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fsd50k: an open dataset of human-labeled sound events.IEEE/ACM Transactionson Audio, Speech, and Language Processing, 30:829–852, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.073690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.073690Z digest=sha256:9514d0a45c9cac344df1dd07d7a52175c04528258e905fd613f7454249c44fd8

Observation 0518d6f5-8d75-48c3-b78a-d633529e7370 · outbound

This paper cites ACE-Step: A Step Towards Music Generation Foundation Model.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ACE-Step: A Step Towards Music Generation Foundation Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.128810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.128810Z digest=sha256:3421137ceb4a14460072c657f8d85b6b49b90ba54d56cba22c7501c3b828e5e7

Observation 4c21aa88-2213-4606-98eb-e49dc3c1c900 · outbound

This paper cites Gemini 2.5 Pro model card, 2025.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gemini 2.5 Pro model card, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.186318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.186318Z digest=sha256:ad1daddda8ac88b770a74b35a1e39d031b458c402f2075dd3ebc7c3f19b71357

Observation a1ba9792-4aab-40b9-9f65-7f4614df886a · outbound

This paper cites Gemini 3 Pro model card, 2025.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gemini 3 Pro model card, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.234916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.234916Z digest=sha256:2d5783f0d1bb0d6900fabb38cf253e71f7d8d4173efebf3cdc10661cd5ca4029

Observation 8f428483-1094-474a-94e4-b9c1f8069bcc · outbound

This paper cites Gemini 3.5 Flash model card, 2026.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gemini 3.5 Flash model card, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.287513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.287513Z digest=sha256:f67e2603f4881d8acd1591ddadc93b89e16286f17dc913c21476c05fc7c17874

Observation 398c03d8-690b-4022-b369-76dbf280892e · outbound

This paper cites Gray and John D.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Gray and John D

Reference 33

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T00:14:53.099472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:43.350871Z digest=sha256:de1f44400017528443095b8eca40c50fb2160b32a7b20435ad687379f294cdd4

Observation 13cb153f-3be4-453e-b761-803c1b53ce09 · outbound

This paper cites MRSAudio: A large-scale multimodal recorded spatial audio dataset with refined annotations.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MRSAudio: A large-scale multimodal recorded spatial audio dataset with refined annotations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.425923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.425923Z digest=sha256:72e1de5af6b228a2ec9252bbf9c32b44f639dcd217529d297a209fe7026f6e7f

Observation 3aaee3ad-b073-4f0a-8d1b-5cdf76fa6c1b · outbound

This paper cites TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.486294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.486294Z digest=sha256:24a6cc88290bf975f597dbae9002364fb916f49f86b8d14a9eb12e595b865a1d

Observation 0675b870-9e75-46ce-a1c9-36c8511e58e6 · outbound

This paper cites STARS: A unified framework for singing transcription, alignment, and refined style annotation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks STARS: A unified framework for singing transcription, alignment, and refined style annotation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.568806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.568806Z digest=sha256:32bd0d76e76ceddcb523ec1b593f027872670b8a9ec0bc0f3be63a5f37405a33

Observation 5b288ef5-2c01-4a1f-824a-60db308d8ab5 · outbound

This paper cites Classifier-Free Diffusion Guidance.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Classifier-Free Diffusion Guidance

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.690618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.690618Z digest=sha256:d989a9ce4e5c89aa0c597779da3c2f53c8039d08edff4d0d70d27beadccf667f

Observation e949b5c7-b8f6-4f56-b981-51364bb9d405 · outbound

This paper cites VoiceSculptor: Your voice, designed by you.arXiv preprint arXiv:2601.10629, 2026.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks VoiceSculptor: Your voice, designed by you.arXiv preprint arXiv:2601.10629, 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.751576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.751576Z digest=sha256:bb60622ec9bcbc0462c6029219abe26a0f130d844a464202a905613e63121d36

Observation 4ccfcd34-cce5-4033-b454-67182e2d72ec · outbound

This paper cites Word Level Timestamp Generation for Automatic Speech Recognition and Translation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.788931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.788931Z digest=sha256:9fe09044a35affabc8710112972b405dd07c9bd93e4277d189a37bf414ced2bd

Observation 9abdc930-d034-4b06-b470-9efcf5667b65 · outbound

This paper cites python-pinyin: pypinyin, 2023.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks python-pinyin: pypinyin, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.865086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.865086Z digest=sha256:7fd5f1973a2b1ceb9513d6e4be5f7753eba35aab85a13223dd5679dee302dbb3

Observation 8ad53429-5f47-4f33-9b3a-5a55b1bb9063 · outbound

This paper cites InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.938774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.938774Z digest=sha256:893bc90efc457778db82c937bba11f8172858cd1f0fd2953fe5692abbb1a31a3

Observation cfba38d7-261f-471b-8c81-8e48249828ab · outbound

This paper cites MOSS-VoiceGenerator: Create realistic voices with natural language descriptions.arXiv preprint arXiv:2603.28086, 2026.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MOSS-VoiceGenerator: Create realistic voices with natural language descriptions.arXiv preprint arXiv:2603.28086, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.013468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.013468Z digest=sha256:b4d5772f157a2f4da842be410a6d04d8a218f2e75fae1da6a7c76a715de51dcc

Observation f76e05c2-caee-4bb0-833f-acfcd4c9bf27 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.072588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.072588Z digest=sha256:ba14dd04ebcc73ad17153415a88752b7816ab83cbe14174520618167438f3ce8

Observation 4f142beb-f802-45d1-952f-0c0e75de527d · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.128295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.128295Z digest=sha256:0cb5a08b6b34cc283b35980914aea3db2943203fb9e9f07fb9493edda22d1d14

Observation 8bd961c4-c379-4c63-bc1a-f899dad96973 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Categorical Reparameterization with Gumbel-Softmax

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.188454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.188454Z digest=sha256:0c7037c5688b514fa96599079e92522e44ef71d312e8b2868c63a2f20bc3ea75

Observation de7de596-658d-464a-afbe-8d5945ac41c2 · outbound

This paper cites UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.291965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.291965Z digest=sha256:f26916313fa629eb8a9dd8549ae7502fe724b29d6c84742feb8cc0092a065023

Observation 65bf9d44-5877-453f-a5c2-dbfd4a626359 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.413014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.413014Z digest=sha256:47b1de0252c41524540f00914c474570770770eaa920b04f5a475c29062c93ad

Observation 145311e9-5980-45c6-83c6-68dd594e8f83 · outbound

This paper cites Latent-domain predictive neural speech coding.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Latent-domain predictive neural speech coding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.516351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.516351Z digest=sha256:880afedd448ab587dd4aa53c239afe1ecb880f7b1608517819689f0ee4db5933

Observation f1d65616-ac24-4a57-9f4f-cfd2485a6be1 · outbound

This paper cites Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.603836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.603836Z digest=sha256:2b43aa0eab4dc3ee83bfbeaa7e6b98aa554df2bb13da0bcbbe6db91da86d1f6d

Observation bf3d068e-ceb0-4675-abbf-063188f005fa · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.716159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.716159Z digest=sha256:beeb64b4f785dbfd5d6c6ac09c257089c1c14daae534b8e473d4cfd5c55b538d

Observation 913701e3-7d01-4e72-8f4d-9caa3d9c1a73 · outbound

This paper cites MoonCast: High-Quality Zero-Shot Podcast Generation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MoonCast: High-Quality Zero-Shot Podcast Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.813923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.813923Z digest=sha256:69450258adf04e8561a8dab7d0c84e7669f91d1a58185b19fccb0b01f4b0ed7e

Observation 3aee4add-7bbf-4338-8bcb-e00b8f05c8a9 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Libriheavy: A 50,000 hours asr corpus with punctuation casing and context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.915123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.915123Z digest=sha256:b8d91d0e499eaed62f8ce14130e9e63b064eaf55f1eee9ec69d685b4abcff334

Observation 2372d0f5-d560-4c5a-a9f2-25412638214b · outbound

This paper cites AudioCaps: Generating captions for audios in the wild.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AudioCaps: Generating captions for audios in the wild

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.013023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.013023Z digest=sha256:298113d18f5d4f9aee63ea8f7edf09e750d80a745dda88e5fe933476caacc377

Observation 36bcbb4a-06ce-49ff-8ac1-e50046674a45 · outbound

This paper cites Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.110663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.110663Z digest=sha256:93ab3a9e16fab21b22be2675564e719f0e12e12d904d42e8f89298a5c92a46bd

Observation 31010d63-83be-4fe6-a15e-d470da4b3634 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems, 33:17022–17033, 2020

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.209549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.209549Z digest=sha256:550d2ce8729bd3670b88e0b3938908658ce3afd9e38c3a7739558a717e3e6b8b

Observation 3a0e8681-68a2-4b02-8ca6-9a95ba03e568 · outbound

This paper cites Kubichek.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Kubichek

Reference 56

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T00:14:52.634038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:45.299385Z digest=sha256:8d00853e236572f9201ae55da8ccbfd94c8cea25fb9cfe91cf1d7da57feed8ba

Observation 89fb67de-fd30-4c14-b4fc-fa250198c48f · outbound

This paper cites Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.422509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.422509Z digest=sha256:55ad6c224c054523f23fafb4e70beb3e9adf5f30493a58b6acf9abc78b40254f

Observation 59066311-494b-4bd7-b822-fd5f4f5b1025 · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks High-Fidelity Audio Compression with Improved RVQGAN

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.767421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.767421Z digest=sha256:c42f8a334909a40775edf8091d7cccde6558bc01bbc2479851ba6cd9ab0f9abc

Observation 274b58e5-e371-4c0c-bf41-58c2f4508c73 · outbound

This paper cites Parler-TTS.GitHub repository, 2024.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Parler-TTS.GitHub repository, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.897887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.897887Z digest=sha256:d49bc606472f11574bcb9d7fba3a04676c836587bb700e47b1a5cfe42199cd5a

Observation 5d9dbd49-c119-4d3b-b0cf-76080032ac4d · outbound

This paper cites Towards streaming synchronized spatial audio generation via autoregressive diffusion transformer.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Towards streaming synchronized spatial audio generation via autoregressive diffusion transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.972331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.972331Z digest=sha256:aa5c3bb145f4df85e6029e04bb2fb8f2e7af207d679cc991b8314300bc1a4e0b

Observation 64f66b5f-a00e-4e28-b50c-be9cec58a48d · outbound

This paper cites Robust Singing Voice Transcription Serves Synthesis.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Robust Singing Voice Transcription Serves Synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.067633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.067633Z digest=sha256:59b2dcb3d2faf0e5a79abffc06c49f8d681d9c7f2ee0ab9bdd9c0622a7931856

Observation f35ae698-ac55-4c9a-a820-2a26e6f09cd4 · outbound

This paper cites SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:52.340147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:46.125240Z digest=sha256:ce04cdf5c30ba8a51d6476c7d480f985b8c74854dc305e10c1da51be8db5828a

Observation a4190413-a002-4726-96a9-ee972d6a1490 · outbound

This paper cites Flow matching for generative modeling.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Flow matching for generative modeling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.222847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.222847Z digest=sha256:8a089a85fa7bdf70426dd84c51b58f13f9ef651863550b1944efb61dcee12876

Observation f24b2542-123b-447a-a743-6345e1ce0e7a · outbound

This paper cites UniMoE- Audio: Unified speech and music generation with dynamic-capacity MoE.arXiv preprint arXiv:2510.13344, 2025.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniMoE- Audio: Unified speech and music generation with dynamic-capacity MoE.arXiv preprint arXiv:2510.13344, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.318767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.318767Z digest=sha256:f6755263546ec1ed4607522ea6f1a015f189f617f193f05e9de17c8883e4fa13

Observation 48989418-7daa-453e-99e1-7a3eebd4211b · outbound

This paper cites Decoupled weight decay regularization.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Decoupled weight decay regularization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.430188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.430188Z digest=sha256:e0c55202fcbdc8ed159a78d46493181cd9e504eb6c4ee06b42fcc3345fd90019

Observation d33c18b7-0caa-4bd3-9a9e-2e5011ebca0f · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.524756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.524756Z digest=sha256:b012bf7ce36964f1536f3bf3890a6b69e72d3ad4afc9ca4d899af58e9ea8c2b3

Observation f5bb9fd7-b8dd-44bc-9ec0-195f0b295945 · outbound

This paper cites MOSS-TTSD.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MOSS-TTSD

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.600213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.600213Z digest=sha256:caeb427edc9db07aeeff418fb4aafd9e6776ef076bdc8bad6e9c1b0c69790189

Observation 528e3f56-ed6d-458b-8d18-2d5f0e93ab3e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Representation Learning with Contrastive Predictive Coding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.711486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.711486Z digest=sha256:c150e7814de3f8a1e4c7257b9f54e891df60e2471f5ae92960bc1175f74774e4

Observation 0397d39e-6d29-4f14-96eb-005d53cb285a · outbound

This paper cites A multimodal evaluation framework for spatial audio playback systems: From localization to listener preference.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks A multimodal evaluation framework for spatial audio playback systems: From localization to listener preference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:46.787489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:46.787489Z digest=sha256:f14435b8d06387561afc6c6bc7780d3d5d473f5835166c660b508d9854c31a22

Observation 131edcb5-940e-44b9-96ce-f9d35819cbee · outbound

This paper cites Audio Editing in the Era of Foundation Models: A Survey.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Audio Editing in the Era of Foundation Models: A Survey

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:52.074519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:46.893801Z digest=sha256:df6fc4d29e2ae8c79a6e503ecd4fa766c9c3e7286b8efea2d5b30c99a3631485

Observation 26b3406d-42d4-44ca-bf3e-58a62695a12e · outbound

This paper cites Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:51.835378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:46.946227Z digest=sha256:34a4e24d8dd7410045eca6b494dff0ffb2f671a55ee3a3d7a61be3c8df3b1d9b

Observation ac8b1340-c8f1-4778-87e3-5c8b05add0ea · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Librispeech: an asr corpus based on public domain audio books

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.015110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.015110Z digest=sha256:c661fc25d05db730b020984ed273ae5bf927d3ff5a3acef4f0cb051764c16110

Observation 7cecd235-87b3-40c3-8f52-32a819df9aab · outbound

This paper cites SAME: A Semantically-Aligned Music Autoencoder.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks SAME: A Semantically-Aligned Music Autoencoder

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.082431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.082431Z digest=sha256:fb57283fd3f645a6781dfcdc514efe3393488be804b7b9c5f3ecd9a1dc758e35

Observation 85f2379a-1267-444b-ae27-72f2a45d69cd · outbound

This paper cites Scalable diffusion models with transformers.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Scalable diffusion models with transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.166984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.166984Z digest=sha256:06f7e583bc3ef00124c3e0525b7968f6cd79f5f9db4a4b96a6d9d546dbf2cf72

Observation b72284d2-4250-44f5-b7c9-cbabb5aba932 · outbound

This paper cites VibeVoice Technical Report.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks VibeVoice Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.221328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.221328Z digest=sha256:24d7887198bedd647f09e6e8f16317180d593a8d77b8109d701d84ba57e90f9a

Observation 8fe2f0d6-c32e-4975-87b4-98fba54c03ec · outbound

This paper cites Qwen3-TTS Technical Report.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Qwen3-TTS Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.314501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.314501Z digest=sha256:8b75b0752833382644abc8706ad9031244517437cf44638cf0f51c024ce14296

Observation 70c51ebe-666a-45fd-ab12-5885f1515156 · outbound

This paper cites Nemo forced aligner and its application to word alignment for subtitle generation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Nemo forced aligner and its application to word alignment for subtitle generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.388246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.388246Z digest=sha256:9d7a35ecb8b0ca0160b376bc43d830afe9015ecffd343500573ca751203d10fa

Observation a87082ea-d441-4166-be99-8a6cc1be3a89 · outbound

This paper cites an unresolved cited work.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.480628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.480628Z digest=sha256:eabcd76701c3b9ab5a3717ecd003f610ed891bd073ec17208c59a0150d2d4256

Observation 0857a230-ecd5-4c80-aacf-0b37d099e946 · outbound

This paper cites OV-InstructTTS: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459, 2026.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks OV-InstructTTS: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459, 2026

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.571640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.571640Z digest=sha256:9161b1e892244065cb61ab1aed21e2edd2e592aae2892ff1bae9b98d7634723a

Observation bb35624b-16cb-42de-b677-ebf4912623bf · outbound

This paper cites an unresolved cited work.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.640432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.640432Z digest=sha256:a0ab31a6c0668057b3ffcbac78a9b870adf48162301ad406660e7b8fbefc428d

Observation 7e388abd-9eed-4c5f-bef9-13498a556294 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.719103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.719103Z digest=sha256:dca9143fb143fa14a7a697e38eac398c027a97f5e558861b798a559b72fd2f02

Observation 265648e2-a0c1-4985-8af3-e4ba1b4ce885 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.793573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.793573Z digest=sha256:37db1d38a24d07c4a357b94171f6201b2e63438eaf902fc3e212f9e1e648122f

Observation a95b1244-7755-434f-a7bb-b75e33b5c60f · outbound

This paper cites AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.871653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.871653Z digest=sha256:58b3e5032753b4bc07951457d07071c5d824d4b0a0c03c2225f107d657518f3e

Observation 5dd704b5-fe2f-401b-9b97-bb112eec6049 · outbound

This paper cites Improving the Diffusability of Autoencoders.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Improving the Diffusability of Autoencoders

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.940224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.940224Z digest=sha256:3dc8bf9932e09c18cce125e07bc2bfa47f622a0b00f0a4379610e82e80804eb0

Observation c175b092-2a48-4944-b38b-757f4ae97ea6 · outbound

This paper cites The 2018 signal separation evaluation campaign.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks The 2018 signal separation evaluation campaign

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.024601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.024601Z digest=sha256:fdac58e5c80ee432658e38fabbcce40b1629ccbb9cfaf20cf472cb48f8ac8d63

Observation 1db877a3-803c-4a61-bfc3-812b220892e1 · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech.IEEE Transactions on audio, speech, and language processing, 19(7): 2125–2136, 2011.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks An algorithm for intelligibility prediction of time–frequency weighted noisy speech.IEEE Transactions on audio, speech, and language processing, 19(7): 2125–2136, 2011

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.098240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.098240Z digest=sha256:a08559b7456969f748379811558a496e024bfcd3ac0db4203c2de00b38819219

Observation 5326b022-af52-4660-bbc9-502176ebfb76 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Seedance 2.0: Advancing Video Generation for World Complexity

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.166617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.166617Z digest=sha256:74d99e847733d63153516a727336937fdd1a24f983b66b4de8c190a5b6a94441

Observation d9facae7-a25b-4d97-bbf6-c3f310d05aae · outbound

This paper cites FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.234028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.234028Z digest=sha256:663b337c3bb6e6d907a02aac451faa3b4f50c2581502cbff38da62364b997489

Observation dde9d30c-28e3-48db-bfa2-a2f7981db857 · outbound

This paper cites CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.305543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.305543Z digest=sha256:fca5a2ba8185e8703f3e56feca452c572dd18f148c1756d2e5467b5970fb8fca

Observation b45fb951-7073-44ac-a50a-849f3e67cf70 · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.376493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.376493Z digest=sha256:749bd5aa41e5fc7d4ed4467b09ae444b40e2a4faf40b3a4a8cb299bc38a8a9d6

Observation c9cabb25-a778-4980-9293-5f01de63f73c · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.455337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.455337Z digest=sha256:4c6046f4c6bf2e374c6fbe745935fd0a76e47e9fa4f9f255c8cd9042995bef3b

Observation c59ab591-ccf8-488e-80f9-9d7ac157491b · outbound

This paper cites Soulx-podcast: Towards realistic long-form podcasts with dialectal and paralinguistic diversity.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Soulx-podcast: Towards realistic long-form podcasts with dialectal and paralinguistic diversity

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.549099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.549099Z digest=sha256:dbf86536aac78def834afa179acf98556dff4b9289b69781ad6a5b60d34d4512

Observation fc61158c-106d-4638-89d6-178270575d88 · outbound

This paper cites FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.608557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.608557Z digest=sha256:cf4d855e21b65c7ee0dce6b435874314c0d246b93e39fa105054051e8d6fae23

Observation bbc6b733-63c5-4f8b-85b2-a6837ba82714 · outbound

This paper cites Secap: Speech emotion captioning with large language model.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Secap: Speech emotion captioning with large language model

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.682385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.682385Z digest=sha256:057c1f56643adb6860808e61f770136b000946f439fc745fffba8968e486a1d3

Observation 3e934298-5e24-4464-893f-e74432a7a2e8 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), 2019

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:55.631663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:48.748834Z digest=sha256:08b31a4aca27c6073c7f16b3f792c796c61345ae8ec660437a5fd53ca1480385

Observation 2bae1a7d-e835-4a69-95b2-ea307491fa90 · outbound

This paper cites Qwen3 Technical Report.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Qwen3 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:48.819903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:48.819903Z digest=sha256:9da9e2cc787e54b1667dcb46501d329125359fef9572a78c068d53d5984d5c84

Observation 1619c124-c65d-48d3-ac09-b57a83db686b · outbound

This paper cites an unresolved cited work.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:14:55.528938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:48.916590Z digest=sha256:8703e61e2c1bfd37d63ba8444d2d6e3f15b986619bb24a337cd40dfeda7f5938

Observation bcd2f08e-ee7f-4ed1-b3c2-e482f20237cf · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:49.016529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:49.016529Z digest=sha256:7aa6e0010cc98e5ea9ccda547e75113a4ca4ffdd86f9521b7ca829c31f85fe58

Observation 10357896-112d-4d39-bf8e-1e6cb590b7d1 · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:55.429260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:49.125255Z digest=sha256:c69472d8d503f91e5b7da5ef21540f48b9b5b17fb5369445890485f486cfb027

Observation 6cb030b0-22d5-417f-8608-fd026805c0db · outbound

This paper cites Zezario, Szu-Wei Fu, Chiou-Shann Fuh, Yu Tsao, and Hsin-Min Wang.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Zezario, Szu-Wei Fu, Chiou-Shann Fuh, Yu Tsao, and Hsin-Min Wang

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:55.322898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:14:49.220696Z digest=sha256:2bc640814bcb6735d5a62c203160c6f300ef90a5235237f4750d2e02646a5f35

Pith citing papers

No inbound Pith citation observations are available.