Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2402.01912.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:51.142199Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 381e5d4b-ee52-4d40-b293-c6eadfcfbea1 · inbound
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1222f66f-e60a-4da6-a341-18a4d700dfa7 · inbound
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db7c9d04-f6fa-4db1-9bc1-def38a7a896a · inbound
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d643197-ab5f-49a1-96c9-e56bb18d5d6e · inbound
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0230389-f536-413d-b453-4a4d8395e27d · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9f7812-7aa8-4c1b-ba5a-0257ba1e96db · inbound
Optimizing Multilingual Text-To-Speech with Accents & Emotions Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb386c66-e20f-43ce-a4ca-d2c323e05703 · inbound
MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2cf6a8-f131-4464-984c-ec5e78ae53f9 · inbound
Multi-interaction TTS toward professional recording reproduction Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1545d8-6f18-4bb2-92e6-f9bf768217e7 · inbound
SecureSpeech: Prompt-based Speaker and Content Protection Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb68026-84dd-4f38-be9c-6b9913761e75 · inbound
Unlocking Speech Instruction Data Potential with Query Rewriting Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea71d72-4db1-4369-9af7-bc31544e2993 · inbound
BoSS: Beyond-Semantic Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce65cf1-7f38-4863-bf96-b730fdb28a2c · inbound
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604d43fb-5ff3-4590-b857-a3b1e3df899f · inbound
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeae0d3f-9346-482b-8379-4d91f7bfc692 · inbound
Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160ea297-49c7-498b-8f39-1db404dc9e29 · inbound
Qwen3-TTS Technical Report Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 949c6284-1e5f-489c-ab16-b3c6e92f45be · inbound
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91288b67-7835-4021-acc9-f75b037dc95e · inbound
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff90eac-0152-49b9-a645-8644c62c3689 · inbound
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b13ef46-b70c-4ac7-87e6-6a8f1720f315 · inbound
Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8758a34a-8d0e-4f9b-b29b-0d694d21bec0 · inbound
Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12c7d0ed-e83f-4cc2-907b-800c1af311b4 · inbound
Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0b47817-d91b-4f24-b4e5-bb1081c9a2ee · inbound
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f6ae9c-89d1-4459-9a1f-92ba7f33e61a · inbound
GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f00501ae-1501-41ad-b754-7230a5da2bd1 · inbound
VoxCPM2 Technical Report Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b681c0fc-bc11-41d0-9b8e-f75c196e9e85 · inbound
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ad237b5-80b3-46d7-9f7a-40afd9551ec8 · inbound
FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4655fea6-e4ec-4da8-8bdb-f9354cff3497 · inbound
EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e2215f8-42e7-4a5b-9347-715b6a74b32e · inbound
Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f255a2-f5d5-4f26-a467-9c911d7b2c6a · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33c18b7-0caa-4bd3-9a9e-2e5011ebca0f · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8851c956-d783-434b-9400-67f0ca5946d0 · inbound
Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.