Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:24.646239Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.14988.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:24.646239Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T19:19:36.427337Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T19:21:48.097349Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e6bf2c79-2fd0-478e-99f2-f47e54e9c85f · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Naturalspeech: End-to-end text-to-speech synthesis with human-level quality
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a90a7c8-d495-4e6d-ba73-9f4d7631ddf4 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6790d590-3fb0-49b3-9ce4-c95510e3ae4f · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85822a91-7e9e-4dcf-bcd2-6c853645dff4 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ef7c61-4fb3-480a-a09d-9f6dfc26b888 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SpeechAlign: Aligning Speech Generation to Human Preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcbbabe-5463-4063-8e98-a35ab3e4c91d · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Emo-dpo: Con- trollable emotional speech synthesis through direct preference optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd0dc393-7380-4d9b-86b1-41727b768a4f · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Preference alignment improves language model-based tts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca2867c0-d3af-4cd2-835f-d9016de33486 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b7574c-8814-4de6-bddd-7803eb02cd7a · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af530fdd-197d-40be-94d1-7c92f2be9d9a · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529d8da5-bbd3-4a32-acf6-9eec08dbacfa · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8be1b90-e219-4d50-826e-65bebd9db008 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cbe4af-0b4a-4535-b907-5547b8ff3d60 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe3355e-f1d1-41e3-862e-94f5a87894c6 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4ef9cb-c92d-4bf7-85b4-3b9d025929a2 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d9c95b5b-4da3-4d66-b6fc-52980ababa41 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90256e37-c7f1-4c17-ae37-b363e025ac24 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive speech synthesis with next-distribution prediction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ebd70e-7347-4e0b-897d-2364e7a2b092 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5bb2d7-2828-4156-b580-53b0b7c92847 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ffaf0fb-3276-448b-a359-aea5a05b0e7f · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99216364-feee-44e5-8711-18a96f4a7f90 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis V oicebox: Text-guided multilin- gual universal speech generation at scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a96b3cfc-3d52-47d9-b621-0ef2835b38f1 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 658aac14-d570-47f6-b978-9161804f311e · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a85a2f9-0209-44fd-b840-7bb86963c6e9 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622d2823-e56a-4b2f-b553-fbff7e28ba02 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3eef67e-c8b2-4d44-a2e5-8e5a1da552c8 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83257d1b-2c04-450a-a881-9b1e45af6008 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5014061c-b215-40d4-b96a-38023f01d708 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c2d8318d-4d3b-409b-a075-ba6024e82095 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis One-step diffusion with distribution matching distillation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb28080d-6f40-4ad9-9895-eb0438988ac9 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0784fa1d-a261-4485-823e-b268b951316b · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 29eec92c-be00-43ab-9150-773dc27e258b · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5878af4-9419-4991-9343-afa0a734da01 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30f650b0-ab7f-4b22-aef8-e708754981bc · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Meta-stylespeech: Multi- speaker adaptive text-to-speech generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4d5860e-7b97-4dcc-9c27-e2996a404932 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a08099-a6fe-422f-88f9-48959baf15ae · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43874f57-cbcf-41ca-b2b2-6953c5ccbe0c · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de44fb7-35b5-40b1-8310-b586fe374102 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676782b7-0ac9-4013-aa6c-2f897b0f37d4 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e21841-37ab-4286-b335-9f8e8d6b4e2d · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Improved Distribution Matching Distillation for Fast Image Synthesis
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2280739-c13a-4d94-b749-3a0dfca648b6 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0826a8e-d10e-490f-883c-d8a246018847 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93fe65ff-4cfc-4fb0-8797-154f58472b5f · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49409248-3195-4e3d-b17a-5e25c7bb2263 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8bd800-123e-4fda-8558-b0dd1ccf24ed · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc01169-ea41-4a73-8179-04fa94037182 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Didispeech: A large scale mandarin speech corpus
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 606d3b08-2ae2-41f0-9ea0-a39daca8ac6c · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Fixing Weight Decay Regularization in Adam, 2018
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82368ccc-fcac-4ca8-b7f8-17edf5268a33 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Robust speech recognition via large-scale weak supervision
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf1f8fd-29cf-44d3-b2d5-a4c8a7b6e62e · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77cee763-0d59-4476-95e4-694a2bb62951 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Wavlm: Large-scale self-supervised pre- training for full stack speech processing
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f572e7a-fc9e-4d19-beec-832cc732b043 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Denoising diffusion probabilistic models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3027e3a-81de-4530-a0a3-088acc48b096 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Murphy, and Tim Salimans
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8e6f38-969d-4cdd-bd82-df9d885b2315 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee60b90-8877-4d7e-b9bd-fb896a34889c · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28fe12df-87a1-4804-ae77-4c38a942c1c5 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Audio 1" and
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc29348e-d938-487c-bdb5-4faa341c8391 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation faca40a1-ebba-4799-a019-4c2fc8c16c09 · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4af1368f-b726-4c67-8405-54479b81822a · outbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec6fbd44-3fa9-4e9a-be4e-7a90b42d0953 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
Reference 263
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.