Pith. sign in

Paper Citation Record · LEDGER

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

As of 12 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.14988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14988 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:24.646239Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T19:19:36.427337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T19:21:48.097349Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact5
  • verified fuzzy16
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6bf2c79-2fd0-478e-99f2-f47e54e9c85f · outbound

This paper cites Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Naturalspeech: End-to-end text-to-speech synthesis with human-level quality

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:31.639139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:19.733560Z digest=sha256:899f5a1d2049079eef240625358a7b8024b387d07d36ea4d25d070c50dc4278a

Observation 0a90a7c8-d495-4e6d-ba73-9f4d7631ddf4 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:31.364146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:19.789105Z digest=sha256:ba46930b327c3194493570576c1ea4a639a84809e2c518abb924004df950e049

Observation 6790d590-3fb0-49b3-9ce4-c95510e3ae4f · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:19.866141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:19.866141Z digest=sha256:50e97dd7dc496c12c42e10f81fba951fcad303e9b207b7a36285562c10672ce7

Observation 85822a91-7e9e-4dcf-bcd2-6c853645dff4 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:19.939041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:19.939041Z digest=sha256:8654927b135c23e684a8a33ef7cf63dc5e3cf77682b3e7a758876850d5549f04

Observation a6ef7c61-4fb3-480a-a09d-9f6dfc26b888 · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:19.994804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:19.994804Z digest=sha256:03ea7f0a743d64c8ddcd2ac8f0b9f88361da5144dde90ce212afe4aad778ce2f

Observation bfcbbabe-5463-4063-8e98-a35ab3e4c91d · outbound

This paper cites Emo-dpo: Con- trollable emotional speech synthesis through direct preference optimization.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Emo-dpo: Con- trollable emotional speech synthesis through direct preference optimization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:31.081647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:20.065138Z digest=sha256:5fb952f0633bb79fc1bb57713cb4c59033b369f516c5031a5b38a7f7544bb3e7

Observation cd0dc393-7380-4d9b-86b1-41727b768a4f · outbound

This paper cites Preference alignment improves language model-based tts.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Preference alignment improves language model-based tts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.861946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:20.157944Z digest=sha256:bacbe6cda83d568b9e9ae21b8ab191bceb98b59ec43027f0b6df13262b43218e

Observation ca2867c0-d3af-4cd2-835f-d9016de33486 · outbound

This paper cites Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.224849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.224849Z digest=sha256:5c90c4144bd04dd67dc702517a0900a25f6e83fa0be0bfedbe5f339da2dd87f1

Observation 99b7574c-8814-4de6-bddd-7803eb02cd7a · outbound

This paper cites Evaluation of Best-of-N Sampling Strategies for Language Model Alignment.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Evaluation of Best-of-N Sampling Strategies for Language Model Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.282544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.282544Z digest=sha256:aa67ff9e7e35502a43cf95b36ebcc7c4fc838b5716ba0344443f5ffa0994140d

Observation af530fdd-197d-40be-94d1-7c92f2be9d9a · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.353019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.353019Z digest=sha256:5150698e2635634b2f008400139f8c37863a7730155fd12d21d9918dc6b8e278

Observation 529d8da5-bbd3-4a32-acf6-9eec08dbacfa · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.428168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.428168Z digest=sha256:f6fd379d2c1fde4874469ba57899ecb08ab048663733b170c825e1a33dea0c00

Observation c8be1b90-e219-4d50-826e-65bebd9db008 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.513937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.513937Z digest=sha256:aeac94e7e8d446f0d987c748348cd14e864486aae0a82c91186f94bb8f49cd7c

Observation b3cbe4af-0b4a-4535-b907-5547b8ff3d60 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.576312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.576312Z digest=sha256:3ee1be44e80383fb3a27956523904e90bd07b4727a83c1a730611efdddbfa9c9

Observation 5fe3355e-f1d1-41e3-862e-94f5a87894c6 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.644717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.644717Z digest=sha256:7899a133f9582a8876b3bf5d53255184f90914cbad489a68e788eaa1600bac0e

Observation 9b4ef9cb-c92d-4bf7-85b4-3b9d025929a2 · outbound

This paper cites Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.585461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:20.694493Z digest=sha256:e334f1c50b3df8e8f70c2252922626ec1be081cec360c7c5e0799545d245396a

Observation d9c95b5b-4da3-4d66-b6fc-52980ababa41 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.763920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.763920Z digest=sha256:ee10b8a849f7a2951e63e55a4fe734886ae394eabcc72b001db06f28164ba709

Observation 90256e37-c7f1-4c17-ae37-b363e025ac24 · outbound

This paper cites Autoregressive speech synthesis with next-distribution prediction.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive speech synthesis with next-distribution prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.825000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.825000Z digest=sha256:19e77910dcdb1eff3ce6ec586089e240e18a668a2e54cf0c2dfcaa3695ccf964

Observation e4ebd70e-7347-4e0b-897d-2364e7a2b092 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.918034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.918034Z digest=sha256:533f0fd26fdc5916c941fb73b123956331fd1d148753dc924a8adc0ad5875f79

Observation 4b5bb2d7-2828-4156-b580-53b0b7c92847 · outbound

This paper cites TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.978975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.978975Z digest=sha256:8c3f7546d5b9c067ed51016757f2d630054a1ecdaf09d99b2377123c4a8ecce2

Observation 6ffaf0fb-3276-448b-a359-aea5a05b0e7f · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.077843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.077843Z digest=sha256:c9d7843b29d4a4d5c6bdcca06288bfe6fa66bc30bfb999349fe2c977711c6f33

Observation 99216364-feee-44e5-8711-18a96f4a7f90 · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis V oicebox: Text-guided multilin- gual universal speech generation at scale

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.331261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.126247Z digest=sha256:a432cfd637363be6eeda7bbb4247a0a54c2637f54a5eaaa6508b198d4c6c2e92

Observation a96b3cfc-3d52-47d9-b621-0ef2835b38f1 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.184511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.184511Z digest=sha256:4890af62f364a5261700c6b6a0358a8fd7a584c13033f8f85dfded850edf9f4d

Observation 658aac14-d570-47f6-b978-9161804f311e · outbound

This paper cites StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:26.218116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.255061Z digest=sha256:2da6de609dba4904863a346faeb7656b3ceacd78e6971a18a19e4c0bc44d22d1

Observation 6a85a2f9-0209-44fd-b840-7bb86963c6e9 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.292062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.292062Z digest=sha256:4461d844fcc0128fe8c7e428056ea71fb18f8ecf3ba0ce5b2ede68f698bd3bb2

Observation 622d2823-e56a-4b2f-b553-fbff7e28ba02 · outbound

This paper cites SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.980592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.342231Z digest=sha256:e361ee2029e6011c1110583050eb7ce8a0be650efe07b1fa2159ee57475f5c07

Observation b3eef67e-c8b2-4d44-a2e5-8e5a1da552c8 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.419026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.419026Z digest=sha256:d2f93e35f84bc3eaeb7a82529abe33a1128026e18c3ada98ff16f11243bfeb5a

Observation 83257d1b-2c04-450a-a881-9b1e45af6008 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.515476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.515476Z digest=sha256:c1097e5ac3c85124593cef590708dff9d2751143e0ffbd17abe18c764d82e9d1

Observation 5014061c-b215-40d4-b96a-38023f01d708 · outbound

This paper cites DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.669295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.591183Z digest=sha256:17e4f1d158917554ac534e6ed550411aff8fdcd518f8df50d6e67c4359d74c55

Observation c2d8318d-4d3b-409b-a075-ba6024e82095 · outbound

This paper cites One-step diffusion with distribution matching distillation.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis One-step diffusion with distribution matching distillation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.024351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.679831Z digest=sha256:3cceee1526b7645cd604997545f143422f974815df47f9820dd719730c30a41a

Observation cb28080d-6f40-4ad9-9895-eb0438988ac9 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:29.772307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.742937Z digest=sha256:5ad86967d9b788d22362867689fc362abd55337786575e6a147edcac368b88b4

Observation 0784fa1d-a261-4485-823e-b268b951316b · outbound

This paper cites SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.406941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.812155Z digest=sha256:6c665735aa799c8e88a68fb95c4f31eeb1d0b88d1dee41f996a0f6a8598f02e1

Observation 29eec92c-be00-43ab-9150-773dc27e258b · outbound

This paper cites AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.182680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.866669Z digest=sha256:252905b898ab3e7cec13f08efabbfc80e7314f7eeb6927e157d30429ce9f1f39

Observation a5878af4-9419-4991-9343-afa0a734da01 · outbound

This paper cites Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:29.499117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.938505Z digest=sha256:56468b2af9a1c99262b3594c904423599c361695474fc6b200eac9df4fad03dc

Observation 30f650b0-ab7f-4b22-aef8-e708754981bc · outbound

This paper cites Meta-stylespeech: Multi- speaker adaptive text-to-speech generation.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Meta-stylespeech: Multi- speaker adaptive text-to-speech generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:29.215543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:21.998568Z digest=sha256:29306f8b75dfe1c77a1de8adbe7239f021c25ecf59129cdd244f3811d48bc1fb

Observation c4d5860e-7b97-4dcc-9c27-e2996a404932 · outbound

This paper cites StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.067217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.067217Z digest=sha256:9dcc64fc50fc17fd12f6f08b75e228d83a33665b128c405cae34ed2b6f4b2c09

Observation d0a08099-a6fe-422f-88f9-48959baf15ae · outbound

This paper cites NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.157083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.157083Z digest=sha256:6d79d08bd44b9d09a82901c53d6748299a8a76f864e0db5cf1d672cb91de43e2

Observation 43874f57-cbcf-41ca-b2b2-6953c5ccbe0c · outbound

This paper cites DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.252443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.252443Z digest=sha256:7fd85bd5614bfb59074297a7c31e7cd8329e0e23636a18bb141d1fda5a1841fe

Observation 3de44fb7-35b5-40b1-8310-b586fe374102 · outbound

This paper cites F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.347437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.347437Z digest=sha256:e28cb2be72b9aacb9490afed3db360b50aaf30f09bfd56dff078217a1eba8d57

Observation 676782b7-0ac9-4013-aa6c-2f897b0f37d4 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.435245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.435245Z digest=sha256:65fc8776b2f1d011de9ff9e99bce96bd42b0aa2432a788a85137d369fb35f274

Observation 04e21841-37ab-4286-b335-9f8e8d6b4e2d · outbound

This paper cites Improved Distribution Matching Distillation for Fast Image Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Improved Distribution Matching Distillation for Fast Image Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.532258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.532258Z digest=sha256:fcec62873f5a1624d9c2291ef3846c94ddaf1cf432a99228d7ec87e7a7448378

Observation b2280739-c13a-4d94-b749-3a0dfca648b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.624809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.624809Z digest=sha256:4b22f436f0b35a81e592353b35fd100c8f520cb17ef4472091322951dfe7de08

Observation b0826a8e-d10e-490f-883c-d8a246018847 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.709549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.709549Z digest=sha256:ebb4c156af66a68f56999739fa10e6914924f87c102c999459b665bfc8d0a721

Observation 93fe65ff-4cfc-4fb0-8797-154f58472b5f · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.966603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:22.772891Z digest=sha256:ca7c9f62b3505169746740d18ecc877d27ff3c945e1518e63b6725e933701a5f

Observation 49409248-3195-4e3d-b17a-5e25c7bb2263 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.884960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.884960Z digest=sha256:c2f2be1b722f27ae00c1084a429fa5d9ebb1655ba33ab12ee903d82611837b32

Observation 1c8bd800-123e-4fda-8558-b0dd1ccf24ed · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.974182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.974182Z digest=sha256:29aeaac4436673af2983fe396010bd9a089968a0eef8b03ff2ab564dc0c10077

Observation ebc01169-ea41-4a73-8179-04fa94037182 · outbound

This paper cites Didispeech: A large scale mandarin speech corpus.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Didispeech: A large scale mandarin speech corpus

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.666445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:23.060690Z digest=sha256:e4c0fe532df422e261bbe7572491db2b082185ed1d53646e23fa06cfbcab80c5

Observation 606d3b08-2ae2-41f0-9ea0-a39daca8ac6c · outbound

This paper cites Fixing Weight Decay Regularization in Adam, 2018.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Fixing Weight Decay Regularization in Adam, 2018

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.380022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:23.185093Z digest=sha256:29750928c22af7a2b2f5324bfbb4eadc2d7eb3418b533e8f983e8d10de0d3f55

Observation 82368ccc-fcac-4ca8-b7f8-17edf5268a33 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Robust speech recognition via large-scale weak supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.288864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.288864Z digest=sha256:e9dffb5a72db3dc47e0ab0ad27fba8dc68968aaa9b1d7b80b646b5627cc9587c

Observation 5bf1f8fd-29cf-44d3-b2d5-a4c8a7b6e62e · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.415106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.415106Z digest=sha256:dd31791978e8e47752a3fe95a5ac2d5dee603b9d5fd0544ebb4ddeec0db73f22

Observation 77cee763-0d59-4476-95e4-694a2bb62951 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Wavlm: Large-scale self-supervised pre- training for full stack speech processing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.170855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:23.616993Z digest=sha256:9adbfa8df95e72056b91f4afeb9ba02ef1e5ed0f68e0317a9acac9adbd49e1db

Observation 6f572e7a-fc9e-4d19-beec-832cc732b043 · outbound

This paper cites Denoising diffusion probabilistic models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Denoising diffusion probabilistic models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.719566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.719566Z digest=sha256:422da6baa1a2d9d599e0a47b23ff09f63d7f7a6afd1a901119ff2473328b6aa6

Observation f3027e3a-81de-4530-a0a3-088acc48b096 · outbound

This paper cites Murphy, and Tim Salimans.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Murphy, and Tim Salimans

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.802869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.802869Z digest=sha256:768dcf61ad87cdab429c13c5912de4b83374ebbfb4e58b9612b05eda8d44d3d0

Observation 9a8e6f38-969d-4cdd-bd82-df9d885b2315 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.922139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.922139Z digest=sha256:eb4c8fe4c5bb1ce53f695fb74d0bd6e92d54da22faa90566e7ac15416eabbd3e

Observation eee60b90-8877-4d7e-b9bd-fb896a34889c · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:27.875223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:24.092300Z digest=sha256:e4f062312af09373f9c26ac9810eade202e80785e138bbc009205b5b08771205

Observation 28fe12df-87a1-4804-ae77-4c38a942c1c5 · outbound

This paper cites Audio 1" and.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Audio 1" and

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:27.566891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:24.233956Z digest=sha256:a4cbde4d02fe4a74b7d15af17ba1d989986249ad339e7b3f1b523e04ca905717

Observation bc29348e-d938-487c-bdb5-4faa341c8391 · outbound

This paper cites an unresolved cited work.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:27.301734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:24.359674Z digest=sha256:55b98c44c361f668bf5a6728995e55fe6234a718ee3c0f37693ee2e6364e6008

Observation faca40a1-ebba-4799-a019-4c2fc8c16c09 · outbound

This paper cites an unresolved cited work.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:27.093620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:24.480897Z digest=sha256:1629f50ab68b607d7d1ecd407f4d05d7701ae4992c66ff80d32dbd5e5f00d5d2

Observation 4af1368f-b726-4c67-8405-54479b81822a · outbound

This paper cites an unresolved cited work.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:26.837349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T15:48:24.646239Z digest=sha256:97f16d56836b2c30a892816d5a9767bcdb626705fe4e406e26286472c27b4940

Pith citing papers

Observation ec6fbd44-3fa9-4e9a-be4e-7a90b42d0953 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

Reference 263

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.100163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:ab42a66f84d8796a398a9667e0e474f6bfcc9a6552575a7c548f4cb0b0532221