Pith. sign in

Paper Citation Record · LEDGER

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

As of 17 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.14988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14988 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:24.646239Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T19:19:36.427337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T19:21:48.097349Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact5
  • verified fuzzy16
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6bf2c79-2fd0-478e-99f2-f47e54e9c85f · outbound

This paper cites Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Naturalspeech: End-to-end text-to-speech synthesis with human-level quality

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:31.639139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:19.733560Z digest=sha256:281fee02b776745d958dfed307dccf3adc9b938b8a710ef9f9ab3e9ea6bf62cc

Observation 0a90a7c8-d495-4e6d-ba73-9f4d7631ddf4 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:31.364146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:19.789105Z digest=sha256:edecea3af9e1d23ee8cb69ae166b31d46338f348650513c4e425d845399a6ffa

Observation 6790d590-3fb0-49b3-9ce4-c95510e3ae4f · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:19.866141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:19.866141Z digest=sha256:56fd24817bb22a9ccc0f3b8d8723e801aaff15c4206ada4d9141d743243cf465

Observation 85822a91-7e9e-4dcf-bcd2-6c853645dff4 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:19.939041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:19.939041Z digest=sha256:d09f4e59ff6212632d52a8c582a0d2a9cf1a973778f6048ddc8f768635327a28

Observation a6ef7c61-4fb3-480a-a09d-9f6dfc26b888 · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:19.994804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:19.994804Z digest=sha256:d8f878e41923f9492a5f854f940c647ae94cd07c95bc6639f60629cd89806a0e

Observation bfcbbabe-5463-4063-8e98-a35ab3e4c91d · outbound

This paper cites Emo-dpo: Con- trollable emotional speech synthesis through direct preference optimization.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Emo-dpo: Con- trollable emotional speech synthesis through direct preference optimization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:31.081647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:20.065138Z digest=sha256:5816685fc51be72ef54da865887b6f928f53f14070ea1374eb016841122b8b3f

Observation cd0dc393-7380-4d9b-86b1-41727b768a4f · outbound

This paper cites Preference alignment improves language model-based tts.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Preference alignment improves language model-based tts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.861946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:20.157944Z digest=sha256:196e22f363fd12519e29a73ee3d114bbca7be32e13db2edac7cf7d75e1ec7063

Observation ca2867c0-d3af-4cd2-835f-d9016de33486 · outbound

This paper cites Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.224849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.224849Z digest=sha256:354089b8b6b694dd80911ac50decd67e44a496ebc024c2e1c58bf851fc27bc59

Observation 99b7574c-8814-4de6-bddd-7803eb02cd7a · outbound

This paper cites Evaluation of Best-of-N Sampling Strategies for Language Model Alignment.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Evaluation of Best-of-N Sampling Strategies for Language Model Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.282544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.282544Z digest=sha256:15991069cd613cfc6861eac252e2fef7a90b66e68d5f9d5c9bbf9c070696e0bc

Observation af530fdd-197d-40be-94d1-7c92f2be9d9a · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.353019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.353019Z digest=sha256:5150698e2635634b2f008400139f8c37863a7730155fd12d21d9918dc6b8e278

Observation 529d8da5-bbd3-4a32-acf6-9eec08dbacfa · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.428168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.428168Z digest=sha256:a9f23ebcac481cc8f2cfcb9238ac9fa8a3579d93f0537474c30bbf613448ccd4

Observation c8be1b90-e219-4d50-826e-65bebd9db008 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.513937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.513937Z digest=sha256:30987853b6e6c6aaf2d0267f0e1fa46e851ce88553858b61ba451e5a64823c6c

Observation b3cbe4af-0b4a-4535-b907-5547b8ff3d60 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.576312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.576312Z digest=sha256:ee266f3869fabaf2ab470cd348b882b80d3b0472f686e8c2c8721744496f6b35

Observation 5fe3355e-f1d1-41e3-862e-94f5a87894c6 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.644717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.644717Z digest=sha256:7899a133f9582a8876b3bf5d53255184f90914cbad489a68e788eaa1600bac0e

Observation 9b4ef9cb-c92d-4bf7-85b4-3b9d025929a2 · outbound

This paper cites Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.585461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:20.694493Z digest=sha256:539f56ef651cf00debefda3c6770faf187bba229e97c5752bfd9c753f276b54f

Observation d9c95b5b-4da3-4d66-b6fc-52980ababa41 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.763920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.763920Z digest=sha256:c19f7a445697cdedaaf35531121a604120954c555c4512edb846a3535621ac56

Observation 90256e37-c7f1-4c17-ae37-b363e025ac24 · outbound

This paper cites Autoregressive speech synthesis with next-distribution prediction.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive speech synthesis with next-distribution prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.825000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.825000Z digest=sha256:19e77910dcdb1eff3ce6ec586089e240e18a668a2e54cf0c2dfcaa3695ccf964

Observation e4ebd70e-7347-4e0b-897d-2364e7a2b092 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.918034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.918034Z digest=sha256:d270c01c482760afa811d089a183786df76b521cd832f40f6de17f46c233934c

Observation 4b5bb2d7-2828-4156-b580-53b0b7c92847 · outbound

This paper cites TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.978975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.978975Z digest=sha256:0f5e444344d0eb2b76482be9730d1fbd68e3fdfb63bb9bd99cda4f5a215e4f64

Observation 6ffaf0fb-3276-448b-a359-aea5a05b0e7f · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.077843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.077843Z digest=sha256:c9d7843b29d4a4d5c6bdcca06288bfe6fa66bc30bfb999349fe2c977711c6f33

Observation 99216364-feee-44e5-8711-18a96f4a7f90 · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis V oicebox: Text-guided multilin- gual universal speech generation at scale

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.331261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.126247Z digest=sha256:7dcde49808509e25e3473b4c84e87f125419648e5b7ad93fafe19a19281e44ae

Observation a96b3cfc-3d52-47d9-b621-0ef2835b38f1 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.184511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.184511Z digest=sha256:3204fd1271f8e66122f61ea486b6a5098d1f993c2dab0e716d0dd3310a4049f0

Observation 658aac14-d570-47f6-b978-9161804f311e · outbound

This paper cites StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:26.218116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.255061Z digest=sha256:76663c7b8ad5a43de7a6807971fed2054143f470ceec347ce3ff3c2df0309ec7

Observation 6a85a2f9-0209-44fd-b840-7bb86963c6e9 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.292062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.292062Z digest=sha256:4e6e098b968641d4341af4d4abcab3db6dfb5719c0905784713e6043009bef2f

Observation 622d2823-e56a-4b2f-b553-fbff7e28ba02 · outbound

This paper cites SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.980592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.342231Z digest=sha256:1c6cbc936bec66c9d939307b3ed48edca4d241bc9590f9be6132753b24707152

Observation b3eef67e-c8b2-4d44-a2e5-8e5a1da552c8 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.419026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.419026Z digest=sha256:3cf6b6959d0a17e77317af47726e5d3ed40b48911163e2cec230b16c3550301e

Observation 83257d1b-2c04-450a-a881-9b1e45af6008 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.515476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.515476Z digest=sha256:89116978b9d20a4003922df749c5f862273cbff0c3e7f6566fda666518815b84

Observation 5014061c-b215-40d4-b96a-38023f01d708 · outbound

This paper cites DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.669295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.591183Z digest=sha256:1cb937f5918ffbdf5c2106b0bebb82caa8278220a124355f6225a992bf4f7bd9

Observation c2d8318d-4d3b-409b-a075-ba6024e82095 · outbound

This paper cites One-step diffusion with distribution matching distillation.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis One-step diffusion with distribution matching distillation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:30.024351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.679831Z digest=sha256:84bbfb960c0bf63ae73236dc3be321eadac991586728c74486ee630dc9db241d

Observation cb28080d-6f40-4ad9-9895-eb0438988ac9 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:29.772307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.742937Z digest=sha256:80c22ff6152e463cfb67b5ca7fa1e5bdd18b7e893abba6fa8fe55ccf2bf98544

Observation 0784fa1d-a261-4485-823e-b268b951316b · outbound

This paper cites SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.406941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.812155Z digest=sha256:2252fad132dc1d14d81809737164d02d070035b37febba2e86883344c30e667f

Observation 29eec92c-be00-43ab-9150-773dc27e258b · outbound

This paper cites AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:25.182680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.866669Z digest=sha256:e351c228ae1fd9eab000d61e6134b6158d8e99fea0032dfc78e116c1037647d6

Observation a5878af4-9419-4991-9343-afa0a734da01 · outbound

This paper cites Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:29.499117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.938505Z digest=sha256:5036a4dc367b187ce0321f52b7a2d85ac41f4a8b60a8b5b014a877aabccaab81

Observation 30f650b0-ab7f-4b22-aef8-e708754981bc · outbound

This paper cites Meta-stylespeech: Multi- speaker adaptive text-to-speech generation.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Meta-stylespeech: Multi- speaker adaptive text-to-speech generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:29.215543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:21.998568Z digest=sha256:045e546f911de1e6c8a7e616923d982ed0c6c81e0eee24e28560eded7407ec57

Observation c4d5860e-7b97-4dcc-9c27-e2996a404932 · outbound

This paper cites StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.067217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.067217Z digest=sha256:824c48f74d94a18974f33ed6c01492a39706a3074adf61846a13de971c27dac2

Observation d0a08099-a6fe-422f-88f9-48959baf15ae · outbound

This paper cites NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.157083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.157083Z digest=sha256:3a68bb00a91a6604adfe7066f2d8489b4fa882e492e6968235d930059ecfa69d

Observation 43874f57-cbcf-41ca-b2b2-6953c5ccbe0c · outbound

This paper cites DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.252443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.252443Z digest=sha256:54b4d9a0e6b409464cc4116419b96a55aa7dd8fe616b469e6326b4c2e08aa8b7

Observation 3de44fb7-35b5-40b1-8310-b586fe374102 · outbound

This paper cites F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.347437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.347437Z digest=sha256:90046a61ad428de39f7111ef01c113b7cb649f042f4f96ab0283a8cb41198101

Observation 676782b7-0ac9-4013-aa6c-2f897b0f37d4 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.435245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.435245Z digest=sha256:4734f8312a013cc259f9a66fd27ee34942a86ef44bf86f1fa79738e5f8238bb1

Observation 04e21841-37ab-4286-b335-9f8e8d6b4e2d · outbound

This paper cites Improved Distribution Matching Distillation for Fast Image Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Improved Distribution Matching Distillation for Fast Image Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.532258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.532258Z digest=sha256:2ba988029918e19dd191f0fc515634b88b25b8839513f69a8fe189eb68db164e

Observation b2280739-c13a-4d94-b749-3a0dfca648b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.624809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.624809Z digest=sha256:4b22f436f0b35a81e592353b35fd100c8f520cb17ef4472091322951dfe7de08

Observation b0826a8e-d10e-490f-883c-d8a246018847 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.709549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.709549Z digest=sha256:0369637ba30e97c7391af5a94a0f588334b3cabf368fbebe647217363ccbfb97

Observation 93fe65ff-4cfc-4fb0-8797-154f58472b5f · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.966603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:22.772891Z digest=sha256:316f18a50da2f96d1c022855a6fab4d5d5b9c427b06d36908b88fd1b0942e9ab

Observation 49409248-3195-4e3d-b17a-5e25c7bb2263 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.884960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.884960Z digest=sha256:f9c9be90cb876c781e5e3d1d05eb160d9612536d72b21a62acb671f5ef4a11f9

Observation 1c8bd800-123e-4fda-8558-b0dd1ccf24ed · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.974182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.974182Z digest=sha256:26ee2cfd8b5b3bef5a0d3f5cf7c4be3cc4b8331230a7e1894176c530b38ea605

Observation ebc01169-ea41-4a73-8179-04fa94037182 · outbound

This paper cites Didispeech: A large scale mandarin speech corpus.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Didispeech: A large scale mandarin speech corpus

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.666445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:23.060690Z digest=sha256:194cebf128d3f8cdc0c62a334f2faeec349a01aa5e064c338104d90320ca9a85

Observation 606d3b08-2ae2-41f0-9ea0-a39daca8ac6c · outbound

This paper cites Fixing Weight Decay Regularization in Adam, 2018.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Fixing Weight Decay Regularization in Adam, 2018

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.380022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:23.185093Z digest=sha256:dc2dd0e3bee732179f6d57405e4fba9a7b75cddee802e3aa06f33c5d97d7d86e

Observation 82368ccc-fcac-4ca8-b7f8-17edf5268a33 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Robust speech recognition via large-scale weak supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.288864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.288864Z digest=sha256:e9dffb5a72db3dc47e0ab0ad27fba8dc68968aaa9b1d7b80b646b5627cc9587c

Observation 5bf1f8fd-29cf-44d3-b2d5-a4c8a7b6e62e · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.415106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.415106Z digest=sha256:633f6e6858eee02c6ccf6a18b2939f5e0b00a7ef5141a56146edf498a09a6897

Observation 77cee763-0d59-4476-95e4-694a2bb62951 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Wavlm: Large-scale self-supervised pre- training for full stack speech processing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:28.170855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:23.616993Z digest=sha256:f1d8515cdda421b31bd70a9e55d19b3b870fe3b838bb3a84fdf0a81533dd9fe0

Observation 6f572e7a-fc9e-4d19-beec-832cc732b043 · outbound

This paper cites Denoising diffusion probabilistic models.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Denoising diffusion probabilistic models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.719566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.719566Z digest=sha256:422da6baa1a2d9d599e0a47b23ff09f63d7f7a6afd1a901119ff2473328b6aa6

Observation f3027e3a-81de-4530-a0a3-088acc48b096 · outbound

This paper cites Murphy, and Tim Salimans.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Murphy, and Tim Salimans

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.802869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.802869Z digest=sha256:768dcf61ad87cdab429c13c5912de4b83374ebbfb4e58b9612b05eda8d44d3d0

Observation 9a8e6f38-969d-4cdd-bd82-df9d885b2315 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.922139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.922139Z digest=sha256:6933fc705b4320c2c19e22376355063b8439cb4912d2c83b70c9b599693f5a1a

Observation eee60b90-8877-4d7e-b9bd-fb896a34889c · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:27.875223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:24.092300Z digest=sha256:f81b8d3f1c718dadf41d51c93a9c232329e0c0ed69e778036f962e021b6add20

Observation 28fe12df-87a1-4804-ae77-4c38a942c1c5 · outbound

This paper cites Audio 1" and.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Audio 1" and

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:27.566891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:24.233956Z digest=sha256:c374ab1169bc443eead5a429b00464e7a5f01a91162d6f5146c6cc11640b70ac

Observation bc29348e-d938-487c-bdb5-4faa341c8391 · outbound

This paper cites an unresolved cited work.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:27.301734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:24.359674Z digest=sha256:4167e98ae0d13f2bb7b4623b5f036333cd1a0799351dfdc88ce9035b3019bd0b

Observation faca40a1-ebba-4799-a019-4c2fc8c16c09 · outbound

This paper cites an unresolved cited work.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:27.093620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:24.480897Z digest=sha256:ebe100bef3bd72afaba63920e128116b5163f0d73b41ffa87d747900b5534e20

Observation 4af1368f-b726-4c67-8405-54479b81822a · outbound

This paper cites an unresolved cited work.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:48:26.837349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:48:24.646239Z digest=sha256:a87a3814c1a1f35856e1129d170baca0024526f364c20403a074a8be0218c959

Pith citing papers

Observation ec6fbd44-3fa9-4e9a-be4e-7a90b42d0953 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

Reference 263

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.100163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:79cd0166674a857a74512df5d636a0708373e789cd7890d3c5a007b21c29150e