Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:34.291292Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2509.00685.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:34.291292Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:34.127489Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:39:03.102169Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c1474778-faef-41c8-8129-2ccfac66fa38 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech LM-based TTS systems convert speech waveforms into sequences of discrete tokens using neural audio codecs [1, 2, 3, 4, 5] and operate in a discrete space [6, 7]
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38e3ca5f-8b4e-4350-8104-c02830c7d831 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Preference Alignment Preference alignment is often formatted as a reinforcement learning problem
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 96266f2d-f5fc-4393-93df-12378db3151d · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech MPO involves constructing a multidi- mensional preference dataset and incorporating additional reg- ularization during training to prevent model degradation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a2a5532-4b7e-4db1-b97d-9796034dec0d · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ef14480-6bc2-4ddc-bf8e-f8dc453829ba · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 801c9947-2627-41c8-acd0-93c70e48918f · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77f167f-eacc-4b15-befe-d7600d999574 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cc4a7448-e733-4358-ac93-788c0a435def · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7474ce0d-c671-43e8-ba9e-1894bebeb606 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Soundstream: An end-to-end neural audio codec,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d467e4ce-079c-4266-8cfc-768027922cbc · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech High fidelity neural audio compression,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a34a4908-4f10-4779-af63-17074227dd82 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Funcodec: A funda- mental, reproducible and integrable open-source toolkit for neural speech codec,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aef01163-c347-47e7-98ce-39d5fe2a0876 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Speechtok- enizer: Unified speech tokenizer for speech language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e8a331b1-aeb0-4d71-9d8b-05979977370d · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Single-codec: Single-codebook speech codec towards high-performance speech generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09c8a629-f739-4481-b3c8-32a1a3d5b399 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ec8996-46a1-4a65-8e85-d84b6076b1aa · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Simpo: Simple preference opti- mization with a reference-free reward,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a4029945-90a5-4bdd-bf1a-6fef626d8fd9 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e241448-ffe0-4bc6-a5e5-a40fb9a69705 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech V oice- craft: Zero-shot speech editing and text-to-speech in the wild,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f8d55cdd-b927-4d12-b893-ac9fee5a77e7 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d6e5a8-229c-4010-ab08-c9020fe60d4e · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ccaec03-6510-4a97-9f6b-eec943eeb06d · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Training language models to follow instructions with human feedback,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 605e7fe9-3802-4c3f-a11e-f5307488ef54 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Model alignment as prospect theoretic optimization,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dbc466ba-c57f-495e-83fb-6dc2ed89498c · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Rank analysis of incomplete block designs: I. the method of paired comparisons,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5981849-b668-4319-8159-8d5f58ab9545 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c8dba464-44dd-4ca3-a11e-fea4c77780e1 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Direct preference optimization: Your language model is secretly a reward model,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a986b73c-56d6-4262-83eb-83801e5a483d · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech This fine-tuned model serves as the baseline for our experiments
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d36b11b8-b95d-4739-8938-71b33d34de74 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Speechalign: Aligning speech generation to human preferences,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4cf35e4-5319-4912-bffc-3276d6180c0d · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f7e82a-5545-42ae-8daf-d8ba21770135 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d9a4cfd-7c26-4547-84b5-c8a37e367c2a · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Dynamic time warping is employed to align the generated and reference speech features of different sequential lengths, following the evaluation script in ESPnet [29]
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ecaaff75-c909-468a-8a05-64daf778a292 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Preference Alignment Improves Language Model-Based TTS
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8134630d-ed0f-48c1-b116-42f1f1a20d78 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Preference Optimization through Reward Model Distillation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f67ff59-a87c-4426-aeda-8256a16813fe · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Libriheavy: A 50, 000 hours ASR corpus with punctuation casing and context,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26dd2f97-3dae-4597-922b-cb06a62ab168 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech The ISCSLP 2024 con- versational voice clone (covoc) challenge: Tasks, results and find- ings,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8aa513c0-2240-4aff-a744-0d9fe6ed5c15 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech LLaMA: Open and Efficient Foundation Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfd8375-19f8-4303-a044-ce387eba3bae · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 16509b4a-4dd7-4961-a565-588a3a90e487 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Large-scale self-supervised speech representation learning for automatic speaker verification,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aaa69b3d-f608-476b-b123-6fa3bc4dbe28 · outbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech ESPnet2-TTS: Extending the Edge of TTS Research
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4029945-90a5-4bdd-bf1a-6fef626d8fd9 · inbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce65c129-29fa-4979-8e1e-73b6d3e69823 · inbound
DDPO-VC: Speaker De-Identification via Diffusion Denoising Policy Optimization MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.