Pith. sign in

Paper Citation Record · LEDGER

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2509.00685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00685 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:34.291292Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:34.127489Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:39:03.102169Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1474778-faef-41c8-8129-2ccfac66fa38 · outbound

This paper cites LM-based TTS systems convert speech waveforms into sequences of discrete tokens using neural audio codecs [1, 2, 3, 4, 5] and operate in a discrete space [6, 7].

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech LM-based TTS systems convert speech waveforms into sequences of discrete tokens using neural audio codecs [1, 2, 3, 4, 5] and operate in a discrete space [6, 7]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.888239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.122679Z digest=sha256:c98b3ba6a4b9f521f8e6d1c371cd1f747d0997865090aab2344cf58c1e60c97c

Observation 38e3ca5f-8b4e-4350-8104-c02830c7d831 · outbound

This paper cites Preference Alignment Preference alignment is often formatted as a reinforcement learning problem.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Preference Alignment Preference alignment is often formatted as a reinforcement learning problem

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.872374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.132737Z digest=sha256:759a35393c2cee79b5636428a1a72545641b610722c16da0b6e0b4442b0a3eb1

Observation 96266f2d-f5fc-4393-93df-12378db3151d · outbound

This paper cites MPO involves constructing a multidi- mensional preference dataset and incorporating additional reg- ularization during training to prevent model degradation.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech MPO involves constructing a multidi- mensional preference dataset and incorporating additional reg- ularization during training to prevent model degradation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.857606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.137462Z digest=sha256:addce40c07b33f79f7dd04377563d42b09e8b60affc8adaf5da18fee5b4ad721

Observation 8a2a5532-4b7e-4db1-b97d-9796034dec0d · outbound

This paper cites an unresolved cited work.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T13:24:34.843343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.142629Z digest=sha256:8ca09ce879636fed628bd1cb8ea98cb0c6a22eb3121bd549b653547065959b14

Observation 4ef14480-6bc2-4ddc-bf8e-f8dc453829ba · outbound

This paper cites an unresolved cited work.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T13:24:34.801594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.156929Z digest=sha256:69fa8c97aa69ebd2f4bce069d03f80020fabc0ddf25c8f73a1bf2acbad02d700

Observation 801c9947-2627-41c8-acd0-93c70e48918f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.185099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.185099Z digest=sha256:d9c08f522c276d122ed350f0bfcab88c53cd97ddda9fc21f40f12023cf579029

Observation d77f167f-eacc-4b15-befe-d7600d999574 · outbound

This paper cites The Interspeech 2024 Challenge on Speech Processing Using Discrete Units.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech The Interspeech 2024 Challenge on Speech Processing Using Discrete Units

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:24:34.500670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.189566Z digest=sha256:223365d1c685050377455c41d8603c18927685e4632e298a8026537f421ab92f

Observation cc4a7448-e733-4358-ac93-788c0a435def · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.194579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.194579Z digest=sha256:6f004ed1e0dd8251c9566138bf0f1550b88c7919489d7cc67954a238e373386d

Observation 7474ce0d-c671-43e8-ba9e-1894bebeb606 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Soundstream: An end-to-end neural audio codec,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.787102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.162131Z digest=sha256:a73b00381006fe766106236a36b92e675f3dd84ed7e353318aa1b25f40b0b3f4

Observation d467e4ce-079c-4266-8cfc-768027922cbc · outbound

This paper cites High fidelity neural audio compression,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech High fidelity neural audio compression,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.772980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.166865Z digest=sha256:ead9b14f0c9b4ed50dd7190ec62d49c6e533f40d1332702373377350d9fe814f

Observation a34a4908-4f10-4779-af63-17074227dd82 · outbound

This paper cites Funcodec: A funda- mental, reproducible and integrable open-source toolkit for neural speech codec,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Funcodec: A funda- mental, reproducible and integrable open-source toolkit for neural speech codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.758699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.171218Z digest=sha256:74efebccec7ff8d440897d9e9faeccd5db1635ed1a7f7ddd45ccac0fe1ab2377

Observation aef01163-c347-47e7-98ce-39d5fe2a0876 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.740531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.175646Z digest=sha256:19e39d42b396ad9c2cb6cb95c83c33636764c8daa56fff2d4d4be565f5b3b89a

Observation e8a331b1-aeb0-4d71-9d8b-05979977370d · outbound

This paper cites Single-codec: Single-codebook speech codec towards high-performance speech generation,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Single-codec: Single-codebook speech codec towards high-performance speech generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.725039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.180968Z digest=sha256:1d000bbc6b858ac7b6e533297dae1316290933f4e8940d0574c69206b137f470

Observation 09c8a629-f739-4481-b3c8-32a1a3d5b399 · outbound

This paper cites GPT-4 Technical Report.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.223367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.223367Z digest=sha256:7daf0e5defb77942a3f1d7c4079791e37da1a059fea55fe6708fcbf0947053eb

Observation 83ec8996-46a1-4a65-8e85-d84b6076b1aa · outbound

This paper cites Simpo: Simple preference opti- mization with a reference-free reward,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Simpo: Simple preference opti- mization with a reference-free reward,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.664942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.227798Z digest=sha256:bfec97ff6d6357adf9520d5162cb99871db200a83b13a1c42b51de90232d3444

Observation a4029945-90a5-4bdd-bf1a-6fef626d8fd9 · outbound

This paper cites MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.127489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.127489Z digest=sha256:3f4d2d3cf1b45b66f22a7df28f5a32c2b26a32ae0eaa934063040883c5c85857

Observation 5e241448-ffe0-4bc6-a5e5-a40fb9a69705 · outbound

This paper cites V oice- craft: Zero-shot speech editing and text-to-speech in the wild,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech V oice- craft: Zero-shot speech editing and text-to-speech in the wild,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.709223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.199300Z digest=sha256:a18d592ec89128b20db81f8a3b080856180bf9b64dec544abb6c47900ea1b48e

Observation f8d55cdd-b927-4d12-b893-ac9fee5a77e7 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.204193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.204193Z digest=sha256:52e16aa106ade5b888f8945a86869c1e699c6b78617a80acb6d4f3ed5a5a3605

Observation c9d6e5a8-229c-4010-ab08-c9020fe60d4e · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.208839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.208839Z digest=sha256:ee381e7d4c7bdb86ac76797280706e3725ece79134e0dcaaf86588bb8b648a75

Observation 2ccaec03-6510-4a97-9f6b-eec943eeb06d · outbound

This paper cites Training language models to follow instructions with human feedback,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Training language models to follow instructions with human feedback,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.694020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.213556Z digest=sha256:02aa6c32148e606521cc0f355854ace66a1bf15e6a49f4f3c2808c062f1bca22

Observation 605e7fe9-3802-4c3f-a11e-f5307488ef54 · outbound

This paper cites Model alignment as prospect theoretic optimization,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Model alignment as prospect theoretic optimization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.679116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.218517Z digest=sha256:ceb20605442bee04fc371b98679c6b8d3040bf9096ef82391b61c151e8670e6b

Observation dbc466ba-c57f-495e-83fb-6dc2ed89498c · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.259985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.259985Z digest=sha256:baf016ed7b959135d38e9bc1d11c2856184f92d25c89e7ce0d8debc020abce4d

Observation c5981849-b668-4319-8159-8d5f58ab9545 · outbound

This paper cites Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.611695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.264903Z digest=sha256:44dde46e4b218eeeb3ed46d1ec8abd1289b9757bd7ffd044a1f296bbbc0f92ae

Observation c8dba464-44dd-4ca3-a11e-fea4c77780e1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Direct preference optimization: Your language model is secretly a reward model,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.648710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.232766Z digest=sha256:d001ffc4a0fca4fd33d14eb6e1364a73b4c11843146cfe0625a81b1c042e4276

Observation a986b73c-56d6-4262-83eb-83801e5a483d · outbound

This paper cites This fine-tuned model serves as the baseline for our experiments.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech This fine-tuned model serves as the baseline for our experiments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.829916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.147812Z digest=sha256:a98d6fca505ada1e696d6530ae24ae916e90c8319a7d0739955a40c1b8ee0988

Observation d36b11b8-b95d-4739-8938-71b33d34de74 · outbound

This paper cites Speechalign: Aligning speech generation to human preferences,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Speechalign: Aligning speech generation to human preferences,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.237058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.237058Z digest=sha256:9637c4a18c750f082223d1c5d7ca88351aaa1a1de2ca5d9bc076450f79468fcf

Observation d4cf35e4-5319-4912-bffc-3276d6180c0d · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.241314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.241314Z digest=sha256:f9c3eb5603418e741374bdf917d54e047b46cd9ff5d767647de15a6dc7b73557

Observation 35f7e82a-5545-42ae-8daf-d8ba21770135 · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.245816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.245816Z digest=sha256:6b3945de0f54c4a1ed5b545433635b8929e2501e8ce8afab935a8439c62559a3

Observation 0d9a4cfd-7c26-4547-84b5-c8a37e367c2a · outbound

This paper cites Dynamic time warping is employed to align the generated and reference speech features of different sequential lengths, following the evaluation script in ESPnet [29].

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Dynamic time warping is employed to align the generated and reference speech features of different sequential lengths, following the evaluation script in ESPnet [29]

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.816105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.152451Z digest=sha256:ce7ec31f2e60a4e8ba30ab2987218d0140ff5cfd9dd5cb46011e9dda104708e1

Observation ecaaff75-c909-468a-8a05-64daf778a292 · outbound

This paper cites Preference Alignment Improves Language Model-Based TTS.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Preference Alignment Improves Language Model-Based TTS

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.250279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.250279Z digest=sha256:c54fcf5c0f77be221df11be51f8080e9cbfbf9ff4389eb9af6dfceec731c7183

Observation 8134630d-ed0f-48c1-b116-42f1f1a20d78 · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Preference Optimization through Reward Model Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.255346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.255346Z digest=sha256:092d69fbad537b4a8b5e48b9d1b48e0cdd672862d62a5a2db768090dd2d50bfa

Observation 0f67ff59-a87c-4426-aeda-8256a16813fe · outbound

This paper cites Libriheavy: A 50, 000 hours ASR corpus with punctuation casing and context,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Libriheavy: A 50, 000 hours ASR corpus with punctuation casing and context,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.595842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.269278Z digest=sha256:2d50f47725091fdad392069462e65c88fd6e5b85082cde63a29c4f9db4607143

Observation 26dd2f97-3dae-4597-922b-cb06a62ab168 · outbound

This paper cites The ISCSLP 2024 con- versational voice clone (covoc) challenge: Tasks, results and find- ings,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech The ISCSLP 2024 con- versational voice clone (covoc) challenge: Tasks, results and find- ings,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.580999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.273751Z digest=sha256:5236716bbfe2c6d1ba15d7617aca9b5a5cce512a372964efaed25fe8fde5547f

Observation 8aa513c0-2240-4aff-a744-0d9fe6ed5c15 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.278020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.278020Z digest=sha256:504a1d3efeed359759d2d48d50e3726eecde9196836c36af32f45c5d5aad443f

Observation 9bfd8375-19f8-4303-a044-ce387eba3bae · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.565726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.282319Z digest=sha256:3e57125c355b7ef8b687c2a53658d5c6d16feb19f12161dacbd98d322a4861c2

Observation 16509b4a-4dd7-4961-a565-588a3a90e487 · outbound

This paper cites Large-scale self-supervised speech representation learning for automatic speaker verification,.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Large-scale self-supervised speech representation learning for automatic speaker verification,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:24:34.550700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T13:24:34.286950Z digest=sha256:67e7efeebd3ddc69905fe459b7c2ccbdfae47b883b561841fd5793833c4c02a3

Observation aaa69b3d-f608-476b-b123-6fa3bc4dbe28 · outbound

This paper cites ESPnet2-TTS: Extending the Edge of TTS Research.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech ESPnet2-TTS: Extending the Edge of TTS Research

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.291292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.291292Z digest=sha256:fb83a82c5c1d352069cb509fd58645f3b6f924f7628731aad7e11d1d30ef2c1d

Pith citing papers

Observation a4029945-90a5-4bdd-bf1a-6fef626d8fd9 · inbound

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech cites this paper.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.127489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.127489Z digest=sha256:3f4d2d3cf1b45b66f22a7df28f5a32c2b26a32ae0eaa934063040883c5c85857

Observation ce65c129-29fa-4979-8e1e-73b6d3e69823 · inbound

DDPO-VC: Speaker De-Identification via Diffusion Denoising Policy Optimization cites this paper.

DDPO-VC: Speaker De-Identification via Diffusion Denoising Policy Optimization MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.104994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T23:35:59.957825Z digest=sha256:75e1b9cc20a0b8cd700eab1976d14b0ce60ed7dc093d50dacc29925ae32b6f6f