Pith. sign in

Paper Citation Record · LEDGER

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

As of 19 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2505.04113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04113 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:42:53.805578Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T03:50:26.873406Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1a7df6e4-fbf8-4191-8a8a-c79b921f0ec1 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.382881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.382881Z digest=sha256:f07b94c5a56fdde3438817f06cf06f21658ea02be9d5be02ac0e82bacf3cd9a9

Observation 4e61aeb9-fa24-4000-8acd-1f5de6c1e587 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.429120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.429120Z digest=sha256:7cf5126958100da6d987a5e1ffbaf3a81df46b856dc71cdfb1b37982fc60c4ca

Observation 0fec1e81-ec53-4289-8fb0-9aeafa1c6848 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.435184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.435184Z digest=sha256:60cd48134e5cb2f42714a0f28630b44466311013ec7eb296aba247fe59922fc4

Observation c054b9b9-5f4f-40f6-8551-940fe0138671 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.440665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.440665Z digest=sha256:5246a407ee96e6cdfc4ff04d82e3e674e6e2dde6dd6bf13c6c43d3ef56a73c43

Observation 4a136a66-6ecf-46cf-aa10-4af35ee3b71c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.090896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.446506Z digest=sha256:4d817aa958f4b2715275de9a56fa299e0626c90ffde875ec35ca373726fbd3dc

Observation 8a4fdccd-eb9b-4ad5-81a9-c3a33c7f376e · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment SoundStorm: Efficient Parallel Audio Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.451504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.451504Z digest=sha256:869396b9319631c94260eca20d7f03fa51db62597e7540a245a0030a2adfefe9

Observation 8d8b4098-e038-4e14-9b6a-916c39232cf1 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.074759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.457377Z digest=sha256:cb459dac5e0310dbab7e184620b46ad4e2397f6757fab838bd805676eae2f200

Observation 08898a01-1af3-4201-a8ae-8b94ada6e9b9 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.462636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.462636Z digest=sha256:bbe83f7b6ba2296d1192e0cb95834d7867362cbbf58b7f1bbe5a64db4a9328f7

Observation 550ed16d-b6da-440e-87b5-8eb46b30387d · outbound

This paper cites DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.468382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.468382Z digest=sha256:14a33208c821436a2e5bb0edb4ccbcadd455f3f9ce7f891bea908deafb3f40aa

Observation 3f7da104-808d-4685-a7f4-e6cbb78ac27c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.473616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.473616Z digest=sha256:7c7c98d3dd06133e64f66333ee97b5ea51fc7e955fe4513e7d046e08cd262f82

Observation 0b384649-ebfd-481b-aa21-a76293ff8383 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.478975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.478975Z digest=sha256:9f91cb2ea21cfae0577e66ffe051d433da4e72c44dd8fee755b3696ddcec2fbf

Observation 9c450d1e-8039-4c40-a742-23278d67cec3 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.047622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.484988Z digest=sha256:da970bda1bfaa7514220b16aec159d5fab177172f3769ca7297b7593a5e2d419

Observation 4193e79f-63bf-4d06-b964-160b40e1f1d4 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.489615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.489615Z digest=sha256:87c33d7d9b5bcdcbeeaa53413f620df0a5ab24a0790cd66a81b828d8179acb66

Observation adaeb182-26eb-4d52-8343-0b7745adbe83 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.020747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.508501Z digest=sha256:c578c05fa8479027c7f632d371a68919e8acc9defe49d3d7cddc20ae3da16b83

Observation 227841fa-f122-4f82-a3c6-48375cb356ea · outbound

This paper cites DeepSeek-V3 Technical Report.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.515086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.515086Z digest=sha256:7cbe21a259b836a4d5fe5429e09ec7e1d3fe38066e5f44088891a581156a4edc

Observation bc12a36a-9d15-4775-ba5b-8aec19a18020 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:55.004563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.519735Z digest=sha256:97c889f531cd3f84b55a6856b6d927c6f61ee4af0aa72f513243f2f9633c4cbe

Observation 4b9d0206-f088-47d9-b4de-4fc0bb4b75d1 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.525148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.525148Z digest=sha256:effed821239ed5be37c3fc9de9709b50336066f41f0b05d9a2ccd234bf93266d

Observation b974ab77-f39f-4698-aa0a-3819e789bd47 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.531276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.531276Z digest=sha256:9c21d38b101a836a66d60610d6f56c7338fef4b74b2fdedd8a973e43faa1ebdb

Observation 548bbc03-0ef8-4a2a-8ffa-0f6d91197ae8 · outbound

This paper cites The Llama 3 Herd of Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.536759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.536759Z digest=sha256:7423478163f2e337cf56778f02c3c730797d8af00bcffbf4ea9891fea7d8f6d3

Observation cabe3b75-9f3d-4e20-b3e4-bf3f270eb961 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.989490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.542162Z digest=sha256:3f8bffca875572bda3b993f3924ba33a31864f165cb4fa2971d821d01e5dfe81

Observation 8dd4a8ed-5495-406d-a597-ff41b1f19ef2 · outbound

This paper cites TLDR: Token-Level Detective Reward Model for Large Vision Language Models.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment TLDR: Token-Level Detective Reward Model for Large Vision Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.547349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.547349Z digest=sha256:fba4769e5e4577c9650a1627b251ac2d1ff4d404a529c52ea30eedd35c23e0ca

Observation b4a844c4-906c-45a7-835b-d867ddb0714e · outbound

This paper cites Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.552909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.552909Z digest=sha256:d2c9b31b8dc0e95c24f3a0a072d00bf363e85a804fa68383cac8628af1831341

Observation c374b99a-d696-4c37-92a7-28badd20c2da · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.973191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.559248Z digest=sha256:1b0c88e3dae1e93e83496b4d1081199c60cbe1524a829d98e98b5cfd1ed529d3

Observation f7d061ae-9426-4c05-9c45-c5d549e87bed · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.957150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.564993Z digest=sha256:9ffa88a5dd65124ca9b1373c8de185644ec8a6be3044d654ba44fbe3f83bbe29

Observation 10f9e75e-9f22-4a58-99ac-41179a7d52cd · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.570692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.570692Z digest=sha256:c56b045d9c271d87d431305dc6ecbd44edd1311fe746335aef7e5ecc580c9525

Observation fd9dc5b2-1d60-4564-a35a-0034003851e0 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.941613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.576247Z digest=sha256:1fd5444b47ebd45d1c689c0e7f50cd44a4095a7e1a49881627a57bd76568eee0

Observation f3450993-6148-4fad-91e6-89f8786619ea · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.925550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.581917Z digest=sha256:ae398cd8cffff52f7b85b0f3dcd5e35c763f3f981b478c6f373ea2a6b216e36b

Observation 400dfd54-299e-4dba-825a-95e6c4df6a19 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.587276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.587276Z digest=sha256:9636adad1a6c17d4ba6136c0e70579b050923645a44c6f83361b52244a684852

Observation 089ac67b-246b-4415-813f-8e7d7cc319da · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.592249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.592249Z digest=sha256:ab7e353826aebcb48e4ba48233bebf157267af689bd7dd4004023a279e21a134

Observation f49712de-77ee-4112-ba43-b53ef2517a5e · outbound

This paper cites Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.597794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.597794Z digest=sha256:0b2ca5b92a706c157c3983ad8ea7a8346c771e76b6ad74a651eda28ccb239401

Observation d3b2ccae-7020-4ddd-acfd-201b60c0d5ef · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.909712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.602750Z digest=sha256:c10ca81dc685d4b6f719174177776cbc0b05e782d1cc115b774a44fdafebb57f

Observation e9c17c1b-7f1f-4802-a957-c16d7ebf814f · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.893366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.607183Z digest=sha256:03a87b3908feba60d360c7fc66bd41f78fba7d964f1cbb8c504f4d1704b1a652

Observation 03548c78-377f-4458-afc5-2db73b8d1b13 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.876839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.612006Z digest=sha256:193cd13cba1c0d57a1496fa751a1a76993096c671a3e528bfbd576783921a4fb

Observation bdb722d8-41d7-492d-acb8-b1e44dd4a63c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.859191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.616509Z digest=sha256:eb237515186d85013155af30123e5c2e18db2b105dd7c7a340fee3c033fb84d0

Observation e2abf169-c2ff-415a-a94e-9440f8ff8bed · outbound

This paper cites Overview of the Amphion Toolkit (v0.2).

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Overview of the Amphion Toolkit (v0.2)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.621398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.621398Z digest=sha256:edb4b365967dcf60951b35d83db5d281de3126eaaf9d5a6dee9fe231ac80bc70

Observation c4d66328-e242-42d8-bfc4-c8d92eee597e · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.842209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.627729Z digest=sha256:ef947fef8312b4127507649060c8993f5151fa7f63c879839b92034ffef1db52

Observation feb0c688-75ec-4004-9c61-584a25c57995 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.823771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.633521Z digest=sha256:956e6683484cd826f5ae6c37f7f4d28ca1289184b541b1cc7b115b4fa89d3ea3

Observation a2782b34-7f3c-4f5c-aaf6-ea0c4c459a6e · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.807908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.639014Z digest=sha256:43c89727c63908d36d00cecfa47b67a8129abe30a7e06114e1bbfa8fe5393527

Observation 90679bdc-6143-4dc5-863f-534d97529010 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.791659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.644218Z digest=sha256:ff4e3c39cc6b8c56eede48f01ce6f2e7b4479705887f9ceb9809f20858a0c52c

Observation 2e67078e-9dbf-48ae-b7fc-11c2a8cf29d0 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.775959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.649412Z digest=sha256:9810116bbe660c5af1877b5395ba64a69103e7a9ee88b7f94a741a786b45fccd

Observation fe585f51-30b8-4557-ac6b-81cd3419d739 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.758884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.654555Z digest=sha256:22335378f045b4465c2eaac9111906a233f3ef6dac2cb77c1c838265d96e01f2

Observation 39689887-aefa-408b-9359-f2cfac8868eb · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.659510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.659510Z digest=sha256:d73e0a9d5ca7a24cb8d0620dd7063b51c512848c70a70ab5c583b693e9b72908

Observation c96d2b90-d07b-408b-80bd-4d0e5ec3e0af · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.664978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.664978Z digest=sha256:ac6455265ced4f647044c6df41ce748b0dca041adec6d2f614c5b22b629c2669

Observation 479152ae-c86a-4608-990d-c04855847f0b · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.670578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.670578Z digest=sha256:6a964f036d2ef3274275fcf72cf2748d8e02817a96df84ee2217f99560002d6d

Observation 440dfa31-0be3-409f-b196-f0a88836e22b · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.677136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.677136Z digest=sha256:80cd15add59bf97381c5564d5c0a50fe3c77d2a4fad1c9c034b8b85d6c545e12

Observation e7219b7a-c1e8-4a69-bf7a-79d6ed03e093 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Manning, Stefano Ermon, and Chelsea Finn

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.682470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.682470Z digest=sha256:e0a343ea82a2973edff86caf3bb2c899832f505e669edfdb8b0c85d0ee6ae92f

Observation a4a24be6-86ce-4818-bf2e-2667107f0043 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.698960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.687405Z digest=sha256:bf99e05423ff516782465fdb0ade1d7665c992d26e40925e045c408110ad5b9d

Observation 7b849058-daaf-4820-93e2-2be380a934c7 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.680850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.692676Z digest=sha256:e43ae7af03886d44ebfb360e6821ec718078252379421bcc745d842f74938ae1

Observation e6a42e10-03c9-4faf-8a9e-96841871c8e9 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.663032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.697796Z digest=sha256:cf96c5c93e3198e88ad4993eaf77f15e53bacc7e6bb083d7cea69f0aac21b6c0

Observation 4778406c-3704-4c23-a49d-c691c86f6bb8 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.647101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.702994Z digest=sha256:827cab77e1e1b314eb99a8d292a823357d1d820769297984d779b7aad97d22bc

Observation ac7cc6b3-b93d-415c-ae03-2e4ea62851af · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.631142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.708270Z digest=sha256:9c4ba17e2d2713427639edd95c75fb8edd01a590d6167879b4c881818c453d34

Observation 1f5076a4-30fc-41fd-9eb9-160db611f4c9 · outbound

This paper cites Preference Alignment Improves Language Model-Based TTS.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Preference Alignment Improves Language Model-Based TTS

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.713408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.713408Z digest=sha256:630e8635876d7b2a2a46c2efa4d42a01278d7d88eedfe7618fc79a878666a46c

Observation 5472b71c-30cc-4302-8aa7-a70f2d34de2a · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.615562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.718479Z digest=sha256:2a8e1044223d1c997aadc2a816b129318792a441ac20e3f804abea5281a0f89f

Observation f2993fc9-7472-4e96-87e7-c7adc9d5e8b2 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.723565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.723565Z digest=sha256:18dcd3e1fa1b13675eedd91d1226f2e587083a30f130593abd0ee24ae58a767e

Observation 07b4463b-47b6-46e3-89e5-055217e8dce1 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.599497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.729052Z digest=sha256:21c2eb264c234dc0c8be5e5ce2364676af377c2bb8846a825efb5d421ffd5fdd

Observation 448c2c86-a5cb-4783-bbc0-b612765e4522 · outbound

This paper cites Metis: A Foundation Speech Generation Model with Masked Generative Pre-training.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.733849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.733849Z digest=sha256:a2e647413a74fca348615ca82485c6993983116cf46f749bc3b29519597182da

Observation 8737ed2e-2e1e-4492-bac2-4446a732a009 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.739056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.739056Z digest=sha256:9a400dd85b8bd1ae9b3dafe5ccf36e8f110614f8582554c63372e2a0f5b799a0

Observation 2fdcfe42-5cb4-448a-961c-a2f136e1253a · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.573360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.744411Z digest=sha256:eee6ec332adea394a57ea07777b6b108502309f07616c0b8643838af8b96d341

Observation a01d4a64-ce51-4701-ba2b-85d3e5b7c7bc · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.557479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.749155Z digest=sha256:4ef53b77d163ba1f4849dc40535bd274cbf57516bf763027b456cc7b63458304

Observation eceea7f9-6814-4320-ad10-83d7f16dc911 · outbound

This paper cites Qwen2 Technical Report.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Qwen2 Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.754630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.754630Z digest=sha256:8e7fbc6588f60e198bb4d80c5bcc6a6834683ef37861948bc91c5e7cdfe2979f

Observation a2e45ab6-71ea-43e4-834e-d74a9a321df3 · outbound

This paper cites Qwen2.5 Technical Report.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.759449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.759449Z digest=sha256:4641857c1d79a2b28ac880c0b54ad4be2a63d5d610f6caa5e8bd0359c555a87b

Observation 2b37bf12-725d-486b-9c8b-11c6521e5cb4 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.764204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.764204Z digest=sha256:eb0bca8b2e16b148f8a799c8d5093fb9a5de8dd584e69e7308167a87ab9d6973

Observation 065e654a-1738-4839-844a-b7ec9c4832b8 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.768989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.768989Z digest=sha256:478c1ba4e82429b3885d842f19e3d6efba04713bd6e766ee3be8ee5fc222fbe9

Observation 6fc6d839-8263-4e84-a789-1b169e45c526 · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.531673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.774266Z digest=sha256:23c0896e1869cf57722418fe1e20940dc0afd7ef23868f135a5a55c7ea6a99f9

Observation 1bcea1b1-1a04-429e-a40c-5ebef4a0ce7f · outbound

This paper cites Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.779240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.779240Z digest=sha256:f3893008c20cca1239847197110f9926ba5efb95a48df6cb6de42a414bfe3578

Observation 2186a10a-f63c-4d32-8698-d5870c802a3d · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.515708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.784486Z digest=sha256:1ebb3694f8704483ab414c45b879b92b9c4e5b4faeb6090e9c2646dfcb9664fc

Observation c79f1747-c83a-41f8-b5a5-7e6cb985cebc · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.499871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.789633Z digest=sha256:2e5d51b38bc325c0afefecb119ecf72c9f8b21f389025a901f3d0c0ede66fdd3

Observation b80548c6-48ef-4bf0-84f1-aaadf282322c · outbound

This paper cites an unresolved cited work.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:42:54.483979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T23:42:53.795052Z digest=sha256:8df46f161b024b63099235daf590501251bbc4edf755db65d1853c45ad8432f9

Observation f78a5245-eab7-4d94-bb44-e49322bd21d1 · outbound

This paper cites online" 'onlinestring :=.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment online" 'onlinestring :=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.800202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.800202Z digest=sha256:13c82b6e72f59f2ca5d16b8655a57121a86f721cc5e832f507d63567201b0ec8

Observation 70d85a7e-1406-4e45-af06-e32a602b289d · outbound

This paper cites write newline.

Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment write newline

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T23:42:53.805578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:42:53.805578Z digest=sha256:df6ea7c3dd1ccd6c59086d4a9f11fccbcbcc08de26015a51d376e142765e9ae2

Pith citing papers

Observation b719acc5-8566-4fb8-a1e6-94edf174f7f8 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.542666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:eda50835cbc0651e3ff7082552bfc9ef78bb82cb1589de12d71e51653396d103

Observation d3e9bbc8-f5f1-4916-948d-b4831abe5629 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment

Reference 237

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.266206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:874cd0105b97223f75a7383effa6eaa255145d0ba8492dab9ce86699c04d4741