Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:42:53.805578Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2505.04113.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:42:53.805578Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T03:50:26.873406Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
70 of 70 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 1a7df6e4-fbf8-4191-8a8a-c79b921f0ec1 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e61aeb9-fa24-4000-8acd-1f5de6c1e587 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fec1e81-ec53-4289-8fb0-9aeafa1c6848 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Common Voice: A Massively-Multilingual Speech Corpus
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c054b9b9-5f4f-40f6-8551-940fe0138671 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a136a66-6ecf-46cf-aa10-4af35ee3b71c · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8a4fdccd-eb9b-4ad5-81a9-c3a33c7f376e · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment SoundStorm: Efficient Parallel Audio Generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8b4098-e038-4e14-9b6a-916c39232cf1 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 08898a01-1af3-4201-a8ae-8b94ada6e9b9 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550ed16d-b6da-440e-87b5-8eb46b30387d · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7da104-808d-4685-a7f4-e6cbb78ac27c · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b384649-ebfd-481b-aa21-a76293ff8383 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c450d1e-8039-4c40-a742-23278d67cec3 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4193e79f-63bf-4d06-b964-160b40e1f1d4 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adaeb182-26eb-4d52-8343-0b7745adbe83 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 227841fa-f122-4f82-a3c6-48375cb356ea · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment DeepSeek-V3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc12a36a-9d15-4775-ba5b-8aec19a18020 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4b9d0206-f088-47d9-b4de-4fc0bb4b75d1 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b974ab77-f39f-4698-aa0a-3819e789bd47 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 548bbc03-0ef8-4a2a-8ffa-0f6d91197ae8 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cabe3b75-9f3d-4e20-b3e4-bf3f270eb961 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8dd4a8ed-5495-406d-a597-ff41b1f19ef2 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment TLDR: Token-Level Detective Reward Model for Large Vision Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a844c4-906c-45a7-835b-d867ddb0714e · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c374b99a-d696-4c37-92a7-28badd20c2da · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f7d061ae-9426-4c05-9c45-c5d549e87bed · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 10f9e75e-9f22-4a58-99ac-41179a7d52cd · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd9dc5b2-1d60-4564-a35a-0034003851e0 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3450993-6148-4fad-91e6-89f8786619ea · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 400dfd54-299e-4dba-825a-95e6c4df6a19 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089ac67b-246b-4415-813f-8e7d7cc319da · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f49712de-77ee-4112-ba43-b53ef2517a5e · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b2ccae-7020-4ddd-acfd-201b60c0d5ef · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9c17c1b-7f1f-4802-a957-c16d7ebf814f · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 03548c78-377f-4458-afc5-2db73b8d1b13 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bdb722d8-41d7-492d-acb8-b1e44dd4a63c · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2abf169-c2ff-415a-a94e-9440f8ff8bed · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Overview of the Amphion Toolkit (v0.2)
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d66328-e242-42d8-bfc4-c8d92eee597e · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation feb0c688-75ec-4004-9c61-584a25c57995 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a2782b34-7f3c-4f5c-aaf6-ea0c4c459a6e · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 90679bdc-6143-4dc5-863f-534d97529010 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2e67078e-9dbf-48ae-b7fc-11c2a8cf29d0 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fe585f51-30b8-4557-ac6b-81cd3419d739 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39689887-aefa-408b-9359-f2cfac8868eb · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96d2b90-d07b-408b-80bd-4d0e5ec3e0af · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479152ae-c86a-4608-990d-c04855847f0b · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 440dfa31-0be3-409f-b196-f0a88836e22b · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7219b7a-c1e8-4a69-bf7a-79d6ed03e093 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Manning, Stefano Ermon, and Chelsea Finn
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a24be6-86ce-4818-bf2e-2667107f0043 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7b849058-daaf-4820-93e2-2be380a934c7 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e6a42e10-03c9-4faf-8a9e-96841871c8e9 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4778406c-3704-4c23-a49d-c691c86f6bb8 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ac7cc6b3-b93d-415c-ae03-2e4ea62851af · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1f5076a4-30fc-41fd-9eb9-160db611f4c9 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Preference Alignment Improves Language Model-Based TTS
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5472b71c-30cc-4302-8aa7-a70f2d34de2a · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f2993fc9-7472-4e96-87e7-c7adc9d5e8b2 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b4463b-47b6-46e3-89e5-055217e8dce1 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 448c2c86-a5cb-4783-bbc0-b612765e4522 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8737ed2e-2e1e-4492-bac2-4446a732a009 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdcfe42-5cb4-448a-961c-a2f136e1253a · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a01d4a64-ce51-4701-ba2b-85d3e5b7c7bc · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eceea7f9-6814-4320-ad10-83d7f16dc911 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Qwen2 Technical Report
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e45ab6-71ea-43e4-834e-d74a9a321df3 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Qwen2.5 Technical Report
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b37bf12-725d-486b-9c8b-11c6521e5cb4 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 065e654a-1738-4839-844a-b7ec9c4832b8 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fc6d839-8263-4e84-a789-1b169e45c526 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1bcea1b1-1a04-429e-a40c-5ebef4a0ce7f · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2186a10a-f63c-4d32-8698-d5870c802a3d · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c79f1747-c83a-41f8-b5a5-7e6cb985cebc · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b80548c6-48ef-4bf0-84f1-aaadf282322c · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f78a5245-eab7-4d94-bb44-e49322bd21d1 · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment online" 'onlinestring :=
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d85a7e-1406-4e45-af06-e32a602b289d · outbound
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment write newline
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b719acc5-8566-4fb8-a1e6-94edf174f7f8 · inbound
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d3e9bbc8-f5f1-4916-948d-b4831abe5629 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
Reference 237
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.