Pith. sign in

Paper Citation Record · LEDGER

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2607.06461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06461 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T05:09:34.802612Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T05:09:34.802612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T05:14:35.199833Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact10
  • verified fuzzy30
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 438cc1ce-dfd5-4814-80e4-8e8e06519130 · outbound

This paper cites WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.202276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:0d8f5840ad89ade1681e161536eb3efcf880466a9efc3a452763257f5778feb0

Observation 7dc65520-1057-424d-a7cd-7b2c80570886 · outbound

This paper cites Overview of WordV oice-5A Several large-scale open-source speech corpora have sig- nificantly advanced the development of zero-shot TTS [14, 15, 16].

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Overview of WordV oice-5A Several large-scale open-source speech corpora have sig- nificantly advanced the development of zero-shot TTS [14, 15, 16]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.349750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:fd8e56ca7b0ef6db9f49f655f25880f0af11a8b2d60ed185861bcd6a2972c5d2

Observation 06bbb9cb-f615-4687-9d80-4c79812d3ac6 · outbound

This paper cites acoustic planning.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS acoustic planning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.388376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:6d7aacc52ecd732e98df3439e36a527730f481d701c1e6ae6bb4bf342f679af5

Observation 476a24fd-c548-412d-8574-f17a63d7eee3 · outbound

This paper cites Experimental Setup We evaluate our method on the WordV oice-5A-test set, com- prising about 2,000 Chinese and 1,500 English unseen ut- terances.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Experimental Setup We evaluate our method on the WordV oice-5A-test set, com- prising about 2,000 Chinese and 1,500 English unseen ut- terances

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:14:35.206410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:366535809f328b715782c33661db1f87dba62aaf79eb666cb97e948e3b7affd2

Observation 03077e0d-be3b-4d6f-b204-ba09b7d90dfa · outbound

This paper cites an unresolved cited work.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-08T05:14:35.390922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:9ec5f08e2a242665fcc1286f5be568d98f54dc2caf98215708c69bb1d3d54f74

Observation 769518fd-403c-4d48-bf25-e5da6fb2c017 · outbound

This paper cites Emosphere++: Emotion-controllable zero- shot text-to-speech via emotion-adaptive spherical vector.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Emosphere++: Emotion-controllable zero- shot text-to-speech via emotion-adaptive spherical vector

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.399994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:595168b42e853db1cc1abda590cdad8607409b8c5cbcbcb0c02c3a41cea10670

Observation 4c152270-885d-4e43-aa32-a5b657ed1a1f · outbound

This paper cites Facespeak: Expressive and high-quality speech synthesis from human portraits of different styles,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Facespeak: Expressive and high-quality speech synthesis from human portraits of different styles,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.374463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:85ba5eadc6ad688bdadff9ab28315201e6404c9ac17b58c8b95917b1183bed78

Observation 66bab643-0bc0-4923-a566-3ccfb6e1bb0f · outbound

This paper cites Scalable controllable accented tts,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Scalable controllable accented tts,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.402098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:060f5251c8c0e9641d1bd8aa75c0652e5cd48d2d26214fcc753cc14507527a72

Observation 004cbd2b-4601-4f9f-ba22-f6f69da22755 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.407108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:adebe389f4628a3f25b6fc17049f23e19d5e01fcfb12c6f1b7c9d36dcb330447

Observation 29ed34aa-4b29-4974-aad7-8f76ccfb7fd0 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.210903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:858454c79f660edb2b4f7961b921008ff8000515681a7939346a7f5a4733e287

Observation cd47ba6f-5544-49b5-a464-0f16e366703c · outbound

This paper cites Hd-ppt: Hierarchical decoding of content- and prompt-preference tokens for instruction-based tts,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Hd-ppt: Hierarchical decoding of content- and prompt-preference tokens for instruction-based tts,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.319449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:03d3da6d16705b70934fd0cca3fb3a90f9d02a6c3a75bd1bae75517fcee3487f

Observation 0e863b7f-d250-427d-98fe-16abf7a70370 · outbound

This paper cites Deep dubbing: End-to-end auto-audiobook system with text-to-timbre and context-aware instruct-tts,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Deep dubbing: End-to-end auto-audiobook system with text-to-timbre and context-aware instruct-tts,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.393361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:31961a8f0e3f03a2b364a5bdef6348a04185704685c268cd7a7584160e6a718a

Observation ec4d5115-d1a6-46da-b9a2-142a33034c98 · outbound

This paper cites MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.191069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:f494b14849e81925919428dfa23ed1c6a8d3b62daab3f350a840ecf66948309e

Observation f8cae5a2-c6b9-4ace-a002-959bb432b877 · outbound

This paper cites Synctalk: The devil is in the synchronization for talking head synthesis,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Synctalk: The devil is in the synchronization for talking head synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.323938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:b8965ce04347fedca68cefc496a70cf19540becb7c8d837ca9c0fb1a97a00c2d

Observation dff54b9b-a21b-42df-813f-a4d38ceddc4e · outbound

This paper cites Di- flowdubber: Discrete flow matching for automated video dub- bing via cross-modal alignment and synchronization,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Di- flowdubber: Discrete flow matching for automated video dub- bing via cross-modal alignment and synchronization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.367832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:9e278f1d23c38aeb7d59f10934e572288525581017b94a9af872142a069f00e9

Observation 9accc0e3-a898-48a0-a184-2f16a4c51616 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to- end text to speech,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Fastspeech 2: Fast and high-quality end-to- end text to speech,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.360398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:58d4f21b8daab45a41cb1f624d6d0c5ecab5aad12a05b0aa938c594b69670d8b

Observation 9a66b4dc-b8d7-4e25-83e5-c6027834b985 · outbound

This paper cites Conditional vari- ational autoencoder with adversarial learning for end-to-end text-to-speech,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Conditional vari- ational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.409227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:37c78b3794678bfc81322784a2b387715fce0d9e6a3c56565e9bf154cdbc117b

Observation 8fef926d-300f-4bee-a1b1-8dbea2b008ad · outbound

This paper cites Word-level emotional expres- sion control in zero-shot text-to-speech synthesis,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Word-level emotional expres- sion control in zero-shot text-to-speech synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.381462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:b2b10796b28f89085ec062c9f0463d53edbb11c2f657164ee573a6c2799df3cc

Observation 752bb62c-b468-419e-9aa4-61162a55f10a · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.362887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:35fd80a6f61c898d91f40a3ac4b0456a5e7acc6ad8f8bf22e3ec62cfd71d561f

Observation fca09a3e-aa5b-4d74-9153-c1f27c158fa4 · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.357410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:8456e70c5599be94bf00b5a8acc3ba0065c4687f1f2ea0767c6f6e3c3487def2

Observation 18dedd07-d8ac-4b5a-b7b2-39a955fd8f4a · outbound

This paper cites Speechcraft: A fine-grained expressive speech dataset with natural language description,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Speechcraft: A fine-grained expressive speech dataset with natural language description,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.397281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:d248dc680d5419aa74bd94685812e88902bfcbd106bc8ce317da0e79127eac68

Observation 9fef3e93-8eb1-45b2-a641-054cdee4c79c · outbound

This paper cites Lemas: Large a 150k-hour large-scale extensible multilingual audio suite with generative speech models.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Lemas: Large a 150k-hour large-scale extensible multilingual audio suite with generative speech models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:14:35.219334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:dd4bc1dd228c1953594837491ef5e9649df773299bfd5fca828c804ecaaa6bac

Observation 20481646-4367-4088-9f35-361e70645fbf · outbound

This paper cites Praat script to detect syl- lable nuclei and measure speech rate automatically,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Praat script to detect syl- lable nuclei and measure speech rate automatically,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.404855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:a9d3ca7e0be8ba81f97ecacfee6f5b5cd392aa933adb09124f614332979db472

Observation 082c4c6b-aa74-41fa-bfed-3b5c34fd1c6c · outbound

This paper cites Tobi: A standard for labeling english prosody,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Tobi: A standard for labeling english prosody,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.377885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:5c91eb590ceb526eb7a053d4f76edee7348427149f600f49262cc0e4c3fb7b7f

Observation a9ddb654-9a17-4432-987c-914523c64529 · outbound

This paper cites Chinese prosody and prosodic labeling of sponta- neous speech,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Chinese prosody and prosodic labeling of sponta- neous speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.384540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:cb9427a76f688b8248a9dd0b2886d494b368255bc0f043ca56039e5d83baab64

Observation c30748ee-b961-4b03-aecf-ea23cfbb2afd · outbound

This paper cites Smoothing and dif- ferentiation of data by simplified least squares procedures,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Smoothing and dif- ferentiation of data by simplified least squares procedures,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.333590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:cda4e51226a2d481c393fe303eaccd17b3d1289d2d080645ce11bb11343de630

Observation 0f59c3ce-b124-4939-8acf-a20b47bd4b4c · outbound

This paper cites Monotone piece- wise cubic interpolation,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Monotone piece- wise cubic interpolation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.321431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:8a3bc7678c910820fa104dd804717302c520ccfd60315e60c10c6dc43743539d

Observation d74c1440-bcef-4cf5-b90a-81004b73dcb5 · outbound

This paper cites Sound level protrusions as physical correlates of sonority,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Sound level protrusions as physical correlates of sonority,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.314036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:c6b41cf9b0c6a95a1329f3e0d530d411abd62d7ba6065416926112cac966ce91

Observation a57a7a16-7d25-4940-b772-c5dd242203e8 · outbound

This paper cites Considerations in the normalisation of the funda- mental frequency of linguistic tone,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Considerations in the normalisation of the funda- mental frequency of linguistic tone,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.353346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:f0ef8addea0ad88832c9dd0d26ff60bc402402346cb071d41d0b89a9d7e4236b

Observation 95f7fee4-bf8d-4dbe-bc47-116a2ed00f32 · outbound

This paper cites Acoustic features of mandarin tone production in noise: A comparison between chinese native speakers and korean l2 learners,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Acoustic features of mandarin tone production in noise: A comparison between chinese native speakers and korean l2 learners,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.304670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:ecedbc374494e1b23faa82eea3a9402e073784e9bb406ef58e06d8c8b794a682

Observation 19dd8176-4bdd-48c7-900b-b99f09205ebb · outbound

This paper cites Modeling tone and intonation in mandarin and english as a process of target approximation,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Modeling tone and intonation in mandarin and english as a process of target approximation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.307634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:d00b257c1542f23f38dbfe118cae4fccb91a70d878d80df33cac64a100f11be8

Observation 5e1ce995-76c2-4fde-bd10-0fa8cd2fea09 · outbound

This paper cites Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.339602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:10f6427d448482cbf5efdde2e3b7fd6d2c673a117623f1760ab10485eae1ef8c

Observation 63524f2f-3004-4f3a-a6ef-20f58fbb9592 · outbound

This paper cites Qwen3-TTS Technical Report.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Qwen3-TTS Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.232183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:590fc2d489b599c01bc651009d5f633200b554bba952fbca50fae777552f60cf

Observation c1c910a5-cc10-4099-8073-99b634d5898f · outbound

This paper cites Moss-tts technical report,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Moss-tts technical report,

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:14:35.182112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:b4107ce2ed595ea762574e5d17a6205ceb80b0be62a41c085e0f79dfdc0edf52

Observation e9165386-2451-42cd-beea-377813d722bf · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.215391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:14f8ae2cdfc98dacfb6142f5296798285df208c2169061abbf3947946179594f

Observation 35794823-e47a-4139-b707-d5cc2a093b73 · outbound

This paper cites HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.225166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:d38ea13f4b5d44b6f91cee1e52a015fe63dc0df04ede87429ea131b08d19c01a

Observation 81016160-7cdc-48f7-bdc2-45f5dd2dcf22 · outbound

This paper cites Qwen2.5 Technical Report.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Qwen2.5 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.196999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:43a9cbee443a53eda7c4a43acee06bf372ddbe41724e33297241adb3fd491a2d

Observation 49189f95-6c17-4aec-81c5-26eff76132e3 · outbound

This paper cites Text- free prosody-aware generative spoken language modeling,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Text- free prosody-aware generative spoken language modeling,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.344091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:84452cea72166f064a358a7814c058b748c436346192143c2d528575b16f76cc

Observation abc989ab-da08-4463-8fc5-671b501b63c4 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geome- try and semantics,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Multi-task learning using uncertainty to weigh losses for scene geome- try and semantics,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.312064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:096492b2cec8f48f29b770047a16b3e18df96caba1d15b1ae6f388be79935cf0

Observation b58c0be2-c677-438b-875e-d2f78ce0c924 · outbound

This paper cites Scaling speech tech- nology to 1,000+ languages,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Scaling speech tech- nology to 1,000+ languages,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.316616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:207cbcf3e31e3ee1d20974682469dde43da54855f90801d6e2475540a3e953dd

Observation 76ec2a85-0826-437e-94af-5c8f54ef1d31 · outbound

This paper cites An overview of voice conversion and its challenges: From statistical modeling to deep learning,.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS An overview of voice conversion and its challenges: From statistical modeling to deep learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:14:35.370869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:e9b19644310f85cf924130abf94ebd9bf079bb91c9ff56d000d2f8b233ea493e

Pith citing papers

Observation 438cc1ce-dfd5-4814-80e4-8e8e06519130 · inbound

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS cites this paper.

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:14:35.202276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T05:09:34.802612Z digest=sha256:0d8f5840ad89ade1681e161536eb3efcf880466a9efc3a452763257f5778feb0