Pith. sign in

Paper Citation Record · LEDGER

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2505.20868.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20868 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:49:13.241404Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:49:08.825623Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:49:13.486015Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dfd6543-a213-4651-9f4d-04afb3ab856c · outbound

This paper cites With recent advancements in deep learning technology [2, 3, 4], the naturalness of synthesized speech has improved significantly [5, 6].

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech With recent advancements in deep learning technology [2, 3, 4], the naturalness of synthesized speech has improved significantly [5, 6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:20.604622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:08.598000Z digest=sha256:13632866a1b9c2d618bf4ed3ec7b83e772e895e7a8255fe62aac777eb0bbbdd8

Observation 37201e95-7427-4e53-a70f-ca8993b94ef7 · outbound

This paper cites Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T13:49:13.614578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:08.825623Z digest=sha256:71023cbf3c586f4ec01a977f1c8a38a23317cb461f922d0643bdc5bd26fd51de

Observation 4fb189b0-c7a9-45e0-9980-15690766f006 · outbound

This paper cites Both are about the same distance.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Both are about the same distance

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:19.947674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:08.939599Z digest=sha256:cdbcbbcf29ed52557bd9050dc987057f01b827813bc3ab8f7d34d92ff194010a

Observation aeec44c5-4f37-4aae-8896-1e08c45796df · outbound

This paper cites V oiced-aware style extraction considers the acoustic character- istics of different speech regions, enabling more detailed style extraction.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech V oiced-aware style extraction considers the acoustic character- istics of different speech regions, enabling more detailed style extraction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:19.626922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.058543Z digest=sha256:b1109c563094d640e03f697f7bd3b2d5dc5225ce5d6a5f329b30bb3f3d816649

Observation d96f3b0e-2712-4887-bdc8-68f6554cd658 · outbound

This paper cites RS-2019-II190079), Artificial Intelligence Innova- tion Hub (No.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech RS-2019-II190079), Artificial Intelligence Innova- tion Hub (No

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:19.344369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.173582Z digest=sha256:36b40970d6f141359b3d00391a169323ebe6bbfb6a7424c39d18c364f0b10630

Observation f5a2dd12-6165-4ae2-916e-8e79ce511c84 · outbound

This paper cites Emosphere-tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to- speech,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Emosphere-tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to- speech,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:18.454764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.800248Z digest=sha256:587d2d3d7d9dc8cb288ee90ed9bc99c7c8d0f9be2f2fc66fcd39216ae2e02d18

Observation b5132ec1-c45e-40d4-a806-0af870f8d5b8 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:09.278767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:09.278767Z digest=sha256:4bdf91d6c5c34b60468e154582da6e1cbfd645adb618377e18c25da51fdb307d

Observation 29571275-8d64-43b2-9ff7-6b1c4801f5b2 · outbound

This paper cites A new recurrent neural-network ar- chitecture for visual pattern recognition,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech A new recurrent neural-network ar- chitecture for visual pattern recognition,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:19.160124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.364915Z digest=sha256:b929315671f61b33681a197cf1c194477b2763b4a9a5de60d783edacf75006a8

Observation c62140b5-f2ed-4f10-a2e9-d7a13ac36582 · outbound

This paper cites Multiresolution recognition of hand- written numerals with wavelet transform and multilayer cluster neural network,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Multiresolution recognition of hand- written numerals with wavelet transform and multilayer cluster neural network,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:18.952551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.466239Z digest=sha256:6aa88b043050a554a50c241a822c15de6955fe1b1642da3048cf1b0c2e26c0b7

Observation 23c905f2-7275-49bc-b83c-26f720a333ff · outbound

This paper cites Multilayer cluster neural network for totally un- constrained handwritten numeral recognition,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Multilayer cluster neural network for totally un- constrained handwritten numeral recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:18.753635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.583141Z digest=sha256:776251fa4f7aac37ccf71f5492f6a6449b5757a6714796560ddc4ce9ae456ffa

Observation 492cd53c-4ab1-4a74-86e3-26f2be4dd38f · outbound

This paper cites Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised represen- tations for speech synthesis,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised represen- tations for speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:18.593849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.695138Z digest=sha256:bd83977c3b2b0a48e8986b06592e203843a0b6b7f7a97cf0102c6572155ffddb

Observation a48a850a-8e74-4922-823f-2a7263d62987 · outbound

This paper cites Neural dis- crete representation learning,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Neural dis- crete representation learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:17.115825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.568416Z digest=sha256:f8d9502eb3cb1e3a20f4baac2aee268a6d379ab577798e95f2dcc6dcb67f8388

Observation c48ee2b7-9b0b-4e3e-b5fd-9096f83f40b2 · outbound

This paper cites Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:18.294039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:09.957838Z digest=sha256:c149aadd0149bdfd1ed44adcb28e0d8875fe47ea68a413bce391d2ee6cbd63e8

Observation 8448b2bf-8a24-430f-b812-f41398451042 · outbound

This paper cites Style tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Style tokens: Unsu- pervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:18.156277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.073323Z digest=sha256:d1f85431a8d8bcfb65e205661dc0dfffdd70447156b6745175a878a95306ae09

Observation 2e288fe4-32ca-4ff3-8ce1-20e4398137b7 · outbound

This paper cites Meta-stylespeech : Multi-speaker adaptive text-to-speech generation,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Meta-stylespeech : Multi-speaker adaptive text-to-speech generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:17.966373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.197878Z digest=sha256:fa257b67ed35980127ae5ef7b0956a14a1a96c06e7145cca8ceafb076fb61549

Observation 87bad46a-6ff1-4fcc-8ed0-00d8ecd8128b · outbound

This paper cites Qi-tts: Questioning intonation control for emotional speech synthesis,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Qi-tts: Questioning intonation control for emotional speech synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:17.696519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.306047Z digest=sha256:2c51f8e3df1517b1ff2b174651f71e7bab2fb5ebda4021eb80fdd7cb6a4971b6

Observation 9c21da54-39e4-42aa-9ac8-76ae8847ed08 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:17.390099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.461481Z digest=sha256:a539047ae94909f0d303625d1676f5809192d3c49e83a6a8d45058645b515978

Observation 650cbabd-599a-43b2-9269-0a03b82b5f8d · outbound

This paper cites Good helper is around you: Attention- driven masked image modeling,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Good helper is around you: Attention- driven masked image modeling,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:16.219360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.337465Z digest=sha256:6bcc261f86199a13d5e86e6bde089eb98a59533ef974bee3892ba6483711c8a0

Observation f2af16f3-3e05-4ec7-84f7-77d67aa20ab5 · outbound

This paper cites Furthermore, we introduce style direction adjustment, which adjusts the ex- tracted style by modifying its angle using content and prosody vectors in the embedding space.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Furthermore, we introduce style direction adjustment, which adjusts the ex- tracted style by modifying its angle using content and prosody vectors in the embedding space

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:20.232580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:08.720482Z digest=sha256:56a8ebd0d8995588bc1a9f9500ccd305d354d75121eaf4457936bced1c6dda4d

Observation 074771a2-cf4f-4efc-8812-f55a96c6d12b · outbound

This paper cites Tsp-tts: Text-based style predictor with residual vector quantization for expressive text-to- speech,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Tsp-tts: Text-based style predictor with residual vector quantization for expressive text-to- speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:16.868842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.708545Z digest=sha256:6417cb429c59ba81214620013eeb6b5b5b6479b5e7f9f6c3354a630f7d2d47a0

Observation 0ea39445-2923-4323-b567-fd8a08025c1b · outbound

This paper cites TCSinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech TCSinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:16.669619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:10.821151Z digest=sha256:485e3e8954bf5f84d0d9d945a31dc0d2cfd6768f7a52b9f31df94a786eb3e506

Observation dab69dc1-f56f-4ad0-8981-43d15e4c9865 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:10.936166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:10.936166Z digest=sha256:4eae4d43fdc93cd6b2c2cdb90a6c79336116aad9a0cae0036fd9e8aa5daca661

Observation 7d9de12d-4f7d-41d8-b735-c0c6fe04e43b · outbound

This paper cites Signal compression based on models of human perception,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Signal compression based on models of human perception,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:16.507721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.040077Z digest=sha256:9fc6cfbf3a07aad8510481643ffedf11898910b6bed4954550bee1fc79c7dba2

Observation 0240fbeb-c308-44f8-b170-4b17c87a7776 · outbound

This paper cites Not all image re- gions matter: Masked vector quantization for autoregressive im- age generation,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Not all image re- gions matter: Masked vector quantization for autoregressive im- age generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:16.392379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.146945Z digest=sha256:140022429100006629f4509c16ca2180bc99a876c118a80a44c7565193ddd42c

Observation 97400640-9587-4f5a-a7f1-440d9a240ea3 · outbound

This paper cites Restructuring vector quantization with the rotation trick,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Restructuring vector quantization with the rotation trick,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:15.976189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.506769Z digest=sha256:821cc0871d47ff6bd1c67ed06326f94edb2dbac7c09bb55a9e518e77d83df0b8

Observation 99691b28-48f4-4916-9eb4-a0ca2170a121 · outbound

This paper cites Autoregres- sive image generation using residual quantization,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Autoregres- sive image generation using residual quantization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:15.694669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.631850Z digest=sha256:d5b88fec311fb775b18af6ffbf09d98340d6c6221e17d2bf15eaad13c01c0a71

Observation 1fe9ced6-5848-4c35-96ab-7802f53d7094 · outbound

This paper cites A convnet for the 2020s,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech A convnet for the 2020s,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:15.396367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.753822Z digest=sha256:60986f0d2377a1c66a9cc81b807ee35787ce6754a45b4326342c13dbcae86698

Observation c1aa1b80-3cd3-4629-a678-65b6d0716cff · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:15.126563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:11.920312Z digest=sha256:49292dfbba32a55d65ae45004dc1f0751e3e186dd64e12213980d33ae4d8a5b8

Observation 7a4cec8a-bac5-435d-a068-c5191f12d452 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Bigvgan: A universal neural vocoder with large-scale training,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:14.856896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:12.102402Z digest=sha256:482e08d3b112cb7f01e6e65481a048c4317c71f864fedc4a7b62ef48a959f6e2

Observation 9f0c0508-fd77-48f2-978e-a74fed2f6b3d · outbound

This paper cites Syntaspeech: Syntax- aware generative adversarial text-to-speech,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Syntaspeech: Syntax- aware generative adversarial text-to-speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:14.520036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:12.273643Z digest=sha256:c50ecede51ea8ad3ae682c82dc35256fd1e7b69bb4a77df75786792ade796e36

Observation 7c89230c-00db-4818-a6cc-7e7873e10f0b · outbound

This paper cites Emotional voice con- version: Theory, databases and esd,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Emotional voice con- version: Theory, databases and esd,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:12.414577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:12.414577Z digest=sha256:3fb1fee4ea5b7c95805751179ede9eb37f452290c9ddfffc09f0198bfff258b0

Observation bb99b960-9367-4744-a57c-e39fcc39f562 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Decoupled weight decay regulariza- tion,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:12.573452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:12.573452Z digest=sha256:772910771314236d66261b467f08047a6557d5313ff8ef8e045527af61303053

Observation f854d94f-9b33-44f6-bdf1-248c9534ecb0 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Gaussian Error Linear Units (GELUs)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:12.683286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:12.683286Z digest=sha256:4e378d148aa9c73491c5e915cbded71c9337eabaa2f91c2c0bf2da5e1d9608f1

Observation 962890a7-b18b-4de6-bc54-d4d2fe32f05f · outbound

This paper cites Attention is all you need,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Attention is all you need,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:12.837000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:12.837000Z digest=sha256:9256c79b6d5c4709cb5cf917a8ed66deeba166294045c068c14c2c510fb6b4b7

Observation 40934d75-70b6-4804-92c4-91e82b081f61 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:14.186471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:12.956194Z digest=sha256:f7fedad200507bd6253c64e833b66244eabf05e72b58a015a42868757cbcf8be

Observation 28783661-d21c-4f50-ba61-08b5993018d3 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Robust speech recognition via large-scale weak su- pervision,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:49:13.892653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:13.090584Z digest=sha256:223eb9d548d45287d04137f2cc737dd92ffecd16b514caff3cc636224ed8babd

Observation 5fa609b9-6ec7-4f85-9a93-9a2989fae3c2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:13.241404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:13.241404Z digest=sha256:0a2ad28f9b717df6e2faf744d166b81dcfb207c810aa9a2a9eb86e770bc22e01

Pith citing papers

Observation 37201e95-7427-4e53-a70f-ca8993b94ef7 · inbound

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech cites this paper.

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T13:49:13.614578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:49:08.825623Z digest=sha256:71023cbf3c586f4ec01a977f1c8a38a23317cb461f922d0643bdc5bd26fd51de