Pith. sign in

Paper Citation Record · LEDGER

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation

As of 19 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2412.20048.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20048 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:41:01.512970Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy69
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c23cb365-eee1-4a1b-bb8a-07ca5bd05705 · outbound

This paper cites The amazing benefits of being bilingual,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The amazing benefits of being bilingual,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.269476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.917403Z digest=sha256:505f1ca25b246f88987de5101ad611f940a2bdf8e4d019ceddfb7d4d08cbb6a3

Observation 5d15a981-3d6f-4e69-a3a7-70b09b9846fd · outbound

This paper cites Disentangled representation learning for multilingual speaker recogni- tion,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Disentangled representation learning for multilingual speaker recogni- tion,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.242873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.923994Z digest=sha256:0c4a66295bf8c3fd016483ac12143c4c617548d360d3a78e6ac55a1afdf81c10

Observation 8216c3ff-4d1a-4620-a9b8-fd20ab69d79a · outbound

This paper cites Crosslingual and multilingual speech recognition based on the speech manifold,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Crosslingual and multilingual speech recognition based on the speech manifold,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.220288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.932355Z digest=sha256:412421c224f10c9f51675b54c4e376c712f6582b88a2b740cbf9f32d1d6e17f5

Observation 2b16526b-174f-48db-b2f1-7ea9f1eaa3fc · outbound

This paper cites Distilling a pretrained language model to a multilingual asr model,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Distilling a pretrained language model to a multilingual asr model,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.199694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.938640Z digest=sha256:f8e32031e6c993d13f2c82de9e25d9133562f57ec373704502609a12e0d2276f

Observation 2b79a606-c8fd-4e75-98e3-43b4e84b1814 · outbound

This paper cites Joint ASR and language identification using RNN-T: An efficient approach to dynamic language switching,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Joint ASR and language identification using RNN-T: An efficient approach to dynamic language switching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.174204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.946711Z digest=sha256:a7420bb16da3cc6a4758498b020e5f2d295ead49cb3add86a585ae6384a8f9d7

Observation eed4ff27-cd21-4963-ba27-574002257665 · outbound

This paper cites Joint unsupervised and supervised learning for context-aware language iden- tification,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Joint unsupervised and supervised learning for context-aware language iden- tification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.151563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.953481Z digest=sha256:323e3b6c5db9d5be0d21a015e654c3f5dcb6e30ab696e14aceadea65a1997a05

Observation 349dcebc-a7fa-433e-97cd-28dc3958c96b · outbound

This paper cites Fastpitch: Parallel text-to-speech with pitch prediction,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fastpitch: Parallel text-to-speech with pitch prediction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.131618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.961403Z digest=sha256:2ce70ca05f31490f602dec9b35023035ad92d0c537b54d11211b0115418cd005

Observation bdb4b98c-6964-4637-9bcc-d30e1d0d7339 · outbound

This paper cites Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.107774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.969012Z digest=sha256:9fd1b4f52f3b70979d398d531082ca9b25fc2c6c0b897245c45489059689964f

Observation a604f3c6-c4ef-4ebd-896c-415750cedae8 · outbound

This paper cites EfficientTTS 2: Variational end-to-end text-to-speech synthesis and voice conversion,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation EfficientTTS 2: Variational end-to-end text-to-speech synthesis and voice conversion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.085372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.976039Z digest=sha256:0cad9d1159e7736fde4bf9210141dc714834a0a690de8c1bb5ff9bc4ab0e4ddc

Observation 05260672-3867-48f6-89a2-bf5c14eae1b5 · outbound

This paper cites Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.059788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.982674Z digest=sha256:6d9c2246e0e62219b73f4e35593033a9f5aa967c9eb714a2959d89119e4838bf

Observation a11dc593-9a18-48e6-9cf7-05aff151bfdd · outbound

This paper cites Improve cross-lingual text-to- speech synthesis on monolingual corpora with pitch contour informa- tion,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Improve cross-lingual text-to- speech synthesis on monolingual corpora with pitch contour informa- tion,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.039156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.989110Z digest=sha256:ad899107a955fe5dc22a1634db2667f02cffa3df6b0bffa4caabb7ce69fcbc51

Observation 8302a362-5913-4d46-8a29-daf86d124842 · outbound

This paper cites Language-agnostic meta-learning for low-resource text-to-speech with articulatory features,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Language-agnostic meta-learning for low-resource text-to-speech with articulatory features,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:03.015542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:00.994932Z digest=sha256:9cac7d9948e39e05f9a057caf5dfe010671e86b824b9c100081559129db73e90

Observation af47961d-8ebb-45e2-a4ea-6e4228fde64b · outbound

This paper cites Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.995913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.004849Z digest=sha256:4ef2b4d670258e17d095d11ddf64ac8eb3775eee24dd91824611856c02839bc2

Observation 8d41d9eb-99b6-4812-a974-2a26a56b0778 · outbound

This paper cites Disentan- gled speaker and language representations using mutual information minimization and domain adaptation for cross-lingual TTS,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Disentan- gled speaker and language representations using mutual information minimization and domain adaptation for cross-lingual TTS,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.973107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.013367Z digest=sha256:e713f673f309d2a45122823031654dccc599372ef680323e545c836226292034

Observation c6f464f7-6ff1-4529-8ab7-2d9d93c1e1a1 · outbound

This paper cites GenerTTS: Pronunciation disentanglement for timbre and style generalization in cross-lingual text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation GenerTTS: Pronunciation disentanglement for timbre and style generalization in cross-lingual text-to-speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.954486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.020861Z digest=sha256:0031f686a978fda34a6003e047d7903ce11b73011f872dae8880d424919f9896

Observation cd8acda7-6487-45a4-82f8-9eba3e9e4432 · outbound

This paper cites DSE-TTS: Dual speaker embedding for cross-lingual text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation DSE-TTS: Dual speaker embedding for cross-lingual text-to-speech,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.929948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.026896Z digest=sha256:136aa96748fbdcb8afe61eda39585c247553e4da8e299663aa01b0ba0e9bdafd

Observation 93e48c2c-23e7-41ff-8d96-4703c38e8ec7 · outbound

This paper cites ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:41:01.687794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.037296Z digest=sha256:dbf69119e68c9b4de93e7c74c1d304fb359360e6ad029c246e6523adc295276d

Observation 24b0fff1-493f-4e57-bbd5-7cbf40f7e594 · outbound

This paper cites Unit selection in a concatenative speech synthesis system using a large speech database,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Unit selection in a concatenative speech synthesis system using a large speech database,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.907351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.048078Z digest=sha256:8098d132d6d2f1704f81a8a0777c8c4bfb2a14342502ffb2b208bde81b6906f7

Observation 91598463-fd05-41fc-9cb8-7aa59d3722e9 · outbound

This paper cites Statistical parametric speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Statistical parametric speech synthesis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.884267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.054553Z digest=sha256:bebf2b34801384615db2f715e0fa1dcd3a033ba9044075cf40e033d4af01078e

Observation 6d1d7ac2-cff8-4500-a8f4-be36d07032a7 · outbound

This paper cites Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.864249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.061028Z digest=sha256:a4f5e0e038c28465d3e098ca6c5f7680e8b11f9ae7d743237796d9dd5ee31910

Observation 30070c5f-f1c9-4403-9a58-cf8fa98d6ec9 · outbound

This paper cites Harmonic-net: Fundamental frequency and speech rate controllable fast neural vocoder,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Harmonic-net: Fundamental frequency and speech rate controllable fast neural vocoder,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.845442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.067703Z digest=sha256:565cc0005ca943923256392eb2956e12ab169942e9b30133931572cfe7dfa6d3

Observation 17f12e01-95b9-44b2-8e78-61efbc5e51c5 · outbound

This paper cites Fregrad: Lightweight and fast frequency-aware diffusion vocoder,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fregrad: Lightweight and fast frequency-aware diffusion vocoder,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.824269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.074760Z digest=sha256:1afcfb73feae837e15fb5624d7f1d4cc5f62cd32a27ee20a5d67f718988d979f

Observation 5b22e679-ee23-42cb-bb09-a32f3753b78f · outbound

This paper cites TriniTTS: Pitch-controllable end-to-end TTS without external aligner.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation TriniTTS: Pitch-controllable end-to-end TTS without external aligner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.800273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.082362Z digest=sha256:756597531e15004745ac69972334b3885104e8f888252154b7ee0c365a001e9a

Observation 86d5654b-fd8c-4a89-b268-c97125f338ff · outbound

This paper cites Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Hierspeech: Bridging the gap between text and speech by hierarchical variational inference using self-supervised representations for speech synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.782213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.087731Z digest=sha256:e8c61c9d9e8dc5c2146e7ba27cacca7c2475ba9a4d5fa5dca3395b8af622e7d5

Observation 614d5c98-5ce2-4392-8be2-08ff8d25809c · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation WaveNet: A Generative Model for Raw Audio

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.095472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.095472Z digest=sha256:4e18456794ccc28e1c2594a55572062286f94def2dfe55c9826df14d1d6411b0

Observation 4bceb40e-42d5-4061-a722-705378184bec · outbound

This paper cites Deep voice: Real-time neural text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Deep voice: Real-time neural text-to-speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.763678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.105039Z digest=sha256:1d91baa634fe29eb20613e5fcfc1d07b90693ac87173b1475de0a7627f47c4f0

Observation c9318edb-cb90-4bc6-b523-4e1e1278648e · outbound

This paper cites Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.740586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.116179Z digest=sha256:5c2ba57e35672775bc4b6f3bf804efe5f52e30e82129766a089f759089ebdd78

Observation 57adf011-d731-4c2d-8505-c9b089076f6b · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.722246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.123118Z digest=sha256:e8f2027142b3a9fb2206634dec87f8ac670d00195aad95e741d10020b5c941c6

Observation 199ac34f-19af-4db2-a874-85977630b03b · outbound

This paper cites Matcha- TTS: A fast TTS architecture with conditional flow matching,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Matcha- TTS: A fast TTS architecture with conditional flow matching,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.703668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.128890Z digest=sha256:1fa6013bc27e741edb9ecb76bd081b23569d6489190e220cf2e986a4d2db24ae

Observation bb6e89df-9fc1-48d8-ac11-f7b31e9a642c · outbound

This paper cites Multispeech: Multi-speaker text to speech with transformer,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Multispeech: Multi-speaker text to speech with transformer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.685786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.135253Z digest=sha256:b76da74674eb454d1c9c6139a448eb9cd725de6dc1270cddb56f368879ca256f

Observation a10192b2-9bf5-40e9-b767-bef9dbf8dbc8 · outbound

This paper cites Lightspeech: Lightweight and fast text to speech with neural architecture search,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Lightspeech: Lightweight and fast text to speech with neural architecture search,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.666778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.141137Z digest=sha256:78c1ba7118e47e12de5579a0f313f7604fd7b6f47e3fac5b3242224b878e9ed8

Observation 92c323f5-604a-4040-ab6b-ff61414a0bcd · outbound

This paper cites Phonological features for 0-shot multilingual speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Phonological features for 0-shot multilingual speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.646583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.147694Z digest=sha256:a326f49ad3580811dc60dec35c4066c2f72e5f06d84d29ecafb7e1e81d270822

Observation 85d6b405-fd35-4a37-adff-08282372a955 · outbound

This paper cites Text-inductive graphone-based language adaptation for low-resource speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Text-inductive graphone-based language adaptation for low-resource speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.626269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.153949Z digest=sha256:b3e3e7ab05d6bc1eb18241a58494f444795051bc19a4659e9bd0527002344417

Observation 7e7c046c-2cf5-42ac-b3f2-7cfea3e0e140 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.601319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.163504Z digest=sha256:af4f145ea33f2e28a3e0397bc13b4fc5ce1873bbfb46ee8e215b8c20e8ee50f6

Observation 7af7583e-39dd-4fd3-acf9-644ea4cd4064 · outbound

This paper cites Domain-adversarial training of neural networks,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Domain-adversarial training of neural networks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.579706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.171886Z digest=sha256:a51db872a942f52aabdf2ebc9189b11e70b2e882bdfb68b4c77e24c48c328329

Observation 785c3b1b-eafd-4333-8647-7bc47c845abd · outbound

This paper cites Learning disentangled representations via mutual information estimation,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Learning disentangled representations via mutual information estimation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.545882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.180485Z digest=sha256:fe3be723ac343f7232102e409a91d8e2a8ae29b11389a9976fd7fd30c65c1ada

Observation 36afed2e-9437-4b04-a4dc-48e4d2cfb84e · outbound

This paper cites SANE-TTS: Stable and natural end-to-end multilingual text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation SANE-TTS: Stable and natural end-to-end multilingual text-to-speech,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.506996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.190551Z digest=sha256:83db277ecbf4870ec76c97db7610f8c99a4ec5d45b3d1b0b46296300b34691c5

Observation 97c9a134-6f77-4bdc-b621-9e75747d8f33 · outbound

This paper cites Crossspeech: Speaker-independent acoustic representation for cross- lingual speech synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Crossspeech: Speaker-independent acoustic representation for cross- lingual speech synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.489429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.198083Z digest=sha256:73b40026717e1dddc4ee48a270bac1be9624114c6df86d5f818543d6fec2f829

Observation a47dc9af-30fa-4784-a4a9-756411a563de · outbound

This paper cites Invariant risk minimization,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Invariant risk minimization,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.471418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.206074Z digest=sha256:e42ffd648ef840892a80de3051f6a06f8ead0df143eca861a820aca518e3661e

Observation dc79d5a0-1db8-437b-8f03-c3ec0ddb7f7d · outbound

This paper cites Domain generalization with mixstyle,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Domain generalization with mixstyle,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.451892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.213292Z digest=sha256:ebe4f1a94402bafba9d6fa96d13ba4e35cc138a98b799357af7e45c2d2784ea9

Observation b1a2639b-98e0-4e28-95e7-d75dd4702376 · outbound

This paper cites One TTS alignment to rule them all,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation One TTS alignment to rule them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.427354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.232108Z digest=sha256:ccc7e03ca6c54ea5da6c545177474dcb5f52bb99612f8783f0e1b33da6d07f9a

Observation d0c5ad9d-09b4-4dde-967d-f6f1541e0898 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Conformer: Convolution-augmented transformer for speech recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.407581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.240834Z digest=sha256:52b45885a0b2de643a17bfc6a5d05bf492560a4d5c8a314799d9180cab32e7ff

Observation 3ba2697f-2546-41d8-ae46-6659e9c2772f · outbound

This paper cites Feature-critic networks for heterogeneous domain generalization,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Feature-critic networks for heterogeneous domain generalization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.378049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.258299Z digest=sha256:5b60bea6c23ae0fb7288a9a3a40ee2681744be06c4bd3078f0dd72169d14d801

Observation a66522b6-11f1-4ab5-b738-483249044bb8 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.349918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.267937Z digest=sha256:2ef17186bb13a74c7b9007ac5e4b2c61a32cdbbcc989ef22c04d1eae68d72302

Observation c4ece250-45d7-4872-beb6-ebd9ae1e1680 · outbound

This paper cites PV AE-TTS: Adaptive text-to-speech via progressive style adaptation,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation PV AE-TTS: Adaptive text-to-speech via progressive style adaptation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.327619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.275765Z digest=sha256:7a7008bc112859fbd1086a4d59ef58a01e0eb700510724c3bfc01851d0d35774

Observation fcfee05f-68af-4a65-998e-6b01c8c3a593 · outbound

This paper cites Style normalization and restitution for domain generalization and adaptation,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Style normalization and restitution for domain generalization and adaptation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.302559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.282956Z digest=sha256:4766b1080b04102faf74f50caf523baecfab86f1582b70f812f167aa7bc55a78

Observation 602d7743-9ae7-4ad6-b30e-d0c9a3de8ea6 · outbound

This paper cites Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.277338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.290550Z digest=sha256:92273a5e9cb591d0dd75fb2eb874804a1d9b54c2bbeb975b1f3ad7af363c7831

Observation 01b8547f-fdc7-467c-a438-a8e4011503cd · outbound

This paper cites Diffprosody: Diffusion-based latent prosody generation for expressive speech synthesis with prosody conditional adversarial training,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Diffprosody: Diffusion-based latent prosody generation for expressive speech synthesis with prosody conditional adversarial training,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.252247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.296819Z digest=sha256:d1c2db00bbb974209918a1bfb8f3d8c47f5aefc363dfac3ed3ef4726d9f13de3

Observation 26cfbfa1-825f-4daa-a197-d4e51f33f2c6 · outbound

This paper cites pYIN: A fundamental frequency estimator using probabilistic threshold distributions,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation pYIN: A fundamental frequency estimator using probabilistic threshold distributions,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.228142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.302355Z digest=sha256:be4cd5a583de1353bbb4e6162d0883f92b8741abe5cfed088349e960b7b4f660

Observation f75cf353-7153-4d34-88dc-12db8c1bde0a · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representations,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Neural analysis and synthesis: Reconstructing speech from self-supervised representations,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.209008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.310006Z digest=sha256:26a48baada07bf8580b33f90c6dad2286e2747a7589f61b77ce24406617d656a

Observation 760857da-4f2d-4ad3-b8d5-19ec24273660 · outbound

This paper cites Exploring wav2vec 2.0 on speaker verification and language identification,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Exploring wav2vec 2.0 on speaker verification and language identification,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.190483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.320135Z digest=sha256:93196d328a16140c8b7918975e72d02d0be9601f111fce26ce8397cbd72e551d

Observation 71799840-ac50-45c1-893c-407a7d1fd3a2 · outbound

This paper cites Let there be sound: Reconstructing high quality speech from silent videos,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Let there be sound: Reconstructing high quality speech from silent videos,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.169374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.326985Z digest=sha256:9014c1e6d91cdae0e5af4646b79c9749dae7fd4ef024026c467d9164b8ade55b

Observation 7cf0cd41-4470-4132-a999-417d4e9ebbae · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Scaling speech technology to 1,000+ languages,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.141990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.334513Z digest=sha256:92611b882cf800c0f4c77a260f28231aeba721fcae5c7f704baee5b389ac9dd5

Observation d3ef05fd-f846-4cc0-a9b2-c146ec400b44 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.123668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.341458Z digest=sha256:3fd7fef75553a0ce5177aa6f9ad942360da8529ef5bf6ef0c9ebf8833ad7a1d1

Observation 0f4223bf-e28b-4eee-8e8c-47859fd98892 · outbound

This paper cites Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.101079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.347868Z digest=sha256:1d7ec6bcc6e3bf59eaa79284a30405f512b8cc5d0993c9f0070e163b3de24a0f

Observation acaa876b-4f89-4927-9574-4d563ef18ac5 · outbound

This paper cites The LJ speech dataset,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The LJ speech dataset,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.082325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.356498Z digest=sha256:3da0e7e14c84738b63aa330cfa13d5be198b912c5a4fdfe4022abb74b593554c

Observation a57298f3-6db4-4c26-bd39-344a6ea1c185 · outbound

This paper cites CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.363187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.363187Z digest=sha256:db7bf4c0fc36be2d6d02001a3b7d3bff171d390e8065e6d139b531ebdebefd27

Observation f5c4ab24-b74c-4038-a6a3-f30616549ad2 · outbound

This paper cites The BIAOBEI dataset,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The BIAOBEI dataset,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.060733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.373919Z digest=sha256:12b92725f765feb78241ce1b6f09d164ac1a79cc447cda31f85dd49ac7bc52d2

Observation 0e24fc1d-ff54-4eac-9a95-82c80d341ab7 · outbound

This paper cites AISHELL-3: A multi- speaker Mandarin TTS corpus,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation AISHELL-3: A multi- speaker Mandarin TTS corpus,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.040412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.384569Z digest=sha256:6c282b18172650e8aac67f71973379cd8849e9dcf8885f7d50bce6814c1313e5

Observation db8ce6fb-d726-4e9a-8b6f-d6c4797a14b3 · outbound

This paper cites CSS10: A collection of single speaker speech datasets for 10 languages,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation CSS10: A collection of single speaker speech datasets for 10 languages,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.022899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.393788Z digest=sha256:5ebed746d0d31299c50cd41a1fd5b793b98227a15b6bdbb8498d2008fadd38a2

Observation b73612f4-d09b-4e05-8e46-eb109fdb70ce · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.404222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.404222Z digest=sha256:c2841696d5e132329a76544446dd77fa976dc94ba67bb903ce12ff90542bf71a

Observation 4a2e9241-8696-4630-ae8c-87cb57ecb2ae · outbound

This paper cites Multi-speaker TTS data,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Multi-speaker TTS data,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:02.001771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.413718Z digest=sha256:b6c5cb13383b3ce72d760b1875d3f88eaeee78dc928ac6247c9d1c664789bf7d

Observation 3972753a-281c-48cd-8d64-58c05baed105 · outbound

This paper cites Phonemizer: Text to phones transcription for multiple languages in python,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Phonemizer: Text to phones transcription for multiple languages in python,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.981863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.419782Z digest=sha256:62bb906beb9e3a3b669bc74691ddcd7bbaaafc89984e596b4f86c2c98a6472a0

Observation 83f93c23-0ea3-49da-bc02-620346aed726 · outbound

This paper cites NANSY++: Unified voice synthesis with neural analysis and synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation NANSY++: Unified voice synthesis with neural analysis and synthesis,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.958515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.426365Z digest=sha256:08a4c1779fee3e73d238b2138733cc7afbda7d174a64cb3324317d18f7fe4a6a

Observation eb81c13e-73ad-42df-bdce-8e2274db605a · outbound

This paper cites Language modeling with gated convolutional networks,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Language modeling with gated convolutional networks,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.934226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.439805Z digest=sha256:b52121a5392105fbd349e834dfcea5638e1cf3118a1d26a357e7b9ee9c41df3a

Observation dd1b4f6f-e28d-4026-b08d-8bcceceaadab · outbound

This paper cites Fre-GAN: Adversarial frequency-consistent audio synthesis,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Fre-GAN: Adversarial frequency-consistent audio synthesis,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.908739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.453287Z digest=sha256:f24c9b3ccd37ac978e57fc4b8f442920cbf282425a0ff6c44b446c092ea1c8c3

Observation 966db735-0d84-4e0b-bec3-82e608b80322 · outbound

This paper cites UTMOS: Utokyo-sarulab system for voicemos challenge 2022,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.888958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.460239Z digest=sha256:b9da90fa1977bf68da5b691eaa0a8419977faad6411c7ee312033d1d45fc5d91

Observation 3704482b-c728-455c-ad3e-694cd1fe8531 · outbound

This paper cites The blizzard challenge 2005: Evaluating corpus-based speech synthesis on common databases,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation The blizzard challenge 2005: Evaluating corpus-based speech synthesis on common databases,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.862661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.467343Z digest=sha256:ef46b8223f36513d9e23159d677ef38d35b553e7d7aebcc70c67e2bda4ff9133

Observation e17948f1-e53e-4637-a248-cd30562989d9 · outbound

This paper cites USAT: A universal speaker-adaptive text-to-speech approach,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation USAT: A universal speaker-adaptive text-to-speech approach,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.841392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.473731Z digest=sha256:cc4fa33399b2825cf257c7f310ca986f11b048653922001c0784afebd71da734

Observation d5973154-1774-4f7a-9d78-36add20b08ae · outbound

This paper cites Dual-branch modeling based on state-space model for speech enhancement,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Dual-branch modeling based on state-space model for speech enhancement,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.803661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.484734Z digest=sha256:dd1b7a5505191dc067f3d58a9214772c48f3b9328f4352ad7db49d0ed188e4db

Observation 5b5b390d-242c-4881-949c-f96c4c20a761 · outbound

This paper cites V oicegrad: Non-parallel any-to-many voice conversion with annealed langevin dy- namics,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation V oicegrad: Non-parallel any-to-many voice conversion with annealed langevin dy- namics,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.779771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.490103Z digest=sha256:2f1c87b7562d64ad3b3f3281aef02428b45d2bdf8b97abb99de9f538de61e1a4

Observation 7f978c48-dc69-4df5-913d-dc6478aadf7d · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Robust speech recognition via large-scale weak super- vision,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.737629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.495439Z digest=sha256:9da75b584df1108df20320fb48743990053f05bcbd3fbadead4494bd8f1debe5

Observation 192c509f-3006-45f7-88c2-f350c5d28d0d · outbound

This paper cites Visualizing data using t-SNE,.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Visualizing data using t-SNE,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:41:01.710005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T23:41:01.500859Z digest=sha256:ee73ec602e8b1f09c89bcea967811734f39490acb90f1d65e431b5dd11a693b6

Observation 96e5ca60-1ce3-47fe-b1c8-a37cf7fd8237 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.505486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.505486Z digest=sha256:8fb763ab4d9ad942c136799159bc4f5ae9b7b2f7153809d7c03f3265e9d674b7

Observation 99f0a98a-20ff-4fe6-ac2b-0712939f2c70 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:41:01.512970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:41:01.512970Z digest=sha256:f8b3674f480e194dd452f8ea14404a498ed03149f968bdcc82f82c94920b28ca

Pith citing papers

No inbound Pith citation observations are available.