Pith. sign in

Paper Citation Record · LEDGER

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.11650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11650 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:21.772183Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 522995db-b11a-42d5-a5e9-e411b986dd1d · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.591905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.591905Z digest=sha256:ab83c45bbbe058653dc5a473afeabd8c1b2bc5b9e0158aa0e87e980d57331a69

Observation 91f0d584-4046-4020-99df-8a05eae7e989 · outbound

This paper cites Better speech synthesis through scaling.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Better speech synthesis through scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.597148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.597148Z digest=sha256:bd94a99524836df060c232bdaaee8becf92622f24188ddb4d123355ebe9644f5

Observation 168de320-8cbc-4adb-9c74-5a0875765363 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.601698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.601698Z digest=sha256:c03d19729309f8235d16feb700f4eef67175c96c2e54ff39cc34322262c5c045

Observation 08e61146-fd25-4f06-8e69-9f784b6de208 · outbound

This paper cites YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.496936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.606031Z digest=sha256:d7b84524bfc3db34102ff7ef21a529a064f4af50bce7ddd2fe4784af79433ef2

Observation 307bb490-e946-4a7c-92c3-5c23997d3d48 · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder WavLM: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.481546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.609726Z digest=sha256:962e19f9c8e48093d4db52d440602b6b153fd35fbec2b7dbdbcb375755d817b9

Observation 5d63568e-4495-4fc0-8521-de52398f9c33 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.613843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.613843Z digest=sha256:d4007289ece912a7f0cd757f1ad34e2dba110306434abe2d7bcace7159398090

Observation 6eaf9818-2954-4232-ad28-8a18644b280f · outbound

This paper cites w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.467348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.618053Z digest=sha256:27c511397b8379f8a7c5ab9a0085aaa0891221431bc64413cf431bca412e4dd2

Observation d1e8297f-5517-4315-803e-1177a6098a76 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.621880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.621880Z digest=sha256:3fae8ad9d1e56509d5b385a08e724c8c511d7796dc0aa0b74312012158c4f5a3

Observation bf21efb6-8216-4119-a169-78a80e1930b9 · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.452706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.626443Z digest=sha256:488168f83c9eb2ead1c5b1de3f915c7be32f5f39a9c0f4d54fbce8ecdd9484ad

Observation 2b3a91b3-eca5-4e8c-8f8c-d74484b897e8 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.630552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.630552Z digest=sha256:a32d27d1e3a0d3df7da39dc2e3f07b89a1f39d0a69ef2e328dd42f8aadf27661

Observation 77a6d468-c0c5-4318-89ad-8d6560c533b1 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.634937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.634937Z digest=sha256:c02e0ca2bc5c2e7177fb029f01fb936f56937b298c4800f87cf2e5d08f60452b

Observation 4c9e8fee-bcc0-4b66-8f6a-1823e3024333 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.639687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.639687Z digest=sha256:defb0d34a9f7f3949aed0e1b7db7017f5364d6b3b3824e3ee919ccedec7f9373

Observation 12761c13-59d4-4838-8a85-fc06de9715be · outbound

This paper cites ElevenLabs multilingual text-to-speech, 2024.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder ElevenLabs multilingual text-to-speech, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.436624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.643733Z digest=sha256:a51cb791aa67892b7d8eb7915004c5fed6f9a452986d362dd1386492658777e1

Observation 5abfd9d1-ad6e-4e63-9de9-ed7c959a217f · outbound

This paper cites an unresolved cited work.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:36:22.423415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.647442Z digest=sha256:2bdc52b987c747428e3d0f1dfc40b1fc13093c5861fc5c747afe93c0aba2e94e

Observation 0ac40676-6dfb-435f-aab8-561f0b362cc6 · outbound

This paper cites Fish audio speech-2, 2024.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Fish audio speech-2, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.410694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.651225Z digest=sha256:61a5a616296aa47bc044dfe54b1635c0476264cc817241397d86ea149c977e63

Observation d1b00463-796a-4ac9-9366-52cceed22412 · outbound

This paper cites CV3-Eval: The cross-lingual evaluation benchmark of CosyV oice 3.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CV3-Eval: The cross-lingual evaluation benchmark of CosyV oice 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.396763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.654719Z digest=sha256:7ae1f888240dc604e33c1a3cc1a2de3abbca500d270ce14c806cddccf7d925e1

Observation a0972a56-f80c-46fc-a82b-3175a160d78c · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition.Proceedings of Interspeech, 2022.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition.Proceedings of Interspeech, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.381557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.658544Z digest=sha256:a3a2721682c73a6aebc5d1ac7ed992db20dbeaa4c65f724b288d87c32bf823f0

Observation 0714bb5c-6e1f-4cac-bf3f-280becf7cd01 · outbound

This paper cites MOSS-TTS technical report.arXiv preprint arXiv:2603.18090, 2026.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MOSS-TTS technical report.arXiv preprint arXiv:2603.18090, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.662138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.662138Z digest=sha256:da46596815ae55e57393d60aa12db28bd694410292eee5d3037a28e87be2a20d

Observation a065e00d-6607-4303-afb9-a04c549ece6d · outbound

This paper cites Classifier-Free Diffusion Guidance.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.665862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.665862Z digest=sha256:48abd5b8299d8f296afb8cccc5684b53fc37f61fa3568e99b2ab3ffa877a85a3

Observation 254a2655-16a3-4a86-a49d-20a7c8408ab7 · outbound

This paper cites Qwen3-TTS Technical Report.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Qwen3-TTS Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.670968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.670968Z digest=sha256:e5fffa0baa92a1b9bd81eb686ead532a9d1d20da3a2ba585558ae42cc4eeba58

Observation a488129c-6f09-40eb-8ab4-ef4ab1757731 · outbound

This paper cites Mistral 7B.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.675743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.675743Z digest=sha256:c1acba9a06ed487f34896ad9607d5af9c0eaf93df55220e30fea38522d971dbf

Observation 457fcc10-450a-4318-bd55-271b957749f9 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.367751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.679369Z digest=sha256:f0b288808b2ba192cf4271c5deffba4bee47effbb67253e61cfdc604356b948c

Observation 0556a7c5-451b-4c06-94bb-3e5db3357557 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.682711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.682711Z digest=sha256:c057d29a392882d6b3fbcc89140de7181f64c86b2315596e11ad3f28f3f14bec

Observation 1ee8c75a-397e-4216-9bbc-02691f87df44 · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder BigVGAN: A universal neural vocoder with large-scale training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.354329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.686299Z digest=sha256:a3d61762b53b728923018170f5db0254decfca35453699aaa24e662020bb4998

Observation e16a0760-7edd-4ee4-b87d-cdc3a09394ed · outbound

This paper cites IndexTTS 2.5 Technical Report.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder IndexTTS 2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.689699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.689699Z digest=sha256:37b6d22c78f51b29dad2460617985b87141160b0ea6db6c0e85cb963765b6fdb

Observation 334e7c7d-3041-48b5-a39b-e8664ebd08f0 · outbound

This paper cites Flow matching for generative modeling.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Flow matching for generative modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.337702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.693711Z digest=sha256:9bb99e56cd9d791578aa335401ea52e4a504b2c2d48a92e2ceed856cf55733b9

Observation eb3de67e-0a51-43a5-b732-6d357ce07bf0 · outbound

This paper cites Cross-lingual F5-TTS: Towards language-agnostic voice cloning and speech synthesis.arXiv preprint arXiv:2509.14579, 2025.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Cross-lingual F5-TTS: Towards language-agnostic voice cloning and speech synthesis.arXiv preprint arXiv:2509.14579, 2025

Reference 27

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:36:21.981418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.697166Z digest=sha256:0802bec8bb4348e99fc151326433874891ea3b260b8492a7be21a9f75203f596

Observation 5dd4ecca-953a-4bad-8f56-b72790ca8ef1 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Zero-shot Voice Conversion with Diffusion Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.700908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.700908Z digest=sha256:5b3a0bb71d2a8b598a205f4b5380b626b4ff14265d053a64ae34b7837a99b781

Observation ddc50cef-a039-4e7c-ae10-aad928a72906 · outbound

This paper cites Decoupled weight decay regularization.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Decoupled weight decay regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.705297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.705297Z digest=sha256:b5bd0cc7cff379f669579527ecefdb183cf1264a5f92f5ea56451d0fba593597

Observation 00901e98-f154-4607-9daa-0868ec04a7d7 · outbound

This paper cites Matcha-TTS: A fast TTS architecture with conditional flow matching.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Matcha-TTS: A fast TTS architecture with conditional flow matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.709109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.709109Z digest=sha256:d53c3ef45546c95463776f4c9ca63e2f65c62a05b628a41e71b87c3dfbb28f0e

Observation 68779867-47e2-48f7-bca3-bd55533e7e13 · outbound

This paper cites Attentive statistics pooling for deep speaker embedding.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Attentive statistics pooling for deep speaker embedding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.307044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.715001Z digest=sha256:512b5c643b4642c949aecaa0665f1e3b014024344005fd29f639a54afe4a8ce4

Observation 74a6e69e-a1f5-4461-8d6a-68c182873978 · outbound

This paper cites Scalable diffusion models with transformers.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Scalable diffusion models with transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.719373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.719373Z digest=sha256:13522b519d42c65eb179bb25b3c6706776147e5fa6e10d4fb128a2fed313e9bf

Observation 73795dc4-a4d7-496c-816f-3f3d6294f3ad · outbound

This paper cites Qwen-audio-3.0-tts: Freely controllable and highly robust speech synthesis with multi-stage training paradigm.arXiv preprint, 2026.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Qwen-audio-3.0-tts: Freely controllable and highly robust speech synthesis with multi-stage training paradigm.arXiv preprint, 2026

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.285727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.723747Z digest=sha256:b92ed4c2968cd6388a8eeb62189001537f2668c849447069986f6a941836dcb2

Observation ee23f655-280d-4ca3-83a9-fd35fa2e5798 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Robust speech recognition via large-scale weak supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.727889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.727889Z digest=sha256:038e2c3f7d182087b5d3ee1a390dad474e185e646290db26ff337414df370912

Observation 417f6fb2-4093-4889-b312-c400a4e0a83d · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 2019.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Language models are unsupervised multitask learners.OpenAI blog, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.731493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.731493Z digest=sha256:b85aa2fd4cf8315bbea1808665425ffd3289ea2271ae3f1d52312780a739816f

Observation 1844908d-efb1-4af3-9bc8-dc0d07380c8b · outbound

This paper cites Chatterbox-TTS.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Chatterbox-TTS

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.257628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.736094Z digest=sha256:990c9d92d43895330f510a8f7d798894ef12a8f6233588e01fae8cef4fd3f944

Observation 4ebfec16-3416-4c58-9d42-04eb0f5a098d · outbound

This paper cites Natu- ralspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Natu- ralspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.243165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.740401Z digest=sha256:a4e877bfae0501e09bf6a776a26c56651519220bd109e4962952e48e4c86558e

Observation 2b95ccd5-ab9b-4332-a760-f7c158ac93cf · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.744717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.744717Z digest=sha256:89e708e84e2edf374080b5a850bfb4f8c3835783fa5158446e9fe33555603455

Observation 8cfb784c-0ac9-4fd6-a210-2d260750e22f · outbound

This paper cites CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.748566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.748566Z digest=sha256:02427e760e847e2bf181d3ab2383296edcbfe895354d7494d66257002fa8de91

Observation d2eecb1a-ea0b-4bf8-93c3-ff37f5c41f39 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.752186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.752186Z digest=sha256:db38f3ce75c685249afb7df9db94441c3cf72df6756d84586f0bab36b467a831

Observation 81cb5a0f-c28f-4f44-b65c-1dff527c1049 · outbound

This paper cites X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:21.848785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.756639Z digest=sha256:6a15fed2afd17823e5a5c09c3f492342b2bd3ea00115dfd818e49e239cc92002

Observation a4408303-cfc5-49ef-bf47-55239e67808e · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.760400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.760400Z digest=sha256:8945a10c138797aeabeb7e796cd5ee156a3962b99541f24355ee38b4b1457337

Observation 846ea7d4-6523-4177-a59c-7b1b2f59c24c · outbound

This paper cites IndexTTS2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder IndexTTS2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:22.224921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:36:21.764481Z digest=sha256:cba95ad363732ca68dea4cc2e498aa0204bae6b14b0dc991f064179a7d7736bf

Observation 242cc360-1cfa-4113-afcc-df98209f47ca · outbound

This paper cites VoxCPM2 Technical Report.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder VoxCPM2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.768454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.768454Z digest=sha256:4a29c74983e31117c520301d30d9eb471781dcded366619440efe6cd637a5e4c

Observation 92b450f3-d50b-4a05-9438-f792287449f3 · outbound

This paper cites OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.772183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.772183Z digest=sha256:e1532cce36b377f63114194953453435fa112784ac54be08cae1e083e97b5420

Pith citing papers

No inbound Pith citation observations are available.