Pith. sign in

Paper Citation Record · LEDGER

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

As of 20 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2607.21042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21042 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:40:31.629008Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fec16d61-d81e-443d-aed7-23f47a919d30 · outbound

This paper cites IndexTTS2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs IndexTTS2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.242144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.242144Z digest=sha256:7125aaeb22e9dbe06b8ffaa5ce02c5d587369ca8b9704082184ccca84b6860a9

Observation d6f1d8cf-1a05-46ed-b690-81d6ebccf83d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.323776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.323776Z digest=sha256:424f033a4c5ec9f5445ace6e95c4b2e02873075fb2c2cf744d6618ca91886f99

Observation b4ee5dff-afe5-4298-a849-0a991a309525 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.396520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.396520Z digest=sha256:772af3d30f846e3b67b2f3e319c1bc3ecee4bb962482f4940878bc1e7ff2d701

Observation 96e8430a-ea47-465d-8339-2ee17845339c · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.509915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.509915Z digest=sha256:cf417b882c4e691e71a989ed5a876127f19a2a11ad34e6c274bd73536d6b5d59

Observation 69ce9690-b3f7-48be-9344-2ff340b19901 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.564412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.564412Z digest=sha256:49c72e069f4cb95706bd85af7a750e4812dce8057cf1ceea0c211bdd4549c4f0

Observation 982ae0b7-0e29-4b7e-909f-bbe192749aeb · outbound

This paper cites Qwen3-TTS Technical Report.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Qwen3-TTS Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.601726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.601726Z digest=sha256:b6a7a9e7433bf4a086681f9ad7b4922763ea76ff410ec8f90588b58317d74b74

Observation c1eae712-ec60-4c69-95c0-911f42e4aa10 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.699814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.699814Z digest=sha256:c6b90e9f7043689c76999176e88a1a615997fdc1fc1a934338262b6dade8cc06

Observation 62330e2b-7265-45bf-b130-cffe8ce10383 · outbound

This paper cites Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.846912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.846912Z digest=sha256:5a93d7698ec0d76b523f4bdb63800302f195e162a2e3cb020cc92c79bde03390

Observation bf11fb27-ee48-4b88-b5fd-f345fae17195 · outbound

This paper cites F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.013513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.013513Z digest=sha256:31ebc08bb3af078963adc98b2da5da21760bbccec0f077eef7591a460e6d8d1c

Observation 66841eb7-e957-4f41-93a6-585e617657eb · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.181140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.181140Z digest=sha256:412672e98327d6abdba045a9703b28093b2a5768fffcb5f23593db4e83768bff

Observation 38630a00-faf5-4be8-9e30-678341d467ab · outbound

This paper cites E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.346583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.346583Z digest=sha256:796de76ca8d12c7c04fcfe0d2a9589486a33bb3c5330b22b151f51e82cb92dc1

Observation 9403fed2-bd57-4861-8c89-584338ba551b · outbound

This paper cites ZipV oice: Fast and high-quality zero-shot text-to-speech with flow matching,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ZipV oice: Fast and high-quality zero-shot text-to-speech with flow matching,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.514609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.514609Z digest=sha256:0ae40ea6eb105588bbf18f8291182f5884559d7bd54b983f77e8a77452b4a26d

Observation 31622478-221e-469f-b918-289070be5762 · outbound

This paper cites InstantSpeech: Instant synchronous text- to-speech synthesis for LLM-driven voice chatbots,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs InstantSpeech: Instant synchronous text- to-speech synthesis for LLM-driven voice chatbots,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.656148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.656148Z digest=sha256:f5477aea0536fcc683b0d363a8642550bfd82d2e7654e2c0fc2d7bcd1f479f73

Observation 19192bbc-4e90-4097-92f5-3afb94e74293 · outbound

This paper cites NVIDIA TensorRT,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs NVIDIA TensorRT,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.824143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.824143Z digest=sha256:34e02deda051c4638d601cede45570998b9eca212bcdf609c932741f4ad7ede5

Observation 86f99114-cb22-410d-baca-080180d3b210 · outbound

This paper cites TensorRT-LLM,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs TensorRT-LLM,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.954933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.954933Z digest=sha256:d77da0e63619e911bbf913b3414fb2b333bfb40102b7c2ade770e36557657363

Observation 24e9842b-7211-4936-bba4-543e5ce34603 · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Efficient memory management for large language model serving with PagedAttention,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.041783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.041783Z digest=sha256:3ddbb6bcef26bc2bdc995b7097221c9b90b9d88567d37da328f034a1b548e60d

Observation 2a0c6fb1-3d97-4155-97db-0d5d453e3ac6 · outbound

This paper cites SGLang: Efficient execution of structured language model programs,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs SGLang: Efficient execution of structured language model programs,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.152701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.152701Z digest=sha256:344da527291e34803f7c4f389eb0f4f27513e445f50b5bd680c1d80392c1fdc1

Observation 98505226-bd62-453b-bae7-f3951934fe91 · outbound

This paper cites ONNX Runtime,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ONNX Runtime,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.228782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.228782Z digest=sha256:3b7d2a0ef90c4b93a1342d7724dfba783d381af3858c69738b938c8b14967293

Observation bf5dbdbf-4d1e-4b8f-8f10-633bd713f2d9 · outbound

This paper cites Language models are unsupervised multitask learners,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Language models are unsupervised multitask learners,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.342098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.342098Z digest=sha256:986d7a10d9157b639c3327b382e0dcd948d39a3cbf1c27b2ac71d88a68220821

Observation 21d5285d-f0ab-4207-b46e-2eeb66e0ea26 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.429613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.429613Z digest=sha256:b84b7fee962328a1608c7b2c2a354961e3dd98e9911542fa68e0303b93dd4c60

Observation 12e67d0c-c9a7-4959-8eae-1defe7b0c0eb · outbound

This paper cites Flow matching for generative modeling,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Flow matching for generative modeling,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.553448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.553448Z digest=sha256:1c97261754b640162bfebb4c0b67bc63775198ca3199b11f408ee4ad3a530470

Observation 178b5e08-1039-4ddc-adf0-a0e355e094af · outbound

This paper cites Scalable diffusion models with transformers,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Scalable diffusion models with transformers,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.648884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.648884Z digest=sha256:01bca4d3581d9fd8ceb691210dda504045c00707a3188e12a7b96a05b6c8852e

Observation f122820f-8cdf-4ba4-877c-474326dcfa7b · outbound

This paper cites Classifier-Free Diffusion Guidance.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Classifier-Free Diffusion Guidance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.739573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.739573Z digest=sha256:650f131a1018ca794c87c3aeb4f3dceac7261e7422473dcbffa3c468ea763829

Observation e5d0f3a8-6f14-4794-95e9-61cf4c3f3ecd · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.851867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.851867Z digest=sha256:14b0811820428d0fa455c00e140d82812b6a127203159b746a86ffc129b43534

Observation 5644a396-7193-45a7-9609-94d1141bbfeb · outbound

This paper cites W2v-BERT: Combining contrastive learning and masked lan- guage modeling for self-supervised speech pre-training,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs W2v-BERT: Combining contrastive learning and masked lan- guage modeling for self-supervised speech pre-training,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.993395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.993395Z digest=sha256:c2c8766971a9d79f47bcbc6e4a42450c0e01b99f9770ab758a0ee2508cd8f17f

Observation b32c882a-d706-4c04-87e0-27129d8f2570 · outbound

This paper cites Neural discrete representation learning,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Neural discrete representation learning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.127935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.127935Z digest=sha256:6e752fa5efb1c526c5672d64175797ae0256d3f0b753517dbc57348084e7798e

Observation e823c9e9-3761-41c9-8ae0-0e6045ecd682 · outbound

This paper cites CAM++: A fast and efficient network for speaker verification using context-aware masking,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CAM++: A fast and efficient network for speaker verification using context-aware masking,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.329231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.329231Z digest=sha256:4ba35fb6fb5fb2b2f94c0449f924bb49fc073d792bb1812b76ee2b9cf608b096

Observation adba63e7-9b82-4b95-9037-b5d133917ecd · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.520790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.520790Z digest=sha256:b5d825efe42a51e513fcf691ea9bb1c4fde4adfa46e3373b26c01be2891cce5d

Observation 93436678-4bc4-41f6-9dae-44af003e61e7 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Common voice: A massively-multilingual speech corpus,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.659984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.659984Z digest=sha256:7b2946e968d9bf03937dc57c3492aff41386953267f00551f4de5857841f00e0

Observation 8f384ed3-a057-4f77-a7bd-93e8e741b587 · outbound

This paper cites DiDiSpeech: A large scale mandarin speech corpus,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs DiDiSpeech: A large scale mandarin speech corpus,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.811354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.811354Z digest=sha256:218319c1af4ad4123553e5dc5b1664664f48659a56ac1e87ac5e9e973c4f2175

Observation d4c71c9f-d584-46f2-b0b7-c67e4afd8fab · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Robust speech recognition via large-scale weak super- vision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.981779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.981779Z digest=sha256:8b5c58ee25fd1cba68cd6afab688cd4b88b24ce24c5cc178b52cde1af6152d77

Observation 67d7bdf6-6f35-4641-a6fc-4d39b8665843 · outbound

This paper cites FunASR: A fundamental end-to-end speech recognition toolkit,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FunASR: A fundamental end-to-end speech recognition toolkit,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.162278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.162278Z digest=sha256:b5675c846039b61d6bc36b4996926731daa2cff42405c283b791b771cf060dca

Observation a80f1a77-2a32-4105-a6a5-d30f411bdd55 · outbound

This paper cites ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.274726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.274726Z digest=sha256:6fb04859d1eee2c4fbf2915935c11d3b843253372ba1c2af0cfc8b4c2c532acc

Observation c09dd944-95cd-4d36-9d19-c2fe9cdbafea · outbound

This paper cites WavLM: Large-scale self-supervised pre- training for full stack speech processing,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs WavLM: Large-scale self-supervised pre- training for full stack speech processing,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.417593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.417593Z digest=sha256:f489512c00f7a0cb4efca7100cf2c3db0558fe6553f84ca92a3b869a635148e9

Observation f26d5304-4e82-4d12-b4ec-f0eca8ea1b43 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.629008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.629008Z digest=sha256:5bc880a2618cb1e1c2bd537c803906afa0fe111b4eb091ef2ec025958a2e72a6

Pith citing papers

No inbound Pith citation observations are available.