Pith. sign in

Paper Citation Record · LEDGER

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2010.11567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.11567 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:20:25.368078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:38.607429Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d85fa9b8-b4ef-480a-a964-8c46945f38de · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:58.008994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:58.008994Z digest=sha256:e9b187b79de2ae144e8b0b0c0c5481e06fa08f94ac8df785eb58a56034e3f65a

Observation 8e2bd706-61c6-48a5-89eb-acf0618e489f · inbound

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection cites this paper.

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T04:35:12.394750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:35:12.394750Z digest=sha256:b358d474b451b290e1a47a8606fda945a1d69e287a4812516cadff903d28ff09

Observation 080ee40c-76b7-47f6-b127-d023a047e778 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.038704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:319d3e865d46eec94235eaaf1b503c1fe56f526fedb086aaf66f71137c59d4f4

Observation b977b766-0fdf-4394-aabe-9214b623aa62 · inbound

Inclusivity of AI Speech in Healthcare: A Decade Look Back cites this paper.

Inclusivity of AI Speech in Healthcare: A Decade Look Back AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:20:25.368078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:20:25.368078Z digest=sha256:307f2b7e49a204b3ec4c45cd3f07871b2b16ebf2f706e3b8261dc2d86d770570

Observation 72b3b378-ef9c-4449-8708-5f8c65c46d46 · inbound

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval cites this paper.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.239809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.239809Z digest=sha256:e0e5df3a4aa49a2c60c8f1287f6ec918b641764153b8aa14dc81659ca39a7166

Observation b1404a95-a522-4169-adbf-291624ac1cc5 · inbound

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training cites this paper.

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:56.536930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:56.536930Z digest=sha256:3bf4e4d9617786d7b1dcfe105ad5ec82259890b1c68a77426dcbcb12e6aa22aa

Observation a99b368e-9179-465a-ada5-1305950229a7 · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:21.272793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:21.272793Z digest=sha256:87e8eb7afb08d51a95e0144767467582f0e056af436cd141089840f8c1e934fd

Observation ae9dbea3-a6f2-40af-abba-98735084747f · inbound

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling cites this paper.

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T22:07:46.302600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:07:46.302600Z digest=sha256:db32d532a51f6a5234ba1d9d2cfe305f1ae45d62122eae1560d58804cc87d47f

Observation 9145723c-61f8-4897-9a5c-95f581107e06 · inbound

SwiftF0: Fast and Accurate Monophonic Pitch Detection cites this paper.

SwiftF0: Fast and Accurate Monophonic Pitch Detection AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:05.595288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:05.595288Z digest=sha256:8459066559efb4c8d5ed879b55cd331e22b091ebea1c486b50e028364dec565d

Observation 3fe8a6d4-6c3c-4f6e-a61e-b88ac4cb3a0d · inbound

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners cites this paper.

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:53.149135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:53.149135Z digest=sha256:27b806c8466a0c7ec99d290d2f8ac181eebf2b81e57a08d8adc21ce42f4247d9

Observation ede1d017-df78-4cc2-bed4-1c84314f5094 · inbound

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance cites this paper.

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.785216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:34:15.220263Z digest=sha256:bd825c3b1646db804e87b67df3b4365079deded838099f262f95ceff45cf4722

Observation 59b7d247-258c-4e58-81c3-665acdc0a55f · inbound

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement cites this paper.

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:59.187097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:59.187097Z digest=sha256:3aaf7c67071207bb60a42a8adba64434ca095480871c8641b060bf210ceb9627

Observation e9434d92-ebe4-4f3f-b63e-32a431447960 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.807719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T20:34:50.539351Z digest=sha256:a770baac8e4dce70c9e839c773b6e2b8e98dfedb4b0d5bb94657c383eeb8eb0a

Observation 660a81a9-150a-4a4f-85c3-0cc6b8fda1c1 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:33:13.615847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:33:13.615847Z digest=sha256:5e3e7b89d87588e0e68be27bd8262b3ee21bb207904b03ecc9fbc5c9b0e71703

Observation 382d7016-5903-49f6-8b9b-bb98d1d4fa4f · inbound

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation cites this paper.

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:01:07.558745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T17:00:59.421441Z digest=sha256:c77ed1b81267c79b8ff3b02a953d5521362f6f8bf72c41f74f8a81549f6c49a3

Observation 99101ee1-bdd1-40b3-8aad-c3249b48d02b · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.656870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:ed3136bda224ba32254e898a2cc9d4d2bce87b52907dc253f49f3923eaaec5bb

Observation d9f90c94-f0b9-4589-90ee-a3386cf88e91 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.000225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:d4f3364825d18e9e4581cd7b89412289196339a8ee6a9824c524e6321d3050fd

Observation 5c0903ec-2150-4288-8a05-51b1f7ae5c0f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 129

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.944674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:c2bfdb8956ab9ac06d79949dd3a876502dcc1dfad5e457d21ae36ff0fb10e832

Observation 9e17326b-787e-4924-a8f8-91a541982b3e · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.259741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:f3496d58ca17719f09c5622d293e8ade49bf9bee6745b793291cc98d20273edd

Observation 1436a86b-d04e-4db9-a1a8-371b85b70769 · inbound

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing cites this paper.

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:06:23.940166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T12:54:20.815371Z digest=sha256:2b47d1b66220483f13cca77f0eed53a8c11bbe66224620ca5c4190c9205860c6

Observation 69a15ad5-1818-4586-ad04-70fbff2f53cc · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:56:51.761438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:42163a0bb2ad2d09088f2f649e74eac638f335a5b6a89b63143e2057473d0ab6

Observation b89d8ed0-6b2e-4c88-b7bd-05fd353eac66 · inbound

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding cites this paper.

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:47:44.574399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T11:51:25.305324Z digest=sha256:f156f0405690fc7e75fa46bdba47402025c17889b875c888c40642b95914dbcc

Observation 2e069813-4b96-4c0a-957a-6beeb27708c7 · inbound

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective cites this paper.

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:12.259675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T08:43:35.279033Z digest=sha256:a45d56b476722f408be51ee6e2716710813c5dc6d2cb13a8f420e53b2986ede1

Observation d58346c9-cf64-4b04-90ef-49a30faf9bf1 · inbound

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction cites this paper.

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.751682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:07:31.795522Z digest=sha256:13de5070bfa253ec9efdd3f969d19699801c6664a8d251c8a91a7a55e4007daf

Observation a949a7e9-5d13-464e-a3a1-a3f94fe8517f · inbound

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation cites this paper.

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:38.609190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T12:53:34.428544Z digest=sha256:c29ac3e4e10c465f58db0a67ebc2b8cb4ffb398ae9ffe89c4f4371e3020862ef

Observation 7f9b7f3b-fe35-4c58-8aef-be0e603c8e3f · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:23.043310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:23.043310Z digest=sha256:5bbd6461f93df73abea465a746cc978eb5de464e893ff211188ed359b3af6adb

Observation 38a312d5-8ddd-48a6-901d-e15527f0d233 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:27.881861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:27.881861Z digest=sha256:a8e86694c80e1ec4e70ef1f03a268ed5cbc39a36757e43a40eed023da52e5241

Observation a95b1244-7755-434f-a7bb-b75e33b5c60f · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.871653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.871653Z digest=sha256:d477e6d977fb1158d6bcfd2b920a1d92cbe088d941792c0906624f5c22b32b75

Observation d0303442-e53b-4903-88a3-74435f2c5e76 · inbound

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction cites this paper.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.679417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.679417Z digest=sha256:cda697992da7d7d09f125d38e4bcfbae6d62f771b2a5dbeabc425ebd2aaae876