Pith. sign in

Paper Citation Record · LEDGER

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2010.11567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.11567 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:56.239809Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:38.607429Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 080ee40c-76b7-47f6-b127-d023a047e778 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.038704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:4589d6f1b2a36b14e18a86d599b4b19e937956f402751d981018d1344f7c5feb

Observation 72b3b378-ef9c-4449-8708-5f8c65c46d46 · inbound

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval cites this paper.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.239809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.239809Z digest=sha256:06024bcd47e4a2018f10f785e49825e46bcbf7a2428b6726b12f948bd4acb0ea

Observation b1404a95-a522-4169-adbf-291624ac1cc5 · inbound

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training cites this paper.

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:56.536930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:56.536930Z digest=sha256:a7e0077bcada0a14e33b30d98b280878243f380b4fee9aaef3ef4db0fc3c1b58

Observation a99b368e-9179-465a-ada5-1305950229a7 · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:21.272793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:21.272793Z digest=sha256:cf83313c979bcdd99b55355508b4d763b42be77424c8dfdc2ee127ede2c54207

Observation ae9dbea3-a6f2-40af-abba-98735084747f · inbound

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling cites this paper.

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T22:07:46.302600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:07:46.302600Z digest=sha256:1528f233115cd2362421581951823e8a867fc5520933eb6996ee8f9392131e84

Observation 9145723c-61f8-4897-9a5c-95f581107e06 · inbound

SwiftF0: Fast and Accurate Monophonic Pitch Detection cites this paper.

SwiftF0: Fast and Accurate Monophonic Pitch Detection AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:05.595288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:05.595288Z digest=sha256:971d4b082c97802b172a5a727bbbb2cb5d93bcface134841bf8f0d539297943e

Observation 3fe8a6d4-6c3c-4f6e-a61e-b88ac4cb3a0d · inbound

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners cites this paper.

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:53.149135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:53.149135Z digest=sha256:446fde46aea3c4533d668e0f509286925d5ff3f9b32bdae8caf237d9d40efb58

Observation ede1d017-df78-4cc2-bed4-1c84314f5094 · inbound

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance cites this paper.

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.785216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:34:15.220263Z digest=sha256:73ca4056383e08fddb7ba7cd4ec784b9cce94e9215037010439c9d57548423ea

Observation 59b7d247-258c-4e58-81c3-665acdc0a55f · inbound

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement cites this paper.

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:59.187097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:59.187097Z digest=sha256:a03bb399c5d5db8ceac15636bcf34c60922d6c74456c16d59d1d3f6242717837

Observation e9434d92-ebe4-4f3f-b63e-32a431447960 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.807719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:34:50.539351Z digest=sha256:130989ca63768b908c289ad4c371b011d08f3f5989d99612db9472db42af96b4

Observation 660a81a9-150a-4a4f-85c3-0cc6b8fda1c1 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:33:13.615847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:33:13.615847Z digest=sha256:c2ff5986e0fa8c7297725ab770d854bcb5df9e6fad164b887403a683fc20b53f

Observation 382d7016-5903-49f6-8b9b-bb98d1d4fa4f · inbound

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation cites this paper.

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:01:07.558745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T17:00:59.421441Z digest=sha256:d44c0859fa23bcda87a41dc2a231e3798802faa546f35d417cb00d53c9d276e2

Observation 99101ee1-bdd1-40b3-8aad-c3249b48d02b · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.656870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:d3ab8485a9058f1cac56527b92cd418989eacea308248134768d2081860c091e

Observation d9f90c94-f0b9-4589-90ee-a3386cf88e91 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.000225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:38870c6c8089442bc82f681751579810b0bd9ce7d9ad65013a10acb6d0f0e022

Observation 5c0903ec-2150-4288-8a05-51b1f7ae5c0f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 129

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.944674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:b3eba180ed54a339104f245a57a2a5daa476eba82e8bca6f9836698b8755bda2

Observation 9e17326b-787e-4924-a8f8-91a541982b3e · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.259741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:a1b42f0d268872a985b590410c3bf29501a4440a70e4799de749b6b0f0eb4592

Observation 1436a86b-d04e-4db9-a1a8-371b85b70769 · inbound

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing cites this paper.

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:06:23.940166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T12:54:20.815371Z digest=sha256:db7f216055be7519189e9d4b9ae08c40b9e7d09c1fdb1d2903421bfbbadabb1a

Observation 69a15ad5-1818-4586-ad04-70fbff2f53cc · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:56:51.761438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:5894a13a3d557cce4c09c587979b5513899fab7a08537cd47564b9767591b163

Observation b89d8ed0-6b2e-4c88-b7bd-05fd353eac66 · inbound

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding cites this paper.

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:47:44.574399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T11:51:25.305324Z digest=sha256:79c169a6e534d3945551ce15d98e6c12b5a0ce52c7da7cae59df4db9003ac54e

Observation 2e069813-4b96-4c0a-957a-6beeb27708c7 · inbound

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective cites this paper.

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:12.259675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:43:35.279033Z digest=sha256:a0098cd60fa17ecc719c9a035c4386899d6fa983ccc22ba4a95025accf6b96d7

Observation d58346c9-cf64-4b04-90ef-49a30faf9bf1 · inbound

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction cites this paper.

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.751682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:07:31.795522Z digest=sha256:fad408bcc1055c3164d8c00ba4bb2299a289489c7817e412f0ce503d2a90caa7

Observation a949a7e9-5d13-464e-a3a1-a3f94fe8517f · inbound

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation cites this paper.

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:38.609190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:53:34.428544Z digest=sha256:3912304d1e90156fce7fdab8f35cd22dba5bffe5915846126e9509ab8e80f6ac

Observation 7f9b7f3b-fe35-4c58-8aef-be0e603c8e3f · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:23.043310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:23.043310Z digest=sha256:b9ed9a25de7d380bcf9ee5951db58d3661616688ce0fe1de0d01a6c5def3dc96

Observation 38a312d5-8ddd-48a6-901d-e15527f0d233 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:27.881861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:27.881861Z digest=sha256:b0d51cec3a3f5fbe160c3c5d63a50edf1a62f78e6249ca6a8996cdde88abf612

Observation a95b1244-7755-434f-a7bb-b75e33b5c60f · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:47.871653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:47.871653Z digest=sha256:f28a1618ab722ee263fe1a225262d8510df6bbe03e7b54bd06c6d8e4bffa3519