Pith. sign in

Paper Citation Record · LEDGER

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.05763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05763 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:16.449290Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.635938Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 138ad91b-1082-4925-afe0-dcdaae9df868 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.458352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:28e0e436595dcdd21d9ddb68eb4e4d464e8df4cbb2ed7eeb5ebf22afa1840352

Observation b5f64cc5-3f35-438e-ac87-c9d6ef9c416b · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.272896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:ef493ab1bc195e3b6385b2e4919f1904e092f13a58964bb82472cf20572ac4a5

Observation 30bb3cbe-ee07-4b8c-98c1-e66a3eca2135 · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.449290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.449290Z digest=sha256:2c9dacf97cc9afeda28c7ab27e947d873c28c157c8b22c555710857118264eb6

Observation 67f74920-be81-43fe-a871-d1ab29ffa4bb · inbound

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition cites this paper.

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:26.135591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:02:26.135591Z digest=sha256:dc8a95522241fccc24f5e8c040e5c3253d16e7d11315bec6028d1113d096e92d

Observation f5a86135-f78e-4daf-8f2e-11ed8cd36458 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.630347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.630347Z digest=sha256:0aca2709e0c211f2100607daa76d5e85e0d93ec521aa21db3e1053aedce4b31f

Observation f022607e-521b-434f-afbb-449dca63c080 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.139824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:a60af676c21a750e6c823dc48070d34cad21733d2a65b59f0f2676ed7990141b

Observation b737ab51-4b48-4e0b-ad2d-e47b00678ded · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:36.184168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:93ec986f4c49c1e2e2f4787ac4efecb6848bff42aa037ff9618ad4a476a3a02b

Observation ed7d087d-a2d1-4db7-9c99-5a4dbbd044ae · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.637490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:d43fa71a0dcbedf5409372180d8efd259a222b10cf5abc02feab7e1ea92645b1

Observation 1315673f-b59b-4c4b-b380-2768e3b7aad5 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:37882c8397977432f4559f95edfdb4073eafc47becc2428761043deb2196e29e

Observation 83f0598a-d331-4df4-ac54-890ddc2508a6 · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.508576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.508576Z digest=sha256:099e861556076522148d5e24c9d1967e6603178d757f9026a75375e149defcfe