Pith. sign in

Paper Citation Record · LEDGER

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.05763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05763 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:09.630347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.635938Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 138ad91b-1082-4925-afe0-dcdaae9df868 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.458352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:c7875025000c861797c2065171b8f40a9eeeb19703b3b7c5238dc801aa8071bd

Observation b5f64cc5-3f35-438e-ac87-c9d6ef9c416b · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.272896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:d2f0f177b935d38f6d6e10a04279f81bf09f1cd3a795cdda3e8d4171e721cfca

Observation f5a86135-f78e-4daf-8f2e-11ed8cd36458 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.630347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.630347Z digest=sha256:0aca2709e0c211f2100607daa76d5e85e0d93ec521aa21db3e1053aedce4b31f

Observation f022607e-521b-434f-afbb-449dca63c080 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.139824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:5f6644004c8a9d5fcd7bf9ffcb2f2c5ad67cfe9b4d92b70916d74775ebce7517

Observation b737ab51-4b48-4e0b-ad2d-e47b00678ded · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:36.184168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:bf763fb0a9148d1c719e5a1766c74c14764e4b1fdcb3cd994db44c56337238ff

Observation ed7d087d-a2d1-4db7-9c99-5a4dbbd044ae · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.637490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:cb210e97343fa7fa332dd817db48cf9be456d3c9324d891699bb93b2eee0f661

Observation 1315673f-b59b-4c4b-b380-2768e3b7aad5 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:37882c8397977432f4559f95edfdb4073eafc47becc2428761043deb2196e29e

Observation 83f0598a-d331-4df4-ac54-890ddc2508a6 · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.508576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.508576Z digest=sha256:1a7cc80874764a59a38b3a13f0a34093d581f69c62bae4420dee7b4bb68e18d7