Pith. sign in

Paper Citation Record · LEDGER

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations

As of 18 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.05056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05056 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:17:54.873072Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2bc0535-d1be-4d1a-b77b-d629698f9d6e · outbound

This paper cites Wenet- speech: A 10000+ hours multi-domain mandarin corpus for speech recognition,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Wenet- speech: A 10000+ hours multi-domain mandarin corpus for speech recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.288705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.731469Z digest=sha256:6ee6be84013e6946da78ed99e14ad85b1d84f12985bb4871a2bba7b34b1cbfdb

Observation 77a4ced1-7334-413d-b218-15832b16259b · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.736082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.736082Z digest=sha256:6d48cc95430f65a983cca15b488c060aaa85370dbe453fe18f5a8cdd59db7f15

Observation b42c57bd-82b5-480f-a914-2489ef6fd6eb · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Common voice: A massively-multilingual speech corpus,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.275363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.740591Z digest=sha256:58972ad8e2c6a01dece51f9c456bf0cbaaf4cd0df4044a4be3197866365e5037

Observation f8f103d4-e055-451c-82be-9733936bc31e · outbound

This paper cites Robust speech recognition via large- scale weak supervision,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Robust speech recognition via large- scale weak supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.744860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.744860Z digest=sha256:e98da42ed53319e7804e36e9d4800aa0d1a546a62b3f8cd339e9a47a982f7fd1

Observation 55023dc9-7f4e-4384-98e4-c7f3ea493964 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.749572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.749572Z digest=sha256:6643a01ad174ce4ccabdc125d7e34a8840e6d0bd46ba04d4a4c689691e4364d3

Observation 3d982d91-c326-445a-9352-0b50476c0277 · outbound

This paper cites Moonshine: Speech recognition for live transcription and voice commands,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Moonshine: Speech recognition for live transcription and voice commands,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.252681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.754270Z digest=sha256:0439d2eed8cc6f84410d3f9e5ccd42a4e75023eaa18127060431e82dfa3633f1

Observation c26f3345-058c-47a6-927e-1fc6efdd23e9 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.758777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.758777Z digest=sha256:759606fef13c6aa88012b7c1cab5f1546666026ba9881f5b7f30a03308444734

Observation 48bc5af6-780d-4404-9472-37f1c386c58f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.763554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.763554Z digest=sha256:92a93784cee9c197284277913a721faa21b84d79aa83ca41580c607a2f72bbd9

Observation 511157f3-7058-45ba-81eb-cc42ad099429 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.768702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.768702Z digest=sha256:228e229ec626439f17011536d4931d7c6b79272008e1a81069f43e5b115e9476

Observation 1970c52d-02db-4ad8-abf7-549d375330ab · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.775167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.775167Z digest=sha256:9e74ae173d49122b10135f73bd9f3d7e602d70a485110241ef0f7ee605ca7341

Observation bf6b5d5b-4d10-46c2-af9a-2efd7e394ffa · outbound

This paper cites Cmu wilderness multilingual speech dataset,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Cmu wilderness multilingual speech dataset,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.240176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.780086Z digest=sha256:aec31b8393dbc92a9f7518eaf9798f50c23a0ffab03523ab6fe701ea00f28944

Observation 27164b1a-e5bd-423b-9acc-2107e568e4d9 · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Scaling speech technology to 1,000+ languages,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.224591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.786384Z digest=sha256:5bf65eec3f53e2c1131e66c74ff0ac7c4520e0271353c2fd61ee03de023cb8db

Observation 26ebfd01-f80c-4bb7-a0d0-596ae32d5516 · outbound

This paper cites Experiments on speech synthesis for teochew, can taiwanese help?,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Experiments on speech synthesis for teochew, can taiwanese help?,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.210405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.792634Z digest=sha256:ca2700ba8073061c2cb9551ad095af964a5e6c29fcd3ede7a211d7c2030f372b

Observation d57dc0b9-6310-4325-b80a-c679a42524c7 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.797759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.797759Z digest=sha256:0571a93d93169efe76d3ee33dcafad6e646f84980a8395c16def006c62c9e140

Observation 52eda01c-a102-4c54-829b-22cfc640800b · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.802856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.802856Z digest=sha256:d2eaeb35dec0dcbe52060bb1d2a69a83748d41d04944772f8aee4e759d4a7324

Observation 66fbd766-3d6d-4ae5-86ca-c0cb870d2574 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.807671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.807671Z digest=sha256:8e437a43cdbc8b5ac230d2623d05b3b1c7df614f45a076ba7dd2e32fa488664e

Observation 7c668f53-36a3-4271-b271-ef9cfe4e33cd · outbound

This paper cites Low-Resource Self-Supervised Learning with SSL-Enhanced TTS.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Low-Resource Self-Supervised Learning with SSL-Enhanced TTS

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:17:54.944368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.812211Z digest=sha256:3f422aaca413f77291c8a5818ebb48d4e6a882086439cb039a3c2938065aec72

Observation 110d9f52-3a06-4157-ae2e-23b637c2c8a4 · outbound

This paper cites Improving automatic speech recog- nition performance for low-resource languages with self-supervised models,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Improving automatic speech recog- nition performance for low-resource languages with self-supervised models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.168230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.817181Z digest=sha256:03a16a48f2a371459a5b67f80580736023817e36f8e9b57d020a8ec6e966d585

Observation ba0c356c-d37c-4ecb-bb76-51d487b49d23 · outbound

This paper cites Lrspeech: Extremely low-resource speech synthesis and recognition,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Lrspeech: Extremely low-resource speech synthesis and recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.155258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.821292Z digest=sha256:1a261aa057d3a076344815bd99706e09453df412750b9de2aa8a7a5b7e023c3a

Observation 1cddaec1-250d-406d-b315-89bd923b0947 · outbound

This paper cites Autoprep: An automatic preprocessing framework for in-the-wild speech data,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Autoprep: An automatic preprocessing framework for in-the-wild speech data,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.143430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.825342Z digest=sha256:4de770160c63575e56d49003545cc15d20728ba023183baf6e028fa102c5ade1

Observation 1d4721a1-6875-4810-84d8-762a6b8947fb · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.131688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.830556Z digest=sha256:f9c7108336b1e1bf9e4c3b89c4d67318d8934aecf806518e23f1587f7d5f8cad

Observation f3b47410-c53d-45d6-b32e-434c150de927 · outbound

This paper cites AISHELL-3: A Multi-Speaker Mandarin TTS Corpus,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations AISHELL-3: A Multi-Speaker Mandarin TTS Corpus,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.118196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.837119Z digest=sha256:f5acf5b87bf7ad190a38ba595cd544f9a20c29474ee096c92b2220089da60e01

Observation 49ff6398-293e-4466-ac2b-7f553dd9d45b · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Libritts: A corpus derived from librispeech for text-to-speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.101730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.841056Z digest=sha256:3b27cb6594c18dd2ee4f0bd6ddb801ce2972c43909141ef77b16ce85e0b7f085

Observation fd6b2216-eb48-4b38-bab8-9ee0fa48e66e · outbound

This paper cites Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.086136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.845041Z digest=sha256:ecf0a9c0d906977c252468f08c673dc8c4438f43cd81f5cb344e98762cc9134c

Observation 9111284c-c78f-44d8-bf6f-df645d1d2a47 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.850247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.850247Z digest=sha256:70b1c3d5836aa9101ecafa6ea791637f57d14184cf31d397128e1dc9f66f4336

Observation 3eff2833-6a50-46d9-8669-9fcd4c4bb481 · outbound

This paper cites Dnsmos p. 835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Dnsmos p. 835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.072356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.854570Z digest=sha256:1caa51f08939e3230efedb747f9ecd8853e44e2ac9daf5f727808ee2792caf6d

Observation 438585af-4b8b-4481-87d2-c974c242c67e · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.059595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.858921Z digest=sha256:6a6e87b83f12702f6b933104b10a71bd9008b916f6558a0441f1ef74519bbb9d

Observation 2f6f1a83-2544-435e-bfe7-da3c7b64f16e · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.046348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.863685Z digest=sha256:22d04d3c3a00455f858317ed2bcbe5d6635076d48cac8a59183163ad0c039349

Observation 22fb8816-16fc-4cc7-943a-19657ee0c397 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:55.034255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:17:54.869295Z digest=sha256:58c7f5fe8778f052ee4bd1d8d46679738b60f9039150a50d6871014157a0dcbb

Observation e14259d4-9735-4a60-ba17-4cd740434d22 · outbound

This paper cites fairseq S2T: Fast Speech-to-Text Modeling with fairseq.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations fairseq S2T: Fast Speech-to-Text Modeling with fairseq

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.873072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.873072Z digest=sha256:85c3958c266e253f0e26e1a7ffc3b8b6f346f7847449e5f3be660608688e44a7

Pith citing papers

No inbound Pith citation observations are available.