Pith. sign in

Paper Citation Record · LEDGER

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2505.17076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17076 v3

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:07.274864Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:57.724078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21a846d2-cdd2-4376-9909-a4cf5715e0bd · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.468219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.468219Z digest=sha256:96d23b2fb34fb76c745a843667fc0b8bdf0c46e4ce06909bdddd8dcc81aa97c2

Observation 98c1dfae-e363-4fb1-98d3-790361e9d7d3 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.525438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.525438Z digest=sha256:a74ddee86b66619153b8d32e059b000da2bb89d3e6fbb8a0cbf6ed83c2f6129f

Observation 79864f0f-c893-49f5-802e-3d94cdd45593 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.596328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.596328Z digest=sha256:e01fce68dfababc2a63d4b65b943754b2a398354ed8da6a171da055b682fda46

Observation d402c68d-3c50-4cf1-a39d-17d51ce622ae · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Moshi: a speech-text foundation model for real-time dialogue

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.765618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.765618Z digest=sha256:84e4bfe50ee0fce71ed46dc8611f989f3fa26b4f4da7e4478d0a8a967e1821e3

Observation c9344925-fe3d-47f6-9de6-681c54e09ec3 · outbound

This paper cites Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.876754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.876754Z digest=sha256:0feb0469b8d1d73ea9d08a981a85035badf38b9c478d94ec5c1ab39cd97a154a

Observation 140e13e9-533c-4fda-9fae-710e5bec802b · outbound

This paper cites GenSE: A Series of Audio Generation Models,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English GenSE: A Series of Audio Generation Models,

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:43:08.046518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:04.967838Z digest=sha256:4821f8334976af247c502132bcf4db5a865e4b6d22c3b05713d967a60d9db934

Observation 44c14ae2-aab0-4e07-b60d-3d12d1b66ed7 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.080883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.080883Z digest=sha256:c88d153c5fe1e6fa5816969673186a85315681ca7ec14092f88df830427271f2

Observation b4435bc8-5bbf-4d90-8269-6431442dd77c · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.226422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.226422Z digest=sha256:98201d4034daef8fc13f60aa9b8acc8c853f3cf0c4a2f177a5d86d2ba17cd3d9

Observation 78e6ee6c-90dc-4461-9a1c-c0fdbf86a18d · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.374927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.374927Z digest=sha256:83a62ebe6b265abb19f3b34a621fa5e57444edfc22148ef2440b26a6ce2c0dfa

Observation 56e0ad82-5987-481e-82ac-daea9cbf5b87 · outbound

This paper cites SMTL: A Stratified Logic for Expressive Multi-Level Temporal Specifications.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English SMTL: A Stratified Logic for Expressive Multi-Level Temporal Specifications

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:07.764249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:05.456608Z digest=sha256:afb17b4992a90bef458018a51feb23a53a8c5a0795f20459c3cd4a6eecb77c77

Observation 89177396-e7e2-43c4-932e-ec319e1bbae8 · outbound

This paper cites High Fidelity Neural Audio Compression.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English High Fidelity Neural Audio Compression

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:05.546290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:05.546290Z digest=sha256:36395d5feb345029245cf977c73525f42954e4241125e633fc1450585454bbb1

Observation cd29d5e4-fd7e-42a2-b6cb-6238da7c5199 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English SoundStream: An End-to-End Neural Audio Codec,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:10.078581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:05.683087Z digest=sha256:c734e98a27148533a3b8e7b108dac85525d10283cfabd6831aa546bd6f23b606

Observation 28d695fd-43d7-46ff-8a3d-090e9a4fe551 · outbound

This paper cites Attention is all you need,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Attention is all you need,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.890487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:05.915626Z digest=sha256:9095ea7e5089fd19af4c04d0fee29791f92453418f533de28fdf04331f179564

Observation 77b97ef8-a72e-4b57-bb9d-3e6fa480b3ad · outbound

This paper cites Neural discrete representation learning,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Neural discrete representation learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.705553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:06.049554Z digest=sha256:72c7cb0ece0fcdcb8d016cd5b9c1e4b8a40250ec0628c91c0b132deb16ff1ec4

Observation bd8c011f-56fb-40f6-8a36-6a36de94f288 · outbound

This paper cites Connectionist temporal classification: labelling unseg- mented sequence data with recurrent neural networks,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Connectionist temporal classification: labelling unseg- mented sequence data with recurrent neural networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.409285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:06.241092Z digest=sha256:8c6c7701ce83b53d022944f74a7e6058ddf43ed59f1cb0f1e8adcaafe22b2aee

Observation 7d4b39c0-8e81-496f-92b7-eccef56ee622 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:06.340117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:06.340117Z digest=sha256:a8c43bd4233342e4b53bee06d96560101597efbf24329d8aa4090f9ef2176b77

Observation 2eb5f9e9-09aa-4949-8b99-a6e9bb932388 · outbound

This paper cites Librispeech: an ASR corpus based on public do- main audio books,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Librispeech: an ASR corpus based on public do- main audio books,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:09.208049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:06.554819Z digest=sha256:575463f07babfb81cbdc7174f5bf8b6b7c2bb6d170299cf38216a34ffb3d6203

Observation d8aafd0a-cd3c-4a4b-85cd-d208cb1617ae · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Robust speech recognition via large-scale weak supervision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.971715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:06.713635Z digest=sha256:4e3c04698cdbe31ca0364c9e50fddaaef1278912af59c3feba7a77e5549944d6

Observation ba1d2e57-1294-46e5-82ce-83a13c71c556 · outbound

This paper cites StepAudio: A framework for pre-trained au- dio models training and reasoning,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English StepAudio: A framework for pre-trained au- dio models training and reasoning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.726633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:06.864778Z digest=sha256:73592120be2920932fa3d48b2328af9c0fde6bf226784bdb06c61251a6e5e034

Observation b997a9c4-edb4-44f9-80d3-942ff55a5ebb · outbound

This paper cites Characterizing Polkadot's Transactions Ecosystem: methodology, tools, and insights.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Characterizing Polkadot's Transactions Ecosystem: methodology, tools, and insights

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:43:07.563131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:06.984573Z digest=sha256:019a2d08b982589cf89b09ae28789dcae6b91c44f0e9c7d7f7f21ee9443d0eb7

Observation 33c999dd-1326-467f-86e1-1f932a255f1e · outbound

This paper cites A three-layered model for expressive speech perception,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English A three-layered model for expressive speech perception,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.525743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:07.126002Z digest=sha256:3e517dd315fe0a8e71ecfbf15e539baaca713d7b518c126c0de44b047df26498

Observation 63460d8c-dfcf-4b75-813b-993897525a1b · outbound

This paper cites Tone recognition in Mandarin Chinese using convolutional neural networks,.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Tone recognition in Mandarin Chinese using convolutional neural networks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:08.297050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:43:07.274864Z digest=sha256:7e97027fbbbc4bd8289d0b1524ca763198285283c65e6f7cdc71fa95d4bd0f57

Pith citing papers

Observation d6e5ce31-033b-4fde-86cd-ac72e05ad0a5 · inbound

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition cites this paper.

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:57.724078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:57.724078Z digest=sha256:a1634ab720bacfa8eb74602b62f921dfbc8a0be1cf51921c3387e29afa34ffb0