Pith. sign in

Paper Citation Record · LEDGER

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2505.17446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17446 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:02.022181Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:49:57.859145Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:50:02.310312Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cf812e2-ced0-49ae-874a-a2004ad39b8e · outbound

This paper cites Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:50:02.409712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:57.859145Z digest=sha256:b1141fa18e1a7ab4a64e1a309d4e2f9a807bc6e076609887bf8f467b03783783

Observation 0c1fb68c-fbc5-45de-b901-2f473c3febeb · outbound

This paper cites Throughout this study, we used HuBERT [7] as an SSL model and extracted representations from the ninth layer.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Throughout this study, we used HuBERT [7] as an SSL model and extracted representations from the ninth layer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:09.461214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:57.887646Z digest=sha256:8be6b3bd238f4b6590e03b949fc430ea2dc8b05196370d1210a49fad8f28360e

Observation 139ffd34-fb6c-4203-bc97-72dd762a371d · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:09.330352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:57.995598Z digest=sha256:451bcb1e0ffdcb21b43e569d003daf78cd620569b4d8df9a7c41de9bfe1cfd6f

Observation fe7a207b-24d1-47c7-aa0f-b692ededf434 · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:09.205464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.128986Z digest=sha256:45379efb7f7181a552a9c37303a583327af402fa260ae9240c3172d892419e97

Observation 80c5be5a-d540-48db-a5fe-c97e88c5ee36 · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:09.076506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.238918Z digest=sha256:50d70a94cb2e8be1dce65339d67d8aab484c41763a27e5d7f86bb83a4ea47766

Observation 6f133b16-38b6-4852-aea0-89f614884e69 · outbound

This paper cites Dataset As a training set for SLM, we used LibriSpeech [17], a 960-hour English audiobook corpus.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Dataset As a training set for SLM, we used LibriSpeech [17], a 960-hour English audiobook corpus

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.949260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.300160Z digest=sha256:9806b1f1367bd282848a7f7fd484f9bd13d52addf90d3c7ebcf2582168de2490

Observation 0845f3d5-355d-4e8a-974b-cb9c7422d12a · outbound

This paper cites Figure 2 shows results on fixed boundary settings.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Figure 2 shows results on fixed boundary settings

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.846595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.407640Z digest=sha256:63620780089e67d71307d03b1ba7203d670e4398abb076fab7f0f9d8473e419d

Observation c225dd39-0b87-4151-a99c-c7a4a2ac90da · outbound

This paper cites yonder".

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models yonder"

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.730832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.481649Z digest=sha256:100a63c9cf8fdafa54356a9542bf2ee826776a883c63e93a544d63c91f03085e

Observation 564c9789-a8e0-4a74-943a-2d439fbab684 · outbound

This paper cites We conducted mul- tiple speech tokenizations based on the combination of the fixed/variable segmentation and the cluster size.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models We conducted mul- tiple speech tokenizations based on the combination of the fixed/variable segmentation and the cluster size

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.591058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.570939Z digest=sha256:d0b660fe00535cb8edbfefa323201594c4c1adc73470e51e8f503544bb857d94

Observation 3fd089f2-4175-4795-b3fc-07373eed3187 · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:08.423077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.640221Z digest=sha256:1c8239edff183724f148e53ca4a92142cce2b1caefde62d2ed03c22d7aa538b3

Observation 87f1e936-e923-42bf-a50c-0b65c1b67fcc · outbound

This paper cites On Generative Spoken Language Modeling from Raw Audio,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models On Generative Spoken Language Modeling from Raw Audio,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.193845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.753979Z digest=sha256:91a98ae515a161dceab3298a95c22785d0bd19b6face59cca8c69f2fbf870fb2

Observation 70af3af0-d481-4544-949c-607c1f9a2449 · outbound

This paper cites Textually Pretrained Speech Language Models,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Textually Pretrained Speech Language Models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.028899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.883691Z digest=sha256:de697bd2ed3b75fd6fe10bcd1dc6cf793e25904d1a3957e8d3a13612c7bea77e

Observation 5a3393b4-1624-43bc-8954-f4f254b0e0ba · outbound

This paper cites Audiolm: A language modeling approach to audio generation,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Audiolm: A language modeling approach to audio generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.884237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:58.963630Z digest=sha256:9b79d916531abfd91cea3ce52bc9250b7fd3c6da158868004f0d500fe8df11a4

Observation 2a8b485e-ab94-43da-9934-9c0545480a8d · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.713715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.048288Z digest=sha256:18ca07c170d902a7820d8adee4c04a3c809a031c45922917da61beea56fb31ec

Observation 3d90e8be-6ab2-4752-bf34-49cb88ce772e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Representation Learning with Contrastive Predictive Coding,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.558127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.162114Z digest=sha256:6ddad526595dc1dae1489fd5da8bb69011969770b71209a2bb47e75252089cb8

Observation 399a0d2d-8807-4c88-86db-a7dbe180c12f · outbound

This paper cites Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.379926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.229667Z digest=sha256:1209edad560208d600c0d607add976feb26466df5da6ed118cbdac51fc1d8d00

Observation a2b91708-4c6d-450d-be67-836a8f2da1a2 · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.145029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.368364Z digest=sha256:72d4bde14c51bd3735099071b107dc186fbea95ca93ea888935b6d271d9b6cd7

Observation 5cb4db3e-8525-4a82-9f68-88280f53a06e · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.928005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.502483Z digest=sha256:42c7b3c4a84feffb1387dda3af1a3077570716b3f8a77bc1e0e5ac227159e90b

Observation 1bc6e752-38ab-46c8-b9ff-bbc5bdd885e5 · outbound

This paper cites Generative Spoken Dialogue Language Model- ing,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Generative Spoken Dialogue Language Model- ing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.699670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.637090Z digest=sha256:f236ad663141e8dafa2645688130f0c4d8ddafcc19ffda6538aa4eb12e15b74f

Observation 0d1665fd-9f77-4ee3-8638-31813885e252 · outbound

This paper cites Direct Speech- to-Speech Translation With Discrete Units,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Direct Speech- to-Speech Translation With Discrete Units,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.423491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.792124Z digest=sha256:941180c5b57d4300f139f1e74bec624eab3115d5bdaa17db77b7f64d5acc5fde

Observation 1c98c5ae-2af5-432e-a24b-6fe9e72d7ae4 · outbound

This paper cites Text-free prosody-aware generative spoken language modeling,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Text-free prosody-aware generative spoken language modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.151214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:59.930171Z digest=sha256:a51262f05b6e8d20134b053d7094f780052222544220cc3cb6f50a0088753a81

Observation 00f4e789-1c8c-404f-ace8-d8e74dbf0b7e · outbound

This paper cites Attention is all you need,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Attention is all you need,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.922539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:00.108485Z digest=sha256:163a3309e8375efe8c7a03f590f2504895414b44cb9839aaf9b39e1c16a93362

Observation 2c6830ad-de9f-453a-9575-21f2fab6f352 · outbound

This paper cites Self-Supervised Speech Representations are More Phonetic than Semantic,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Self-Supervised Speech Representations are More Phonetic than Semantic,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.697064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:00.225350Z digest=sha256:73293032a6b4f7efedcbb8fd788ea4c2f219a2716a01aa82e9b8f0a428569a0f

Observation 3639a5ac-d4ab-4078-8b27-57f965f6b5a0 · outbound

This paper cites Generative Spoken Language Model based on continuous word-sized audio tokens,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Generative Spoken Language Model based on continuous word-sized audio tokens,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.437143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:00.329313Z digest=sha256:324521380ae76a01028b483996193e98b2dd468d809cdd3557738bcd82f72c15

Observation 45e2ab50-0b6a-4966-8c70-02191b2dff11 · outbound

This paper cites SyllableLM: Learning Coarse Semantic Units for Speech Language Models.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models SyllableLM: Learning Coarse Semantic Units for Speech Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:00.432842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:00.432842Z digest=sha256:ef4874f9cd61b6422b486eef8bbe1ffbc095f4f3350be114b682cb1676bf1087

Observation 4367e883-40a3-4b54-b078-4ae4544f7a68 · outbound

This paper cites Sylber: Syllabic Embedding Repre- sentation of Speech from Raw Audio,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Sylber: Syllabic Embedding Repre- sentation of Speech from Raw Audio,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.138108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:00.556666Z digest=sha256:ede7585412d0df5a4097e3ee1d6eb75cd300a32f8f389fe1608af4f4079b4cb3

Observation 4635ecbd-ac80-4160-9225-3d93d1ffc5be · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.894664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:00.664857Z digest=sha256:2ad77e746758518645b3158d0147fb246ed31c322c96566252189a4d223d8db2

Observation 3efbaa60-489f-409b-80d8-cfd65dd5e073 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no super- vision,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Libri-light: A benchmark for ASR with limited or no super- vision,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.615415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:00.802807Z digest=sha256:ca36999c00aab48d347db9e994b3351455a969f7952e4e182da1156f73eace6c

Observation 1e453aac-68e8-44d7-99a5-ce410e69dde7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models OPT: Open Pre-trained Transformer Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:00.957373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:00.957373Z digest=sha256:bffd78108734abeb7cbd65d7ffc1f015d00e713e77f6858ab9ab789bf97e0ded

Observation 7877ab0c-d921-437a-85b6-25eff4f21cbc · outbound

This paper cites ProsAudit, a prosodic benchmark for self-supervised speech models,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models ProsAudit, a prosodic benchmark for self-supervised speech models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.320994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.110052Z digest=sha256:8e0256cc6bbf5309ef64f97acb6aaa9ae5762cfc36257eee121c6b237abcdd7a

Observation 688f3b9f-cfa1-49ac-9e5f-3dba2828f275 · outbound

This paper cites A corpus and cloze evaluation for deeper understanding of commonsense stories,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models A corpus and cloze evaluation for deeper understanding of commonsense stories,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.110815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.237648Z digest=sha256:ca78944fbd3580d7ed289ccc875eea172eb871f505907bef272ffc8912855589

Observation 5021f133-9058-4f65-ae6b-c35cee78d217 · outbound

This paper cites Praat: doing phonetics by com- puter [computer program]. version 6.4.27,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Praat: doing phonetics by com- puter [computer program]. version 6.4.27,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.857528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.363667Z digest=sha256:6bd8a6b3938af663293e8e63e4c2f2e3de6f5fb1381a22f94e302fb528e446f5

Observation fc7b4e91-a9ec-48aa-a83e-67d0b3690784 · outbound

This paper cites Martinet, Elements of General Linguistics, ser.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Martinet, Elements of General Linguistics, ser

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.592334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.513967Z digest=sha256:8d29ddf753a2a2b19eedb81402d010f045c755510ea9f84f8c71d02c396353f5

Observation 3805a6d2-15f9-4b42-9db8-6599ed20e1da · outbound

This paper cites Are Discrete Units Nec- essary for Spoken Language Modeling?.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Are Discrete Units Nec- essary for Spoken Language Modeling?

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.373953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.653076Z digest=sha256:991f91db56091f7a463d904edc2616f90f3354260499bfeaea294d2dc910799d

Observation 05e41d96-2af2-4730-a5b6-ac84e179efd6 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Spirit LM: Interleaved Spoken and Written Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:01.779226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:01.779226Z digest=sha256:ea98b07688ca833a34222ae48e29829ffa9ae399a1014462f9d69dae1221002f

Observation 47015db1-6fa3-4ab6-99af-e34b107b2ff2 · outbound

This paper cites Multi- resolution hubert: Multi-resolution speech self-supervised learn- ing with masked unit prediction,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Multi- resolution hubert: Multi-resolution speech self-supervised learn- ing with masked unit prediction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.120389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.870557Z digest=sha256:ad3d3e439c1f93ea858b5cd47e37fd3c8bf9bf0f562a32a33aba2638b2e228eb

Observation a782d3a6-28f9-42e3-ba2e-e69c11c1d4c7 · outbound

This paper cites Self-supervised contrastive learning for unsupervised phoneme segmentation,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Self-supervised contrastive learning for unsupervised phoneme segmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:02.865867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:01.956551Z digest=sha256:a3cd62f8510018d3cfe4819aba0fe1681198071e8e27626386c0c0d39bada186

Observation abf54504-99ea-4786-be68-a93e630cec9a · outbound

This paper cites Unsupervised word segmentation using temporal gradient pseudo-labels,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unsupervised word segmentation using temporal gradient pseudo-labels,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:02.620676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:50:02.022181Z digest=sha256:29e82826d36df21704d6fdb833d4060b7cbb94caa71fdcb62df1fa99d94c48f2

Pith citing papers

Observation 8cf812e2-ced0-49ae-874a-a2004ad39b8e · inbound

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models cites this paper.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:50:02.409712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:49:57.859145Z digest=sha256:b1141fa18e1a7ab4a64e1a309d4e2f9a807bc6e076609887bf8f467b03783783