Pith. sign in

Paper Citation Record · LEDGER

Spoken Language Modeling with Duration-Penalized Self-Supervised Units

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2505.23494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23494 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:24.419823Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:21.863037Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:50:24.579818Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06608138-857a-45c2-92e7-a53bdda93974 · outbound

This paper cites In contrast to models that combine speech and text [2, 3], SLMs do not use text data in any of their components.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units In contrast to models that combine speech and text [2, 3], SLMs do not use text data in any of their components

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.109714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:21.791105Z digest=sha256:0ff75d776340b424a6eb8c7fc2daeafffc8e7d393cd20a3355a810bedc4a9dc1

Observation 640436c9-b19f-4ec4-a88f-559e3845649b · outbound

This paper cites Spoken Language Modeling with Duration-Penalized Self-Supervised Units.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spoken Language Modeling with Duration-Penalized Self-Supervised Units

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:50:24.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:21.863037Z digest=sha256:25134230bd7dcf1bad63a8220f5fe08c451b0dc919e64ba35fde906f00f6b13f

Observation 3b814737-64e9-4771-bfa9-1716d8f4b7b4 · outbound

This paper cites an unresolved cited work.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:29.932442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:21.932155Z digest=sha256:0ae373b16b8c026933c9ef89b69580facdd74f21ee5e1448196cfb2001aa66a7

Observation 9ddfc6f1-54c3-4ce5-984a-6b91fadfea7b · outbound

This paper cites Tak- ing the results together, it appears that coarser units are not bene- ficial at the lowest levels in discriminating between isolated units like phones or words.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Tak- ing the results together, it appears that coarser units are not bene- ficial at the lowest levels in discriminating between isolated units like phones or words

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.507973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.106834Z digest=sha256:9ab853a111c94a9fd58986e81b205f2d891f4c48610a4fba466b5e70c17d078d

Observation 594ca001-9469-49ee-8a06-c06c3a00fe86 · outbound

This paper cites an unresolved cited work.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:29.282842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.203345Z digest=sha256:7f82cfc064acd7fbe2c12477d2a961d54dddc99c05bf425759d469c348a579c1

Observation 7ef9514b-7589-429e-95e3-b4a64a36e061 · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units AudioLM: a Language Modeling Approach to Audio Generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.086194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.687822Z digest=sha256:31b66687d17c34a82e18ebcdfd95576e99bacc228de95e92be1f394a3a87f2fd

Observation 7db108b4-b79c-4495-8c00-142e86899e3b · outbound

This paper cites The Zero Resource Speech Challenge 2021: Spoken language modelling,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Zero Resource Speech Challenge 2021: Spoken language modelling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.052936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.306799Z digest=sha256:a53c595db3e7665bb7fe2905011ac59fe4eff56ebefb4090a451e1a68bf528e5

Observation b188f64b-b08f-45b9-b9ef-3693ed0f1568 · outbound

This paper cites Spirit LM: In- terleaved Spoken and Written Language Model,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spirit LM: In- terleaved Spoken and Written Language Model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.872823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.386496Z digest=sha256:bdfb375a97d5398d0177a1e88023f35087ec72bd1f2738a4ed39f291813447f0

Observation beded2c7-2ab5-4525-b284-217e82c78857 · outbound

This paper cites Textually Pretrained Speech Language Models,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Textually Pretrained Speech Language Models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.668155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.456920Z digest=sha256:d4bb9b6699519cc0b270218335225f7d861b377e0d73cd0f2feeab88ad75254b

Observation 758ef516-dc6d-45e4-b7a4-c25630680688 · outbound

This paper cites Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.424550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.546163Z digest=sha256:1dfece9f95e8a94114136fea068b74eb15205459074250c979527ed049afbdd4

Observation b33f50b6-8350-4658-8bbd-fddc890e4db7 · outbound

This paper cites Generative Spoken Language Modeling From Raw Audio,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Generative Spoken Language Modeling From Raw Audio,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.229727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.618879Z digest=sha256:642e65cf3872d9321bc5c86a4ae7218a72805b0bfafd87741621184bb164e6c7

Observation f0ae4a99-49fc-45e4-bda1-d03e3914f1a8 · outbound

This paper cites Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.027150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.158666Z digest=sha256:cfc25f98f7dcd9acb9793f2557e1382b2f4dabb2af85008c5b28b5b9e7b75dc3

Observation 47717e75-0284-4bec-86a6-4659c215d4ef · outbound

This paper cites Generative Spoken Language Model based on continuous word-sized audio tokens,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Generative Spoken Language Model based on continuous word-sized audio tokens,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.939433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.764354Z digest=sha256:55266338293b93652a60e85f0121a8780665aa727fff0e63084656aa236a9217

Observation 9b06804b-c2d4-41b3-bd76-678d11e96552 · outbound

This paper cites Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.702573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.836392Z digest=sha256:f8d6a0fd1e6db705d536b690e08b4649248b43d61df9a7295555b9bd42fd1554

Observation 61a5832a-ad4d-4e37-88b5-a91674fd415c · outbound

This paper cites Unsupervised Speech Segmentation and Variable Rate Represen- tation Learning Using Segmental Contrastive Predictive Coding,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unsupervised Speech Segmentation and Variable Rate Represen- tation Learning Using Segmental Contrastive Predictive Coding,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.475713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.943981Z digest=sha256:06965b29ba3b7f4d4f3e7e8f79b515f65d00c39ea272a461afbb68490c5a51c4

Observation 7c268e06-e551-404f-9937-ba858723430e · outbound

This paper cites SyllableLM: Learning Coarse Semantic Units for Speech Language Models,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units SyllableLM: Learning Coarse Semantic Units for Speech Language Models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.276551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.014485Z digest=sha256:086c7b6f053e6a8c607d2d494b4a8aac0bc5c235291f13be4ea2a14907813945

Observation e7d2421c-bdcd-4b3f-8360-55ca2a4c5a25 · outbound

This paper cites Acoustic BPE for Speech Generation with Discrete tokens,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Acoustic BPE for Speech Generation with Discrete tokens,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.149143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.086981Z digest=sha256:0afcda9e140a3934a7add55b1c1dc8bfdceeb1cd79a764069c8877d95b3f0d2b

Observation b94e3f04-6fdd-4c3b-8868-ca3affeb4840 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units RoFormer: Enhanced Transformer with Rotary Position Embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.950311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.725495Z digest=sha256:330a13c123104c62e1846490ece74aa233cc9da0c21826dc15eabc752ced4796

Observation 817d80d3-99c0-48ea-b483-836c5d35d3a0 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.792797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.225585Z digest=sha256:19268c8d769a6e47129639678efa14094c8217329566890c931df8b44baafd63

Observation 8c120ed1-fc07-47aa-9436-a887fba37f76 · outbound

This paper cites Wavlm: Large-Scale Self- Supervised Pre-training for Full Stack Speech Processing,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Wavlm: Large-Scale Self- Supervised Pre-training for Full Stack Speech Processing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.564096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.288383Z digest=sha256:32481abd637910d5d075ab4f6dc9bf7e71354b0af4b56fbeec3e5a08d1d2699d

Observation 90e56f44-1b75-499f-8e4f-1c35eadb1e42 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units LibriSpeech: An ASR corpus based on public domain audio books,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.356942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.356942Z digest=sha256:05142be1b669c68346d45426d2259b8971a750f4ac05381814dce8ae8151e67c

Observation 9977a396-35fd-47c9-b20f-134a59c3be16 · outbound

This paper cites The Faiss library,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Faiss library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.377408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.417379Z digest=sha256:7b472c17ccf0ed37fc8519f29a3d6e60d918a8ecd64ef6a671a9205c423900c5

Observation abf61174-cbd9-4425-9c65-9a5fc5b9cd5a · outbound

This paper cites Evaluating context- invariance in unsupervised speech representations,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Evaluating context- invariance in unsupervised speech representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.112820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:24.208275Z digest=sha256:81a1786d66a447ae6aad2dc0d6d69036fbeeb890bfdd9d50b3070c2eb232ce57

Observation 28368bce-e9c1-45ba-ba6c-c93255fa2da1 · outbound

This paper cites Mistral 7B,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Mistral 7B,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.174289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.566053Z digest=sha256:bbda076d26f15f53a2c70ff123b95eeabfa83c50030dfb527b0ce7ec3609f4b1

Observation 19ffa97b-0239-4ee9-bde9-251cccdb3674 · outbound

This paper cites Libri-Light: A Benchmark for ASR with Limited or No Supervi- sion,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Libri-Light: A Benchmark for ASR with Limited or No Supervi- sion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:24.815306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:24.419823Z digest=sha256:af84218b0dc855cb7801a3c289b2cde6eb185f8dfd345e7aa21dbc3ec053b1c6

Observation 0658866f-b870-4cb7-95e7-211c87262de0 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.703153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.812663Z digest=sha256:5edcb0bf2adb4fb5855e88ba55bdfc19cdbfe8a495d4d32db02b3909a2e799c2

Observation 16d728c9-29b5-4158-bbbf-c953547cb8ec · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.475685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:23.956477Z digest=sha256:c64a326ab782d70829c7f3aff30e1d248fe907c1f5d132e2d714d5812587c1fd

Observation d5230d25-7278-4221-b2cd-a8084f5ad431 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.351109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:24.086601Z digest=sha256:67f5caaf803aa57ca5ba5d42979346ca62eda56614018508299af4dae0c18397

Observation fba35b79-cc95-451e-bba0-ca94dd66fbb1 · outbound

This paper cites Evaluating Speech Features with the Minimal-Pair ABX Task: Analysis of the Classical MFC/PLP Pipeline,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Evaluating Speech Features with the Minimal-Pair ABX Task: Analysis of the Classical MFC/PLP Pipeline,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.226943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:24.146667Z digest=sha256:99769f9690f19af91d5ddec51230e0c56df782851b782aa93806649eea65963c

Observation 6b7b2e66-7507-4d55-9319-a3a5086ba823 · outbound

This paper cites Rapid Evaluation of Speech Representations for Spoken Term Discovery,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Rapid Evaluation of Speech Representations for Spoken Term Discovery,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:24.978118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:24.294980Z digest=sha256:6536c837b68b929500a240c681163fc024583fce418469df6e7b6d246d39400c

Observation 9225c9b3-3786-4681-8f8b-8855d4f6e34c · outbound

This paper cites manu- facture.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units manu- facture

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.700297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:22.019652Z digest=sha256:863f65def71e7c2f7c4a07dca9afb2d74ba44fd20abf7f9595d60da2cba2b530

Observation e5ebe4c4-413c-4be0-862b-44f3b691c8cc · outbound

This paper cites Mistral 7B.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Mistral 7B

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.642487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.642487Z digest=sha256:14550d640238c9937c37fc6ff4b2f74a2524ac2ecdff3cf34448c52b1617122d

Observation 0b8a60cf-a6a6-4d63-b018-23c722a78db3 · outbound

This paper cites The Faiss library.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Faiss library

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.487919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.487919Z digest=sha256:65d38af503cbcfc455f1efb2f6a33e24327248e7ca421d7f8b30bc85f09baab2

Pith citing papers

Observation 640436c9-b19f-4ec4-a88f-559e3845649b · inbound

Spoken Language Modeling with Duration-Penalized Self-Supervised Units cites this paper.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spoken Language Modeling with Duration-Penalized Self-Supervised Units

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:50:24.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:50:21.863037Z digest=sha256:25134230bd7dcf1bad63a8220f5fe08c451b0dc919e64ba35fde906f00f6b13f