Pith. sign in

Paper Citation Record · LEDGER

Spoken Language Modeling with Duration-Penalized Self-Supervised Units

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2505.23494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23494 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:24.419823Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:21.863037Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:50:24.579818Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06608138-857a-45c2-92e7-a53bdda93974 · outbound

This paper cites In contrast to models that combine speech and text [2, 3], SLMs do not use text data in any of their components.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units In contrast to models that combine speech and text [2, 3], SLMs do not use text data in any of their components

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.109714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:21.791105Z digest=sha256:31e5b54aa5a0c1410445f5dc39d6b5fdd89a32467a34a69e1f69bc72e45c2677

Observation 640436c9-b19f-4ec4-a88f-559e3845649b · outbound

This paper cites Spoken Language Modeling with Duration-Penalized Self-Supervised Units.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spoken Language Modeling with Duration-Penalized Self-Supervised Units

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:50:24.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:21.863037Z digest=sha256:c32be1d8a070025b121f97aeb8d5b8ee2aa506ee5d9ae45903563c5133ec453f

Observation 3b814737-64e9-4771-bfa9-1716d8f4b7b4 · outbound

This paper cites an unresolved cited work.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:29.932442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:21.932155Z digest=sha256:d727bcc03e56eaaf15bce665bcecbc12d863576feea13d608d1dac80a9aba2a8

Observation 9ddfc6f1-54c3-4ce5-984a-6b91fadfea7b · outbound

This paper cites Tak- ing the results together, it appears that coarser units are not bene- ficial at the lowest levels in discriminating between isolated units like phones or words.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Tak- ing the results together, it appears that coarser units are not bene- ficial at the lowest levels in discriminating between isolated units like phones or words

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.507973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.106834Z digest=sha256:e159dabeec71f271ae8cb18d120099ee4dacfaa46241f8993c695d33e1fffe73

Observation 594ca001-9469-49ee-8a06-c06c3a00fe86 · outbound

This paper cites an unresolved cited work.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:29.282842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.203345Z digest=sha256:944a0abb9e113ffde43bc53162e81b694c30c9ca8345d80ed29cc7306a449c2d

Observation 7ef9514b-7589-429e-95e3-b4a64a36e061 · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units AudioLM: a Language Modeling Approach to Audio Generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.086194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.687822Z digest=sha256:fbf84b037e39ad8699fbce338727b7c06752546e59eda8e916b3ac760beadef9

Observation 7db108b4-b79c-4495-8c00-142e86899e3b · outbound

This paper cites The Zero Resource Speech Challenge 2021: Spoken language modelling,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Zero Resource Speech Challenge 2021: Spoken language modelling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.052936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.306799Z digest=sha256:40d2d89bb39d096d420d41f25ebef5f41ea9734080457b99b91b9089ce61afa2

Observation b188f64b-b08f-45b9-b9ef-3693ed0f1568 · outbound

This paper cites Spirit LM: In- terleaved Spoken and Written Language Model,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spirit LM: In- terleaved Spoken and Written Language Model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.872823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.386496Z digest=sha256:340e8accd680b5e62db8823ddf2ae1d19a52d14edf6f329d2651abbd144900a6

Observation beded2c7-2ab5-4525-b284-217e82c78857 · outbound

This paper cites Textually Pretrained Speech Language Models,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Textually Pretrained Speech Language Models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.668155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.456920Z digest=sha256:6c00164e0ff6fc6cac82cc592275ed9a7a691331ee4f2eeba40d2b375e6e019c

Observation 758ef516-dc6d-45e4-b7a4-c25630680688 · outbound

This paper cites Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.424550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.546163Z digest=sha256:713e63dea8434ef1964daf00e18662d9037a094fbacf5d3ac398532a3a8f2348

Observation b33f50b6-8350-4658-8bbd-fddc890e4db7 · outbound

This paper cites Generative Spoken Language Modeling From Raw Audio,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Generative Spoken Language Modeling From Raw Audio,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.229727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.618879Z digest=sha256:1c7fc67f8db5a3e732286b892af4443f3f3315018f3fd106493c42c809af12b5

Observation f0ae4a99-49fc-45e4-bda1-d03e3914f1a8 · outbound

This paper cites Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.027150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.158666Z digest=sha256:478ba45f818486f1e1f35748a84353a617992015894b3507bd1f46b1174b52d4

Observation 47717e75-0284-4bec-86a6-4659c215d4ef · outbound

This paper cites Generative Spoken Language Model based on continuous word-sized audio tokens,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Generative Spoken Language Model based on continuous word-sized audio tokens,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.939433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.764354Z digest=sha256:c00f45e5d0810319fbf173a3773dbceb42f85f9756098159ac426db207b15ac0

Observation 9b06804b-c2d4-41b3-bd76-678d11e96552 · outbound

This paper cites Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.702573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.836392Z digest=sha256:a9d4176533683a69d93cf017055ad7b1e043650f57f8c5f83d175602bf5e518b

Observation 61a5832a-ad4d-4e37-88b5-a91674fd415c · outbound

This paper cites Unsupervised Speech Segmentation and Variable Rate Represen- tation Learning Using Segmental Contrastive Predictive Coding,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unsupervised Speech Segmentation and Variable Rate Represen- tation Learning Using Segmental Contrastive Predictive Coding,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.475713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.943981Z digest=sha256:c2380afaa3b0d86c25887621962039b379e639fcac1b84dcc7dc497428901d07

Observation 7c268e06-e551-404f-9937-ba858723430e · outbound

This paper cites SyllableLM: Learning Coarse Semantic Units for Speech Language Models,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units SyllableLM: Learning Coarse Semantic Units for Speech Language Models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.276551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.014485Z digest=sha256:0e9b26af95bd3d73861c4761725847eecc991fce64c5dfe8069b6fd705f1d93d

Observation e7d2421c-bdcd-4b3f-8360-55ca2a4c5a25 · outbound

This paper cites Acoustic BPE for Speech Generation with Discrete tokens,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Acoustic BPE for Speech Generation with Discrete tokens,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.149143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.086981Z digest=sha256:513d5c8bef12de6def1f23d477cd7250f5fce1b09e0d9faadc66344248c80693

Observation b94e3f04-6fdd-4c3b-8868-ca3affeb4840 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units RoFormer: Enhanced Transformer with Rotary Position Embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.950311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.725495Z digest=sha256:a5cf0deaebcefee7a55dd99a5e6fc144815b02fd0148f1e72dc8709a12c73d5f

Observation 817d80d3-99c0-48ea-b483-836c5d35d3a0 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.792797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.225585Z digest=sha256:4e6355259b7706968d69bc70ae2a0604885e3c803d7392eb0ead00e73ffae38b

Observation 8c120ed1-fc07-47aa-9436-a887fba37f76 · outbound

This paper cites Wavlm: Large-Scale Self- Supervised Pre-training for Full Stack Speech Processing,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Wavlm: Large-Scale Self- Supervised Pre-training for Full Stack Speech Processing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.564096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.288383Z digest=sha256:493fcab428ad70d45a529abe622e0c49eb1e9b6c33aa621592f0f38c5fd1836e

Observation 90e56f44-1b75-499f-8e4f-1c35eadb1e42 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units LibriSpeech: An ASR corpus based on public domain audio books,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.356942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.356942Z digest=sha256:05142be1b669c68346d45426d2259b8971a750f4ac05381814dce8ae8151e67c

Observation 9977a396-35fd-47c9-b20f-134a59c3be16 · outbound

This paper cites The Faiss library,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Faiss library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.377408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.417379Z digest=sha256:3fb61e7eb76e1beb1c2368a5bdf666332d1b1c9fe782c9ae699fdb45c0dc1fe1

Observation abf61174-cbd9-4425-9c65-9a5fc5b9cd5a · outbound

This paper cites Evaluating context- invariance in unsupervised speech representations,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Evaluating context- invariance in unsupervised speech representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.112820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:24.208275Z digest=sha256:606549419520cbb227acb31f453944b0f8a55ff4b342656c6d2d07c4956dea54

Observation 28368bce-e9c1-45ba-ba6c-c93255fa2da1 · outbound

This paper cites Mistral 7B,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Mistral 7B,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.174289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.566053Z digest=sha256:dc185ac3c44b9c577c6a9499a83e83c4dd2bd6a94b28263d07b1b74f7f0202b0

Observation 19ffa97b-0239-4ee9-bde9-251cccdb3674 · outbound

This paper cites Libri-Light: A Benchmark for ASR with Limited or No Supervi- sion,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Libri-Light: A Benchmark for ASR with Limited or No Supervi- sion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:24.815306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:24.419823Z digest=sha256:07bed4bbd76a338100212692bf88197bdfc72136b20972eda790fd5f17568a12

Observation 0658866f-b870-4cb7-95e7-211c87262de0 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.703153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.812663Z digest=sha256:ea4ea542d5c8fbf41fa1fc45962c629d1523e7fa8e598cbcc615e83d6eaeaf78

Observation 16d728c9-29b5-4158-bbbf-c953547cb8ec · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.475685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:23.956477Z digest=sha256:221cd9a14a5f3201d304a2ab0b7d88f76a2a8cb4c133b18fb1b6e8aec8290490

Observation d5230d25-7278-4221-b2cd-a8084f5ad431 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.351109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:24.086601Z digest=sha256:5ecf0e59a8767638a2455f3d8c3befe1c53613b59670dde33e6642c7ea471a83

Observation fba35b79-cc95-451e-bba0-ca94dd66fbb1 · outbound

This paper cites Evaluating Speech Features with the Minimal-Pair ABX Task: Analysis of the Classical MFC/PLP Pipeline,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Evaluating Speech Features with the Minimal-Pair ABX Task: Analysis of the Classical MFC/PLP Pipeline,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.226943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:24.146667Z digest=sha256:c0fd9473b35a557e97db48630a0598511c7de309f5be2ff68969b7063460595b

Observation 6b7b2e66-7507-4d55-9319-a3a5086ba823 · outbound

This paper cites Rapid Evaluation of Speech Representations for Spoken Term Discovery,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Rapid Evaluation of Speech Representations for Spoken Term Discovery,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:24.978118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:24.294980Z digest=sha256:e922591f6183ab7bc23160a5a9c45211a180f306c83d0391ad48f1bb77e0a341

Observation 9225c9b3-3786-4681-8f8b-8855d4f6e34c · outbound

This paper cites manu- facture.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units manu- facture

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.700297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:22.019652Z digest=sha256:113901f2805cd5a3f9c887c5470199c80cc8d4cde64a7fc8b4328956079730c9

Observation e5ebe4c4-413c-4be0-862b-44f3b691c8cc · outbound

This paper cites Mistral 7B.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Mistral 7B

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.642487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.642487Z digest=sha256:14550d640238c9937c37fc6ff4b2f74a2524ac2ecdff3cf34448c52b1617122d

Observation 0b8a60cf-a6a6-4d63-b018-23c722a78db3 · outbound

This paper cites The Faiss library.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Faiss library

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.487919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.487919Z digest=sha256:65d38af503cbcfc455f1efb2f6a33e24327248e7ca421d7f8b30bc85f09baab2

Pith citing papers

Observation 640436c9-b19f-4ec4-a88f-559e3845649b · inbound

Spoken Language Modeling with Duration-Penalized Self-Supervised Units cites this paper.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spoken Language Modeling with Duration-Penalized Self-Supervised Units

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:50:24.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:50:21.863037Z digest=sha256:c32be1d8a070025b121f97aeb8d5b8ee2aa506ee5d9ae45903563c5133ec453f