Pith. sign in

Paper Citation Record · LEDGER

Spoken Language Modeling with Duration-Penalized Self-Supervised Units

As of 9 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2505.23494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23494 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:24.419823Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:21.863037Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:50:24.579818Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06608138-857a-45c2-92e7-a53bdda93974 · outbound

This paper cites In contrast to models that combine speech and text [2, 3], SLMs do not use text data in any of their components.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units In contrast to models that combine speech and text [2, 3], SLMs do not use text data in any of their components

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:30.109714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:21.791105Z digest=sha256:d3cc9260d945594ebb0cff9a0f04ce0db73f00b6a4afc8e8ed96fc028ac5e417

Observation 640436c9-b19f-4ec4-a88f-559e3845649b · outbound

This paper cites Spoken Language Modeling with Duration-Penalized Self-Supervised Units.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spoken Language Modeling with Duration-Penalized Self-Supervised Units

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:50:24.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:21.863037Z digest=sha256:d19dcfe893b3c0209e6a2507438031c25a57d4742fbff8b104f8ea049ec0ed40

Observation 3b814737-64e9-4771-bfa9-1716d8f4b7b4 · outbound

This paper cites an unresolved cited work.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:29.932442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:21.932155Z digest=sha256:3de6df7699b64daae4c67a54a261dfea02c262f0c967d239760060a6603b6f16

Observation 9ddfc6f1-54c3-4ce5-984a-6b91fadfea7b · outbound

This paper cites Tak- ing the results together, it appears that coarser units are not bene- ficial at the lowest levels in discriminating between isolated units like phones or words.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Tak- ing the results together, it appears that coarser units are not bene- ficial at the lowest levels in discriminating between isolated units like phones or words

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.507973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.106834Z digest=sha256:cd11e2977f134caa14f85b7e8fe6ac1d6ce25c0a56df6aed95fb45b8d62a5ace

Observation 594ca001-9469-49ee-8a06-c06c3a00fe86 · outbound

This paper cites an unresolved cited work.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:29.282842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.203345Z digest=sha256:9f9da46b81c273a4754a5d2db8bee727b00c3bf3a7b29fb9fe6d524219eb338e

Observation 7ef9514b-7589-429e-95e3-b4a64a36e061 · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units AudioLM: a Language Modeling Approach to Audio Generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.086194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.687822Z digest=sha256:45151a11021fdfccc02df111c8e386723cae04c2516d8673a9d79be5b4cb565b

Observation 7db108b4-b79c-4495-8c00-142e86899e3b · outbound

This paper cites The Zero Resource Speech Challenge 2021: Spoken language modelling,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Zero Resource Speech Challenge 2021: Spoken language modelling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.052936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.306799Z digest=sha256:ac4f9e548af4d65e2ff035d206fd8506daa1fa874352476b011b2e6841bb29a6

Observation b188f64b-b08f-45b9-b9ef-3693ed0f1568 · outbound

This paper cites Spirit LM: In- terleaved Spoken and Written Language Model,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spirit LM: In- terleaved Spoken and Written Language Model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.872823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.386496Z digest=sha256:47a664386b5cf29d6b20e780dd141dc6994d6c5d65045d2627703dd3c8581be8

Observation beded2c7-2ab5-4525-b284-217e82c78857 · outbound

This paper cites Textually Pretrained Speech Language Models,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Textually Pretrained Speech Language Models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.668155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.456920Z digest=sha256:bf15f37d409bc6649750930eaafb5a8b22e9163bd2f5ca73b8dc86e51ce5c6bb

Observation 758ef516-dc6d-45e4-b7a4-c25630680688 · outbound

This paper cites Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.424550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.546163Z digest=sha256:ef68ad876c189d072a34445a5d30c7e9354c264539d7c13854db60a86c420d4f

Observation b33f50b6-8350-4658-8bbd-fddc890e4db7 · outbound

This paper cites Generative Spoken Language Modeling From Raw Audio,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Generative Spoken Language Modeling From Raw Audio,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:28.229727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.618879Z digest=sha256:c65d24db80c02489228ca7d699ea2dc2ec4590a8ec5e7094e046cd6a2a7eaa41

Observation f0ae4a99-49fc-45e4-bda1-d03e3914f1a8 · outbound

This paper cites Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Word Segmentation on Discovered Phone Units With Dynamic Programming and Self-Supervised Scoring,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.027150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.158666Z digest=sha256:aa98b71d847dfabef5eacb6a22c29cf52256504ea78795b4c6caed706a1808f2

Observation 47717e75-0284-4bec-86a6-4659c215d4ef · outbound

This paper cites Generative Spoken Language Model based on continuous word-sized audio tokens,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Generative Spoken Language Model based on continuous word-sized audio tokens,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.939433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.764354Z digest=sha256:65d90f90585b6645e9d016993ab052ea407f909740e55fd06dd28a1a46e68e4c

Observation 9b06804b-c2d4-41b3-bd76-678d11e96552 · outbound

This paper cites Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.702573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.836392Z digest=sha256:12b7552be94070d565a05a6042f34d603525237da1b0ab7f08d413b110177d25

Observation 61a5832a-ad4d-4e37-88b5-a91674fd415c · outbound

This paper cites Unsupervised Speech Segmentation and Variable Rate Represen- tation Learning Using Segmental Contrastive Predictive Coding,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Unsupervised Speech Segmentation and Variable Rate Represen- tation Learning Using Segmental Contrastive Predictive Coding,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.475713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.943981Z digest=sha256:c3eeef574f0036e4c4ff60d63f2dd0c2003538976c9c9d4676e14d62e654690a

Observation 7c268e06-e551-404f-9937-ba858723430e · outbound

This paper cites SyllableLM: Learning Coarse Semantic Units for Speech Language Models,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units SyllableLM: Learning Coarse Semantic Units for Speech Language Models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.276551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.014485Z digest=sha256:08e96ebe928e458d2bce520526a64bad514ca63559b55531f9a78f2b2b462898

Observation e7d2421c-bdcd-4b3f-8360-55ca2a4c5a25 · outbound

This paper cites Acoustic BPE for Speech Generation with Discrete tokens,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Acoustic BPE for Speech Generation with Discrete tokens,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:27.149143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.086981Z digest=sha256:991c8d6f3a8bb98f5d636bbe4929568598967a3fef1025a4b712d4ecc932422c

Observation b94e3f04-6fdd-4c3b-8868-ca3affeb4840 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units RoFormer: Enhanced Transformer with Rotary Position Embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.950311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.725495Z digest=sha256:7944a736a96880978d8fce711e01863b8089896475fbb32693df175add069429

Observation 817d80d3-99c0-48ea-b483-836c5d35d3a0 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.792797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.225585Z digest=sha256:a9be37784a341734513957a80ec708de9d4bb56c35529a57a6f43ae76a705819

Observation 8c120ed1-fc07-47aa-9436-a887fba37f76 · outbound

This paper cites Wavlm: Large-Scale Self- Supervised Pre-training for Full Stack Speech Processing,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Wavlm: Large-Scale Self- Supervised Pre-training for Full Stack Speech Processing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.564096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.288383Z digest=sha256:c4f073a923765de0b101cdac7de3db6be0a319790a0062cd7314f677dec961e1

Observation 90e56f44-1b75-499f-8e4f-1c35eadb1e42 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units LibriSpeech: An ASR corpus based on public domain audio books,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.356942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.356942Z digest=sha256:bc981733e536361836fac650d788724c67b0cbf91d41712074e6d4d522292087

Observation 9977a396-35fd-47c9-b20f-134a59c3be16 · outbound

This paper cites The Faiss library,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Faiss library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.377408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.417379Z digest=sha256:9c4bcc78f3063df40d739254be4b775b019b93e1bb16079bcbf9bfb62ca5a242

Observation abf61174-cbd9-4425-9c65-9a5fc5b9cd5a · outbound

This paper cites Evaluating context- invariance in unsupervised speech representations,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Evaluating context- invariance in unsupervised speech representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.112820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:24.208275Z digest=sha256:5a0e6da46967424f850d0f4337acd422165b1958932ce2a43c3eb349068e8cb8

Observation 28368bce-e9c1-45ba-ba6c-c93255fa2da1 · outbound

This paper cites Mistral 7B,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Mistral 7B,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:26.174289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.566053Z digest=sha256:c7148b0cbe362341511814b074337cdbf04675c353ef4586045a84290580b9ce

Observation 19ffa97b-0239-4ee9-bde9-251cccdb3674 · outbound

This paper cites Libri-Light: A Benchmark for ASR with Limited or No Supervi- sion,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Libri-Light: A Benchmark for ASR with Limited or No Supervi- sion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:24.815306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:24.419823Z digest=sha256:187df3b025281d9c800f16a26517ffcfbef31286a0b5648532074fd0be45c2bb

Observation 0658866f-b870-4cb7-95e7-211c87262de0 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.703153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.812663Z digest=sha256:2555dbc93f2883f1d57f32ea9158d50050b445bd9fb2e6b5ff9cec66fdf7d2ae

Observation 16d728c9-29b5-4158-bbbf-c953547cb8ec · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.475685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:23.956477Z digest=sha256:80fd5fc96ed47a3220dfecdb48fbc14ae387d999068b635074503a38b6a51349

Observation d5230d25-7278-4221-b2cd-a8084f5ad431 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.351109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:24.086601Z digest=sha256:6064edcad4ce51060d63165768c28b848705cea1b3c8a487123cfacf7da57589

Observation fba35b79-cc95-451e-bba0-ca94dd66fbb1 · outbound

This paper cites Evaluating Speech Features with the Minimal-Pair ABX Task: Analysis of the Classical MFC/PLP Pipeline,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Evaluating Speech Features with the Minimal-Pair ABX Task: Analysis of the Classical MFC/PLP Pipeline,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:25.226943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:24.146667Z digest=sha256:13348443ae5def8eb1d7a653963ae6c085aa3058df41ac67e18a1dad0daa6b4e

Observation 6b7b2e66-7507-4d55-9319-a3a5086ba823 · outbound

This paper cites Rapid Evaluation of Speech Representations for Spoken Term Discovery,.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Rapid Evaluation of Speech Representations for Spoken Term Discovery,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:24.978118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:24.294980Z digest=sha256:b3b442c5e1371ddea5b8effd70f6b0db2df82dab5d1263af6dc3a4e35dc0aded

Observation 9225c9b3-3786-4681-8f8b-8855d4f6e34c · outbound

This paper cites manu- facture.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units manu- facture

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:29.700297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:22.019652Z digest=sha256:0cc2a75ea67be0fb7a7f6c816ca89f7b0faddd2afa23f71d0487cd9d342ab443

Observation e5ebe4c4-413c-4be0-862b-44f3b691c8cc · outbound

This paper cites Mistral 7B.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Mistral 7B

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.642487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.642487Z digest=sha256:34f45cccfc333c628113e32dc96cae34a90063a8110071d0570d374ecc3f1b07

Observation 0b8a60cf-a6a6-4d63-b018-23c722a78db3 · outbound

This paper cites The Faiss library.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units The Faiss library

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:23.487919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:23.487919Z digest=sha256:ba8f4843ab67970bc005aaa2669b42f0bd6b8b48b2b2b9e1c0f68feb7f8d6c56

Pith citing papers

Observation 640436c9-b19f-4ec4-a88f-559e3845649b · inbound

Spoken Language Modeling with Duration-Penalized Self-Supervised Units cites this paper.

Spoken Language Modeling with Duration-Penalized Self-Supervised Units Spoken Language Modeling with Duration-Penalized Self-Supervised Units

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:50:24.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:50:21.863037Z digest=sha256:d19dcfe893b3c0209e6a2507438031c25a57d4742fbff8b104f8ea049ec0ed40