Pith. sign in

Paper Citation Record · LEDGER

On-device Streaming Discrete Speech Units

As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.01845.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01845 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:35:38.605780Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:35:33.841116Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T22:21:53.596123Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3040d7-d133-4d99-99b7-a3186741704a · outbound

This paper cites On-device Streaming Discrete Speech Units.

On-device Streaming Discrete Speech Units On-device Streaming Discrete Speech Units

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:33.841116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:33.841116Z digest=sha256:a95abd79f1fffc429c27ed13ebff8656669500259d10eb0bcce80263dfb0c3af

Observation fd7be342-39df-4441-b9a5-af3cb23e0db8 · outbound

This paper cites To evaluate the predicted DSUs, we focus on the discrete ASR system.

On-device Streaming Discrete Speech Units To evaluate the predicted DSUs, we focus on the discrete ASR system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:46.515002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:33.878111Z digest=sha256:3f750e98b823a87962cb8ec92243b6df9912ff36769790be799fac72c3c972a3

Observation b7b9d3b1-8a94-428f-8a22-75b492385509 · outbound

This paper cites However, this issue can be mitigated by limiting the future window size of the S2U module.

On-device Streaming Discrete Speech Units However, this issue can be mitigated by limiting the future window size of the S2U module

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:46.327654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:33.930393Z digest=sha256:a0d33e1a6076f8725018130e7d3713460d3910c8df653febdbfc86e5f0a02343

Observation b195da30-838c-4d25-b2e7-ccdc6181ee2c · outbound

This paper cites Since the number of layers has a linear relationship with computational cost, reduc- ing them benefits resource-constrained on-device applications.

On-device Streaming Discrete Speech Units Since the number of layers has a linear relationship with computational cost, reduc- ing them benefits resource-constrained on-device applications

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:46.174810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:33.987559Z digest=sha256:143acc790f61c0d1ebce7c3957e2eb39ed8c1acb71934f8b0e1b40c26e8aa7ce

Observation 9ee2a7ed-9fc7-4c27-98c5-c9e0f55ab937 · outbound

This paper cites Also, by applying such, we produce a Pareto optimal curve that represents the trade-off between the computational over- head and the downstream performance.

On-device Streaming Discrete Speech Units Also, by applying such, we produce a Pareto optimal curve that represents the trade-off between the computational over- head and the downstream performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:46.016337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.070482Z digest=sha256:9376d27b8ce42c47745c677962509ccc637ee97deec0217ce27b38179c13ef11

Observation ca88d02a-26d2-4804-a677-28f69b6e1438 · outbound

This paper cites Knowledge distillation (KD) is of- ten used to reduce the size of S3Ms, such as DistilHuBERT.

On-device Streaming Discrete Speech Units Knowledge distillation (KD) is of- ten used to reduce the size of S3Ms, such as DistilHuBERT

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:45.865362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.124317Z digest=sha256:43904a6d7538ec95be5fd793d8271e05c034650f7924b189ce392b51eba2e63c

Observation 6a63d6b9-9db6-4ab7-bc41-edc36e66e719 · outbound

This paper cites How- ever, current methods for generating DSUs rely on full speech input and computationally heavy S3Ms.

On-device Streaming Discrete Speech Units How- ever, current methods for generating DSUs rely on full speech input and computationally heavy S3Ms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:45.550467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.288721Z digest=sha256:a139ce92c608d24a69284fabffcf7f15033d4cc0ccf7a8d0bac7f4a2b7db6d2e

Observation 1b6685b7-021a-4d33-ad1e-93762220c499 · outbound

This paper cites an unresolved cited work.

On-device Streaming Discrete Speech Units Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:35:45.371764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.357562Z digest=sha256:1e867f02afebd99153d5ce494f35fa4e5730ece2b360eda9340b93d00d019db5

Observation 883af9d9-3cfa-4aed-a84b-952631c2eb17 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

On-device Streaming Discrete Speech Units AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:35.097715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:35.097715Z digest=sha256:205345d57e6789367ed4200759aab88fdd04059c29324cd9c109559bd4d8fe8d

Observation bb8e25f6-e0ee-4362-a085-1085116cd7e9 · outbound

This paper cites Self-supervised speech representation learning: A review,.

On-device Streaming Discrete Speech Units Self-supervised speech representation learning: A review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:45.203174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.418993Z digest=sha256:19222c4fbeb1535ebdca50bfd12c89ce583937dda7bb6538bd2de3b5b5c60d09

Observation d0349d48-e10b-46b9-b2b3-95aa9672037f · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

On-device Streaming Discrete Speech Units wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:45.014093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.480352Z digest=sha256:3d77b13a979d170c884d68173fc1ca64fcc2a0537bdcfe4f77488daaeeb6a81a

Observation b1a38999-104f-4bd6-b490-f4c4186463eb · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

On-device Streaming Discrete Speech Units WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:44.823996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.541577Z digest=sha256:d6e00a13ed069b5598ed7fc1fe9311ca541d81b9bc23abca08149d91f6f09ddc

Observation a2fefbab-f650-4fac-b5c1-b5778469abb2 · outbound

This paper cites Exploration of efficient end- to-end asr using discretized input from self-supervised learning,.

On-device Streaming Discrete Speech Units Exploration of efficient end- to-end asr using discretized input from self-supervised learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:44.626961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.631478Z digest=sha256:433f6b8f82f5d4956d0aabdf1fa33c2b8ccb733572e455d45a7737e064cbad4d

Observation 8addd65c-ff24-4e71-a487-140226bb057b · outbound

This paper cites Exploring speech recognition, translation, and understanding with discrete speech units: A com- parative study,.

On-device Streaming Discrete Speech Units Exploring speech recognition, translation, and understanding with discrete speech units: A com- parative study,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:44.407510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.698008Z digest=sha256:0661d326c691799d049d888aab2c822f41e9a86683f2c310bfddc9ec4402391a

Observation 2659ddbf-a1be-4096-9366-39b9c74fa523 · outbound

This paper cites The Interspeech 2024 Challenge on Speech Processing Using Discrete Units,.

On-device Streaming Discrete Speech Units The Interspeech 2024 Challenge on Speech Processing Using Discrete Units,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:44.203084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.817297Z digest=sha256:75451561c4c0f91aa44fa820652c7a0df1d7037f6fa69c4ab533f88a358424db

Observation 4a07b900-ead9-493f-b5b7-aeb5743b252e · outbound

This paper cites AudioLM: a language modeling approach to audio generation,.

On-device Streaming Discrete Speech Units AudioLM: a language modeling approach to audio generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:44.001591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.921663Z digest=sha256:739d00eb760518083ae9d781104409109ec092268fb3098b4cc577d9b3e9fb17

Observation e1038a45-d8ee-476a-8752-ae87d5c6a854 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abili- ties,.

On-device Streaming Discrete Speech Units SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abili- ties,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:43.843634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.994945Z digest=sha256:107e38b5244faca785fbe9e6f9643a5f162f76a654dbdec2103905d39c2678a3

Observation 916bc0c0-d73e-4d5f-b63d-516cf1bcf51e · outbound

This paper cites PARP: Prune, adjust and re-prune for self-supervised speech recognition,.

On-device Streaming Discrete Speech Units PARP: Prune, adjust and re-prune for self-supervised speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:42.107692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.170734Z digest=sha256:8ddd30a8e9ceff9b6cad8a1985a8dcb1ed2b0a6d56419f12e8688ed45a0312a5

Observation 0d9a634d-c28b-4afb-b631-224e242920ca · outbound

This paper cites Discrete speech unit ex- traction via independent component analysis,.

On-device Streaming Discrete Speech Units Discrete speech unit ex- traction via independent component analysis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:43.662951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.182471Z digest=sha256:d5810b9867dd94ac93c9355b0619179bcf0a6c808b5ded9e839c53fe252ced19

Observation d30ae050-9d65-48c5-a60a-99dd72168338 · outbound

This paper cites AnyGPT: Unified multimodal LLM with discrete sequence modeling,.

On-device Streaming Discrete Speech Units AnyGPT: Unified multimodal LLM with discrete sequence modeling,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:43.474147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.264669Z digest=sha256:49201a1954f48af7efcc5cc7035956d675b9abd7b64ee15db4246f2ec758346e

Observation 919ad50c-538a-4515-8b99-dd899e645a30 · outbound

This paper cites Anonymizing dysarthric speech: Investigating the effects of voice conversion on pathological information preservation,.

On-device Streaming Discrete Speech Units Anonymizing dysarthric speech: Investigating the effects of voice conversion on pathological information preservation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:43.258749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.358550Z digest=sha256:3d16506067564f8e4b29d3fe215ac75ac868cf62cf5f2e66c493eb536dd233ba

Observation abc90dcf-0751-421d-90f5-3809f408c2a0 · outbound

This paper cites Self-supervised speech representations are more phonetic than semantic,.

On-device Streaming Discrete Speech Units Self-supervised speech representations are more phonetic than semantic,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:43.039086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.465539Z digest=sha256:0e554aa85ed455e816f8d2a107a6e7cb228be05d84d45212fbe4707ef8012537

Observation cc2f1b35-f9d0-476f-8d31-9ac1e0d3da60 · outbound

This paper cites Comparative layer-wise analy- sis of self-supervised speech models,.

On-device Streaming Discrete Speech Units Comparative layer-wise analy- sis of self-supervised speech models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:42.819444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.579769Z digest=sha256:321b96cb08cb43128b68887b195af3166b327beb9ee5c1526aa936c7733fe5d9

Observation 85d66725-ab20-470e-abfb-8c5fc79aa576 · outbound

This paper cites Leveraging Allophony in Self- Supervised Speech Models for Atypical Pronunciation Assess- ment,.

On-device Streaming Discrete Speech Units Leveraging Allophony in Self- Supervised Speech Models for Atypical Pronunciation Assess- ment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:42.684788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.730501Z digest=sha256:25fe4173555fefce250871bf171f52c252f09663838a0ea1375857f3da77b284

Observation 59c1d868-181c-4fe9-88c8-cba105d943c4 · outbound

This paper cites Understanding probe be- haviors through variational bounds of mutual information,.

On-device Streaming Discrete Speech Units Understanding probe be- haviors through variational bounds of mutual information,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:42.498374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:35.897040Z digest=sha256:cfb63201861e79e662196fecb9a5c6848f080dd941035286b71e2e8bbef7aa7e

Observation 93cfc809-ae91-446a-be94-c1dee10b9a23 · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,.

On-device Streaming Discrete Speech Units HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:42.315454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.021547Z digest=sha256:dd36c5e60dc0ee53a3f68313d6c1c02b82b3a9de5235c332db4f120ce5cb3eb7

Observation cad73ed6-e25f-47ee-a5fa-70c4061fbdc0 · outbound

This paper cites Streaming automatic speech recog- nition with the transformer model,.

On-device Streaming Discrete Speech Units Streaming automatic speech recog- nition with the transformer model,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:40.673957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.351221Z digest=sha256:2acb9806dacb7e139005c9f07c18eecefae1f4f63cc0cebdeb49cba0246d6756

Observation c597c76b-7e9a-4baa-819f-0f44d0b3d551 · outbound

This paper cites Structured pruning of self- supervised pre-trained models for speech recognition and under- standing,.

On-device Streaming Discrete Speech Units Structured pruning of self- supervised pre-trained models for speech recognition and under- standing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:41.942362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.341680Z digest=sha256:c1b0ff7797994f6627ae77c2d40275e801a9b9c9a612303f19db03b57f5ce182

Observation 1c848c50-5e5d-4087-b73b-1cd5c6ed27eb · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

On-device Streaming Discrete Speech Units Lib- rispeech: an asr corpus based on public domain audio books,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:41.728612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.515042Z digest=sha256:991d04b81656aba0246be6744be3dac7d78ec74a175375b14598eade8a17f400

Observation f2c8a963-77d6-4eea-a8c2-d15d2b146235 · outbound

This paper cites Techniques like pruning [18, 19] and quantization [32] is also used.

On-device Streaming Discrete Speech Units Techniques like pruning [18, 19] and quantization [32] is also used

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:45.714856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:34.203012Z digest=sha256:132f77d203bd383fc22c1f70c96a8ec053a5b9395520bd0a4bb7fbe8f114ae5c

Observation 6836055f-0aac-4321-89eb-00f044bc6033 · outbound

This paper cites ML-SUPERB: Multilin- gual Speech Universal PERformance Benchmark,.

On-device Streaming Discrete Speech Units ML-SUPERB: Multilin- gual Speech Universal PERformance Benchmark,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:41.521738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.633996Z digest=sha256:61bd8b3d4005813d4f859e82946bc613d11301a33db05f8cb952f4e521cb6794

Observation df10f0e3-a7f8-4700-bfb4-17008181fa40 · outbound

This paper cites calflops: a FLOPs and params calculate tool for neu- ral networks,.

On-device Streaming Discrete Speech Units calflops: a FLOPs and params calculate tool for neu- ral networks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:41.353450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.761029Z digest=sha256:7c02d1d6efde0fa5d09af66bf21956f86a83c344444eec31e6a40dbe8128d6ca

Observation d5b472c5-242c-40d1-a0e6-e305f3aa9582 · outbound

This paper cites E-branchformer: Branchformer with enhanced merging for speech recognition,.

On-device Streaming Discrete Speech Units E-branchformer: Branchformer with enhanced merging for speech recognition,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:41.169934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.875636Z digest=sha256:6a47e0b2901ace98d0b8e79748002c8f6c5226739b3cb3e08d77e18f02b93026

Observation 81c4a3de-5ea5-46e4-90bc-14a2bacc25d9 · outbound

This paper cites Attention is all you need,.

On-device Streaming Discrete Speech Units Attention is all you need,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:41.018802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:36.991123Z digest=sha256:d732813a7575350cd26670a1a8bcb347bf0d32df3c218feccddba6e1ef2a48d6

Observation e2e3b421-4941-4942-99f8-6c13fe1b1fba · outbound

This paper cites Decoupled weight decay regulariza- tion,.

On-device Streaming Discrete Speech Units Decoupled weight decay regulariza- tion,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:37.116265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:37.116265Z digest=sha256:f29d6a5ef2f398e9428e8b01abe638ab7a4d1e63b75f1effdc4517f9750b1102

Observation 4f22b8d0-8eaf-49b9-a87a-e16a0c9da93e · outbound

This paper cites A time-restricted self- attention layer for asr,.

On-device Streaming Discrete Speech Units A time-restricted self- attention layer for asr,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:40.841618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.234692Z digest=sha256:d4d92655d8ba9e5a5912ece619d9dec76c0495f326358c898f37245f0ec43e0d

Observation 5de08d89-b382-4417-b5a3-8275f762d20c · outbound

This paper cites SUPERB: Speech Processing Universal PERformance Benchmark,.

On-device Streaming Discrete Speech Units SUPERB: Speech Processing Universal PERformance Benchmark,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:40.519694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.446796Z digest=sha256:7a494b84087f05e44988497f8f7188c2cb42d73702f7c6e62d234e80a759efa7

Observation 4e696492-252e-4194-bdfa-935cdea04107 · outbound

This paper cites wav2vec-S: Adapting pre-trained speech models for streaming,.

On-device Streaming Discrete Speech Units wav2vec-S: Adapting pre-trained speech models for streaming,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:40.362670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.536603Z digest=sha256:020803f9ff78a997ab12ad62bf8a315266d35f3e047369f9c6509ac0c5adb9d6

Observation 1a26bfc3-6eba-4726-8fad-12891fd93bcf · outbound

This paper cites DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden- unit BERT,.

On-device Streaming Discrete Speech Units DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden- unit BERT,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:40.230176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.621595Z digest=sha256:c30944927c29046ea4cb20aebada85afa79e9ab1247586528eae93ff753d36ac

Observation aeba3496-4a71-412e-a4af-ef70cab8972a · outbound

This paper cites FitHuBERT: Going thinner and deeper for knowledge distillation of speech self-supervised learn- ing,.

On-device Streaming Discrete Speech Units FitHuBERT: Going thinner and deeper for knowledge distillation of speech self-supervised learn- ing,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:40.040098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.700430Z digest=sha256:51744cb594fda581072f463ac4cb727d5eab8c65eca9c83b6334a3e335d543fe

Observation 3a370e26-a96a-4a7e-858e-ce90a52634d1 · outbound

This paper cites Adaptive compression of supervised and self-supervised models for green speech recognition,.

On-device Streaming Discrete Speech Units Adaptive compression of supervised and self-supervised models for green speech recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:39.892453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.813703Z digest=sha256:2bdbc6368900a055fe80dafebf8b161df33551384d3c05fce69fa65175254390

Observation 5eceb7eb-0847-4ffc-a6f0-8c450b93715f · outbound

This paper cites Soundstream: An end- to-end neural audio codec,.

On-device Streaming Discrete Speech Units Soundstream: An end- to-end neural audio codec,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:39.734384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:37.916027Z digest=sha256:c5072b25014212b4428d8470a010bc1dc8b69c80f90b7b57ca36be555c4fc997

Observation 050abf2a-59eb-45ec-82d8-03ace57cc6e9 · outbound

This paper cites High fidelity neural audio compression,.

On-device Streaming Discrete Speech Units High fidelity neural audio compression,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:38.031930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:38.031930Z digest=sha256:c7a2b5eb30fe1404e3ccf84121a984fecdbcb5b69e710bba83712f59c1827276

Observation 8ea671ed-c414-4979-ac6d-3f853db954b4 · outbound

This paper cites ESPnet-Codec: Comprehensive train- ing and evaluation of neural codecs for audio, music, and speech,.

On-device Streaming Discrete Speech Units ESPnet-Codec: Comprehensive train- ing and evaluation of neural codecs for audio, music, and speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:39.583033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:38.096955Z digest=sha256:125a3489a8e131fed46a95e18a2cc63419771a1368b5161fa2dced6f25134265

Observation 8154b962-78ad-4326-af30-5f2a6dab1602 · outbound

This paper cites Neural discrete representa- tion learning,.

On-device Streaming Discrete Speech Units Neural discrete representa- tion learning,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:39.421530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:38.167171Z digest=sha256:8c1cb8dd4368e877725e4def345f22198461e07009416d3780d67f367be108e2

Observation 6eaa95f1-0deb-4bcf-a909-f9a53f1c2f9b · outbound

This paper cites Distilling hubert with lstms via decoupled knowledge distillation,.

On-device Streaming Discrete Speech Units Distilling hubert with lstms via decoupled knowledge distillation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:39.265094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:38.285199Z digest=sha256:a3c31e460ef9144339340c89794e4fdca4d8487cb69a1ca30a3788e4b00be57d

Observation 70bdf51d-a647-4d23-a17d-dd84c784539e · outbound

This paper cites Knowledge distillation from self- supervised representation learning model with discrete speech units for any-to-any streaming voice conversion,.

On-device Streaming Discrete Speech Units Knowledge distillation from self- supervised representation learning model with discrete speech units for any-to-any streaming voice conversion,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:39.098412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:38.379141Z digest=sha256:3a7fdd2bbc893a918387b155ff7197d9667bb0ae1b02ac3b6418235de11996d5

Observation d1f994b4-9f6f-420d-be44-2b1665970710 · outbound

This paper cites Speechtokenizer: Unified speech tokenizer for speech language models,.

On-device Streaming Discrete Speech Units Speechtokenizer: Unified speech tokenizer for speech language models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:38.958854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:38.448545Z digest=sha256:103556a494c813e1bc554a992b3bb366c01389399679bac6771defead35d9cc3

Observation de6e815b-53b2-4c54-8b57-1be42abe5d1f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

On-device Streaming Discrete Speech Units Moshi: a speech-text foundation model for real-time dialogue

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:38.525748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:38.525748Z digest=sha256:e6ca81383ed03a02eb7c35a1f0e5f1e8ee1fb9cee751f5123aaef590d850088d

Observation 82907e2b-c766-4208-846a-79b7f705f2a1 · outbound

This paper cites We denote various attention window configurations as the number of left, center, and right frames, i.e., [l, c = 1 , r].

On-device Streaming Discrete Speech Units We denote various attention window configurations as the number of left, center, and right frames, i.e., [l, c = 1 , r]

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:35:38.820837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:38.605780Z digest=sha256:0d394887f2b7dcdc27a21fe3e2de4100473509f4768ac431af0768329b2f07da

Pith citing papers

Observation 3b3040d7-d133-4d99-99b7-a3186741704a · inbound

On-device Streaming Discrete Speech Units cites this paper.

On-device Streaming Discrete Speech Units On-device Streaming Discrete Speech Units

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:33.841116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:33.841116Z digest=sha256:a95abd79f1fffc429c27ed13ebff8656669500259d10eb0bcce80263dfb0c3af

Observation 908b712e-f978-4af9-8ab2-a9e54b3f0ad2 · inbound

WhisperRT -- Turning Whisper into a Causal Streaming Model cites this paper.

WhisperRT -- Turning Whisper into a Causal Streaming Model On-device Streaming Discrete Speech Units

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:21:53.598627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T22:16:51.917336Z digest=sha256:631ffe76c90d8a6b0287d27010dd366023a3e97b90f767e27245f45fa6c2f5a8